Electronic device for outputting quality of image as score and control method thereof
The electronic device uses neural networks to identify sub-regions based on pixel importance and summed values for improved I/VQA accuracy and efficiency, addressing the limitations of existing techniques by focusing on user-interest areas and reducing computational burdens.
Patent Information
- Application Number
- PCT/KR2025/001793
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-12
- Filing Date
- 2025-02-06
- Publication Date
- 2025-10-02
Smart Images

Figure KR2025001793_02102025_PF_FP_ABST
Abstract
Description
Electronic device for outputting image quality as a score and its control method
[0001] The present disclosure relates to an electronic device and a control method thereof, and more particularly, to an electronic device that outputs the quality of an image as a score and a control method thereof.
[0002] Advances in electronic devices and multimedia technology have led to a surge in consumer video service usage, and consequently, their expectations for Quality of Experience (QoE) are also rising. Consumers are the ultimate judges of video quality, and providers need to predict the quality consumers will perceive to enhance their QoE.
[0003] Accordingly, robust Image / Video Quality Assessment (I / VQA) techniques are being developed to provide high-quality video services to consumers.
[0004] According to one embodiment of the present disclosure for achieving the above object, an electronic device includes at least one memory storing a first neural network model trained to output a saliency map for an image and a second neural network model trained to output a quality score of an image, and at least one processor connected to the at least one memory, wherein the at least one processor obtains a saliency map including saliency values of each of a plurality of pixels included in the first image through the first neural network model based on a first image, identifies a plurality of first sub-regions respectively corresponding to a plurality of regions included in the first image based on the saliency map, and obtains a quality score for the first image through the second neural network model based on the identified plurality of first sub-regions, and the quality score may be based on a plurality of first quality scores respectively corresponding to the identified plurality of first sub-regions.
[0005] Additionally, the at least one processor can identify a portion of each of the plurality of regions as a first sub-region of the plurality of first sub-regions corresponding to the plurality of regions based on importance values of pixels included in each of the plurality of regions.
[0006] And, the at least one processor determines, for each of the plurality of regions, a plurality of first sum values corresponding to each pixel included in each region, and each first sum value of the plurality of first sum values is a sum of an importance value of each pixel and importance values of pixels surrounding each pixel, and can identify a sub-region of each region including a reference pixel corresponding to the largest first sum value among the plurality of first sum values as a first sub-region corresponding to each region.
[0007] In addition, the second neural network model is further trained to output the plurality of first quality scores based on the plurality of first sub-regions and the plurality of second sum values input to the second neural network model, and the at least one processor determines a plurality of second sum values respectively corresponding to the plurality of first sub-regions, and each second sum value of the plurality of second sum values is a sum of importance values of pixels included in a first sub-region corresponding to the plurality of first sub-regions, and the quality score for the first image can be obtained through the second neural network model based on the plurality of identified first sub-regions and the plurality of second sum values.
[0008] And, the at least one processor can obtain a plurality of importance maps corresponding to each of the plurality of frames through the first neural network model, and identify one of the plurality of frames as the first image based on the plurality of importance maps output by the first neural network model.
[0009] In addition, the at least one processor determines a plurality of third sum values corresponding to each of the plurality of regions, and each third sum value of the plurality of third sum values is a sum of importance values of pixels included in a corresponding region of the plurality of regions, and an region of the plurality of regions corresponding to a third sum value greater than or equal to a preset size among the plurality of third sum values can be identified as an additional first sub-region.
[0010] And, the at least one processor determines a plurality of third sum values corresponding to each of the plurality of regions, and each third sum value of the plurality of third sum values is a sum of importance values of pixels included in a corresponding region of the plurality of regions, and can update the size of the region of the plurality of regions based on the plurality of third sum values.
[0011] In addition, the first image is a first frame among a plurality of frames, the second image is a second frame among the plurality of frames immediately following the first frame among the plurality of frames, and the at least one processor determines a motion vector based on the first image and the second image, identifies a plurality of second sub-regions of the second image corresponding to the plurality of first sub-regions and the motion vector, and obtains a quality score of the second image through the second neural network model based on the identified plurality of second sub-regions.
[0012] And, the at least one processor can perform at least one of upscaling and noise removal on the first image based on the quality score for the first image.
[0013] In addition, the first neural network model can learn a plurality of first sample images and a plurality of sample importance maps corresponding to the plurality of first sample images, respectively, and the second neural network model can learn a plurality of second sample images and a plurality of sample scores corresponding to the plurality of second sample images, respectively.
[0014] And, each of the plurality of sample importance maps may be based on a plurality of users' gazes for a first sample image corresponding to each of the plurality of sample importance maps, and each of the plurality of sample scores may be based on a plurality of users' scores for a second sample image corresponding to each of the plurality of sample scores.
[0015] Meanwhile, according to one embodiment of the present disclosure, a control method of an electronic device including at least one processor, storing a first neural network model trained to output a saliency map for an image and a second neural network model trained to output a quality score of the image, the method including: obtaining, by the at least one processor, a saliency map including saliency values of each of a plurality of pixels included in the first image through the first neural network model based on a first image; identifying a plurality of first sub-regions respectively corresponding to a plurality of regions included in the first image based on the saliency map; and obtaining a quality score for the first image through the second neural network model based on the identified plurality of first sub-regions, wherein the quality score may be based on a plurality of first quality scores respectively corresponding to the identified plurality of first sub-regions.
[0016] In addition, the step of identifying the plurality of first sub-regions may identify, by the at least one processor, a portion of each of the plurality of regions as a first sub-region of the plurality of first sub-regions corresponding to the plurality of regions based on importance values of pixels included in each of the plurality of regions.
[0017] And, the step of identifying the first sub-region may include, by the at least one processor, determining, for each of the plurality of regions, a plurality of first sum values corresponding to each pixel included in each region, and each first sum value of the plurality of first sum values being the sum of the importance value of each pixel and the importance values of pixels surrounding each pixel, and identifying a sub-region of each region including a reference pixel corresponding to the largest first sum value among the plurality of first sum values as a first sub-region corresponding to each region.
[0018] In addition, the second neural network model is further trained to output the plurality of first quality scores based on the plurality of first sub-regions and the plurality of second sum values input to the second neural network model, and the at least one processor determines a plurality of second sum values each corresponding to the plurality of first sub-regions, and each of the plurality of second sum values is a sum of importance values of pixels included in a first sub-region corresponding to the plurality of first sub-regions, and the step of obtaining the quality score for the first image may obtain the quality score for the first image through the second neural network model based on the plurality of identified first sub-regions and the plurality of second sum values.
[0019] In addition, the method may further include a step of obtaining a plurality of importance maps corresponding to each of a plurality of frames through the first neural network model, and a step of identifying one of the plurality of frames as the first image based on the plurality of importance maps.
[0020] In addition, the method further includes a step of obtaining a plurality of third sum values corresponding to each of the plurality of regions by adding up the importance values of pixels included in each of the plurality of regions, and the step of identifying the plurality of first sub-regions may identify an additional first sub-region in an area corresponding to a third sum value greater than or equal to a preset size among the plurality of third sum values.
[0021] In addition, the method may further include a step of obtaining a plurality of third sum values corresponding to each of the plurality of regions by adding up the importance values of pixels included in each of the plurality of regions, and a step of updating the sizes of the plurality of regions based on the plurality of third sum values.
[0022] In addition, the first image may be one of a plurality of frames, and the control method may further include a step of obtaining a motion vector based on the first image and a second image immediately following the first image among the plurality of frames, a step of identifying a plurality of second sub-regions corresponding to the second image based on the plurality of first sub-regions and the motion vector, and a step of obtaining a quality score of the second image through the second neural network model based on the plurality of identified second sub-regions.
[0023] And, the step of performing at least one of upscaling and noise removal on the first image based on the quality of the first image may be further included.
[0024] In addition, the first neural network model may be a model obtained by learning a plurality of first sample images and a plurality of sample importance maps corresponding to the plurality of first sample images, respectively, and the second neural network model may be a model obtained by learning a plurality of second sample images and a plurality of sample scores corresponding to the plurality of second sample images, respectively.
[0025] And, each of the plurality of sample importance maps may be obtained based on the gazes of the plurality of users for the first sample image corresponding to each of the plurality of sample importance maps, and each of the plurality of sample scores may be obtained based on the scores of the plurality of users for the second sample image corresponding to each of the plurality of sample scores.
[0026] Figure 1 is a diagram for explaining Image / Video Quality Assessment (I / VQA) to help understand the present disclosure.
[0027] FIG. 2 is a block diagram showing the configuration of an electronic device according to an embodiment of the present disclosure.
[0028] FIG. 3 is a block diagram showing a detailed configuration of an electronic device according to an embodiment of the present disclosure.
[0029] FIG. 4 is a drawing for explaining the difference in the identification method of a sub-region according to one embodiment of the present disclosure.
[0030] FIG. 5 is a diagram illustrating a method for identifying a sub-region based on an importance map according to an embodiment of the present disclosure.
[0031] FIG. 6 is a diagram illustrating an overall method for identifying the quality of an image according to one embodiment of the present disclosure.
[0032] FIG. 7 is a diagram illustrating an operation of identifying multiple sub-regions in one region according to one embodiment of the present disclosure.
[0033] FIG. 8 is a drawing for explaining an operation of differentiating the sizes of multiple areas in an image according to one embodiment of the present disclosure.
[0034] FIG. 9 is a flowchart for explaining a control method of an electronic device according to an embodiment of the present disclosure.
[0035] The purpose of the present disclosure is to provide an electronic device and a control method thereof for performing Image / Video Quality Assessment (I / VQA) while reducing computational burden and taking into more consideration the user's area of interest.
[0036] It should be understood that the various embodiments and terms used in this document are not intended to limit the technical features described in this document to specific embodiments, but rather to include various modifications, equivalents, or substitutes of the embodiments.
[0037] In connection with the description of the drawings, similar reference numerals may be used for similar or related components.
[0038] The singular form of a noun corresponding to an item may include one or more items, unless the context clearly indicates otherwise.
[0039] In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" may include any one of the items listed together in that phrase, or all possible combinations thereof.
[0040] Terms such as "first," "second," or "first" or "second" may be used simply to distinguish one component from another and do not qualify the components in any other respect (e.g., importance or order).
[0041] When a component (e.g., a first component) is referred to as being “coupled” or “connected” to another component (e.g., a second component), with or without the terms “functionally” or “communicatively,” it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.
[0042] The terms “include” or “have” are intended to specify the presence of a feature, number, step, operation, component, part or combination thereof described in this document, but do not preclude the presence or addition of one or more other features, numbers, steps, operations, components, parts or combinations thereof.
[0043] When a component is said to be “connected,” “coupled,” “supported,” or “in contact with” another component, this includes not only cases where the components are directly connected, coupled, supported, or in contact, but also cases where the components are indirectly connected, coupled, supported, or in contact through a third component.
[0044] When we say that a component is “on” another component, this includes not only cases where the component is in contact with the other component, but also cases where there is another component between the two components.
[0045] The term “and / or” includes any combination of a plurality of related described elements or any one of a plurality of related described elements.
[0046] The operating principle and embodiments of the present invention will be described with reference to the attached drawings below.
[0047] Figure 1 is a diagram for explaining Image / Video Quality Assessment (I / VQA) to help understand the present disclosure.
[0048] Image / Video Quality Assessment (I / VQA) techniques can be used to predict video quality. I / VQA techniques can include Full-Reference Video Quality Assessment (FR-I / VQA), which analyzes the differences between original and degraded images, and No-Reference Video Quality Assessment (NR-I / VQA), which assesses quality based solely on degraded images.
[0049] Specific examples of FR-I / VQA techniques include PSNR, SSIM, MS-SSIM, FSIM, and MAD, but research on NR-I / VQA techniques is actively being conducted due to the limitation of requiring original images.
[0050] The initial NR-I / VQA technique was a method to predict quality for specific distortions based on hand-crafted features. Because it utilized hand-crafted features, it was successful in certain areas, but had limitations for in-the-wild video.
[0051] Recently, with the development of Deep Neural Networks, NR-I / VQA techniques have also developed significantly, but end-to-end learning has become difficult due to running time and memory issues as the resolution increases.
[0052] To address this issue, methods utilizing pre-trained models, naive cropping, and resizing have been studied. However, cropping and resizing methods incur significant feature loss, while methods utilizing pre-trained models suffer from accuracy loss due to the inability to fully train the model.
[0053] Subsequently developed FAST-I / VQA, DOVER, and FASTER-I / VQA introduced the concept of fragments utilizing Grid Mini Patches (GMS). For example, as illustrated in Figure 1, the new technique divides an image into multiple regions (grids), identifies sub-regions (fragments) within each of the multiple regions, and identifies quality using only the identified sub-regions from the multiple regions. Accordingly, processing time is reduced, end-to-end learning is enabled, and performance is improved, enabling effective NR-I / VQA at all resolutions.
[0054] However, since random sampling is performed for each sub-region when configuring multiple regions, performance may vary depending on the selected sample. If meaningless samples are selected, NR-I / VQA performance may deteriorate. In particular, due to random sampling, robustness is reduced, and there is a problem that all sub-regions have the same weighting, even though the human eye may react differently to each region.
[0055] FIG. 2 is a block diagram showing the configuration of an electronic device (100) according to one embodiment of the present disclosure.
[0056] The electronic device (100) is a device that identifies the quality of an image and can be implemented as a TV, desktop PC, laptop, video wall, LFD (large format display), Digital Signage, DID (Digital Information Display), projector display, smartphone, tablet PC, etc.
[0057] However, it is not limited thereto, and the electronic device (100) may be any device that identifies the quality of an image.
[0058] According to FIG. 2, the electronic device (100) includes a memory (110) and a processor (120). However, the present invention is not limited thereto, and the electronic device (100) may be implemented in a form in which some components are excluded.
[0059] Memory (110) may refer to hardware that stores information such as data in an electrical or magnetic form so that a processor (120) or the like can access it. To this end, memory (110) may be implemented as at least one piece of hardware from among non-volatile memory, volatile memory, flash memory, hard disk drive (HDD), solid state drive (SSD), RAM, ROM, etc.
[0060] At least one instruction required for the operation of the electronic device (100) or processor (120) may be stored in the memory (110). Here, the instruction is a code unit that instructs the operation of the electronic device (100) or processor (120), and may be written in machine language, which is a language that a computer can understand.
[0061] The memory (110) may store data in bit or byte units that can represent characters, numbers, images, etc. For example, neural network models may be stored in the memory (110).
[0062] The memory (110) is accessed by the processor (120), and reading / writing / modifying / deleting / updating instructions, instruction sets, or data can be performed by the processor (120).
[0063] The processor (120) controls the overall operation of the electronic device (100). Specifically, the processor (120) is connected to each component of the electronic device (100) and can control the overall operation of the electronic device (100). For example, the processor (120) is connected to components such as a memory (110), a display (not shown), etc. and can control the operation of the electronic device (100).
[0064] The processor (120) may be implemented with one or more processors. In this case, the one or more processors may include one or more of a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), an APU (Accelerated Processing Unit), a MIC (Many Integrated Core), a DSP (Digital Signal Processor), an NPU (Neural Processing Unit), a hardware accelerator, or a machine learning accelerator. The one or more processors may control one or any combination of other components of the electronic device (100) and perform operations related to communication or data processing. The one or more processors may execute one or more programs or instructions stored in the memory (110). For example, the one or more processors may perform a method according to an embodiment of the present disclosure by executing one or more instructions stored in the memory (110).
[0065] When a method according to an embodiment of the present disclosure includes a plurality of operations, the plurality of operations may be performed by one processor or by a plurality of processors. For example, when a first operation, a second operation, and a third operation are performed by a method according to an embodiment, the first operation, the second operation, and the third operation may all be performed by a first processor, or the first operation and the second operation may be performed by a first processor (e.g., a general-purpose processor) and the third operation may be performed by a second processor (e.g., an AI-specific processor). For example, a process of quantizing a neural network model according to an embodiment of the present disclosure may be performed by a general-purpose processor, and a process of learning or inferring the quantized neural network model may be performed by an AI-specific processor.
[0066] One or more processors may be implemented as a single core processor including one core, or may be implemented as one or more multicore processors including multiple cores (e.g., homogeneous multicore or heterogeneous multicore). When one or more processors are implemented as a multicore processor, each of the multiple cores included in the multicore processor may include internal processor memory, such as cache memory or on-chip memory, and a common cache shared by the multiple cores may be included in the multicore processor. In addition, each of the multiple cores (or some of the multiple cores) included in the multicore processor may independently read and execute a program instruction for implementing a method according to an embodiment of the present disclosure, or all (or some) of the multiple cores may be linked to read and execute a program instruction for implementing a method according to an embodiment of the present disclosure.
[0067] When a method according to an embodiment of the present disclosure includes a plurality of operations, the plurality of operations may be performed by one core among the plurality of cores included in a multi-core processor, or may be performed by the plurality of cores. For example, when a first operation, a second operation, and a third operation are performed by a method according to an embodiment, the first operation, the second operation, and the third operation may all be performed by a first core included in the multi-core processor, or the first operation and the second operation may be performed by a first core included in the multi-core processor, and the third operation may be performed by a second core included in the multi-core processor.
[0068] In embodiments of the present disclosure, one or more processors may refer to a system on a chip (SoC) in which one or more processors and other electronic components are integrated, a single-core processor, a multi-core processor, or a core included in a single-core processor or a multi-core processor, wherein the core may be implemented as a CPU, a GPU, an APU, a MIC, a DSP, an NPU, a hardware accelerator, or a machine learning accelerator, but the embodiments of the present disclosure are not limited thereto. However, for convenience of explanation, the operation of the electronic device (100) is described below using the expression processor (120).
[0069] The processor (120) may obtain a saliency map corresponding to the saliency value of each of a plurality of pixels included in the first image through the first neural network model. For example, the processor (120) may input the first image into the first neural network model to obtain a saliency map indicating the saliency value of each of a plurality of pixels included in the first image. Here, the first neural network model may be a model obtained by learning a plurality of first sample images and a plurality of sample saliency maps respectively corresponding to the plurality of first sample images. Each of the plurality of sample saliency maps may be obtained based on the gazes of a plurality of users with respect to the first sample images respectively corresponding to the plurality of sample saliency maps.
[0070] In other words, a significance map can contain information on the degree to which users' gaze is directed at each pixel in an image. For example, when a user views an image, the area where their gaze is primarily directed may have a high pixel value in the significance map.
[0071] The processor (120) can divide the first image into a plurality of regions (grids). For example, the processor (120) can divide the first image into a plurality of 7x7 regions, and the plurality of regions can all be of the same size, and each region can be referred to as a grid.
[0072] However, the present invention is not limited thereto, and the processor (120) may divide the first image into a plurality of regions based on at least one of the resolution, screen ratio, type, or importance map of the first image. Alternatively, the processor (120) may divide the first image into a plurality of regions such that at least some of the regions have different sizes.
[0073] The processor (120) may identify a plurality of first sub-regions (fragments) from each of the plurality of regions included in the first image based on the importance map. For example, the processor (120) may identify a portion of each of the plurality of regions as a first sub-region based on the importance values of pixels included in each of the plurality of regions.
[0074] For convenience of explanation, a method of identifying a first sub-region of one of a plurality of regions will first be described. The processor (120) may obtain a first sum value corresponding to each pixel by adding up the importance values of each pixel and the surrounding pixels of each pixel included in one of the plurality of regions. For example, assuming that the size of one of the plurality of regions has a resolution of 200×200 and the size of the region including each pixel and the surrounding pixels of each pixel has a resolution of 32×32, the processor (120) may obtain a first sum value for each of 169×169 pixels. Due to the size of the region to be summed, the first sum value may be obtained for 169×169 pixels rather than 200×200 pixels. The processor (120) may identify a pixel corresponding to the largest first sum value among the first sum values corresponding to each pixel as a reference pixel, and may identify a region including the surrounding pixels of the reference pixel as a first sub-region corresponding to one of the plurality of regions. The processor (120) can apply this method to each of the plurality of regions to identify a plurality of first sub-regions corresponding to each of the plurality of regions. Here, each of the plurality of first sub-regions may be an area where the user's gaze is most drawn in the corresponding region (grid).
[0075] The processor (120) can identify the quality of the first image based on a plurality of first scores corresponding to a plurality of first sub-regions obtained through the second neural network model. For example, the processor (120) can input a plurality of first sub-regions into a second neural network model to obtain a plurality of first scores. Here, the second neural network model may be a model obtained by learning a plurality of second sample images and a plurality of sample scores corresponding to each of the plurality of second sample images. Each of the plurality of sample scores may be obtained based on a plurality of user scores for the second sample images corresponding to each of the plurality of sample scores.
[0076] In addition, the processor (120) can identify the quality of the first image based on a plurality of first scores. For example, the processor (120) can identify the quality of the first image by averaging or summing the plurality of first scores. However, the present invention is not limited thereto, and the processor (120) can also identify the quality of the first image using only the first scores within a preset range among the plurality of first scores.
[0077] The processor (120) may obtain a plurality of second sum values corresponding to each of the plurality of first sub-regions by adding up the importance values of the pixels included in each of the plurality of first sub-regions, and input the plurality of first sub-regions and the plurality of second sum values into a second neural network model to obtain a plurality of first scores. Here, the second neural network model may be a model trained to further consider not only the plurality of first sub-regions but also the plurality of second sum values obtained by adding up the importance values of the pixels included in each of the plurality of first sub-regions. Through this operation, the degree to which the user's gaze goes to each region may be further considered in the process of identifying the quality of the first image.
[0078] The processor (120) may obtain a plurality of importance maps corresponding to a plurality of frames through the first neural network model, and identify one of the plurality of frames as the first image based on the plurality of importance maps. For example, the processor (120) may input a plurality of frames into the first neural network model, obtain a plurality of importance maps corresponding to the plurality of frames, and identify one of the plurality of frames as the first image based on the plurality of importance maps. For example, the processor (120) may input a plurality of frames into the first neural network model, obtain a plurality of importance maps corresponding to the plurality of frames, add up the importance values included in each of the plurality of importance maps, and identify the frame with the largest sum value as the first image, and may also identify the quality of the first image as the quality of the remaining frames.
[0079] Meanwhile, although the above description describes identifying one first sub-region in one region, it is not limited thereto. For example, the processor (120) may obtain multiple third sum values corresponding to each of the multiple regions by adding up the importance values of pixels included in each of the multiple regions, and identify additional first sub-regions in regions corresponding to a third sum value greater than or equal to a preset first size among the multiple third sum values. That is, the processor (120) may identify multiple first sub-regions in regions corresponding to a third sum value greater than or equal to a preset first size among the multiple third sum values.
[0080] The processor (120) may obtain a plurality of third sum values corresponding to each of the plurality of regions by adding up the importance values of pixels included in each of the plurality of regions, and update the sizes of the plurality of regions based on the plurality of third sum values. For example, the processor (120) may reduce the size of a region whose value is greater than or equal to the average value of the plurality of third sum values, and may enlarge the size of a region whose value is less than or equal to the average value of the plurality of third sum values.
[0081] Meanwhile, the above has described an embodiment in which one frame among a plurality of frames is identified as a first image and the quality of the first image is identified based on the quality of the remaining frames, but the present invention is not limited thereto. For example, the first image is one of a plurality of frames, and the processor (120) obtains a motion vector based on the first image and a second image immediately after the first image among the plurality of frames, identifies a plurality of second sub-regions corresponding to the second image based on the plurality of first sub-regions and the motion vectors, inputs the plurality of second sub-regions into a second neural network model to obtain a plurality of second scores, and identifies the quality of the second image based on the plurality of second scores.
[0082] When the quality of the first image is identified in the above manner, the processor (120) can perform at least one of upscaling and noise removal on the first image based on the quality of the first image.
[0083] The functions related to artificial intelligence according to the present disclosure can be operated through a processor (120) and a memory (110).
[0084] The processor (120) may be composed of one or more processors. In this case, one or more processors may be a general-purpose processor such as a CPU, AP, DSP, etc., a graphics-only processor such as a GPU or VPU (Vision Processing Unit), or an artificial intelligence-only processor such as an NPU.
[0085] One or more processors are controlled to process input data according to predefined operating rules or artificial intelligence models stored in the memory (110). Alternatively, if one or more processors are dedicated artificial intelligence processors, the dedicated artificial intelligence processors may be designed with a hardware structure specialized for processing a specific artificial intelligence model. The predefined operating rules or artificial intelligence models are characterized by being created through learning.
[0086] Here, "created through learning" means that a basic artificial intelligence model is learned using a learning algorithm using a plurality of learning data, thereby creating a predefined set of operating rules or an artificial intelligence model set to perform a desired characteristic (or purpose). This learning may be performed on the device itself on which the artificial intelligence according to the present disclosure is performed, or may be performed through a separate server and / or system. Examples of learning algorithms include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.
[0087] An artificial intelligence model may be composed of multiple neural network layers. Each of the multiple neural network layers has multiple weight values and performs neural network operations by calculating the results of previous layers and the multiple weights. The multiple weights of the multiple neural network layers can be optimized based on the learning results of the artificial intelligence model. For example, the multiple weights may be updated during the learning process to reduce or minimize the loss or cost values obtained by the artificial intelligence model.
[0088] Artificial neural networks may include deep neural networks (DNNs), such as, but not limited to, convolutional neural networks (CNNs), deep neural networks (DNNs), recurrent neural networks (RNNs), restricted boltzmann machines (RBMs), deep belief networks (DBNs), bidirectional recurrent deep neural networks (BRDNNs), generative adversarial networks (GANs), or deep Q-networks.
[0089] FIG. 3 is a block diagram showing a detailed configuration of an electronic device (100) according to one embodiment of the present disclosure.
[0090] The electronic device (100) may include a memory (110) and a processor (120). In addition, according to FIG. 3, the electronic device (100) may further include a display (130), a communication interface (140), a user interface (150), a camera (160), a microphone (170), and a speaker (180). Among the components illustrated in FIG. 3, a detailed description of the overlapping parts with the components illustrated in FIG. 2 will be omitted.
[0091] The display (130) is a component that displays content and can be implemented as a variety of displays such as an LCD (Liquid Crystal Display), an OLED (Organic Light Emitting Diodes) display, a PDP (Plasma Display Panel), etc. The display (130) may also include a driving circuit, a backlight unit, etc. that can be implemented as a form such as an a-si TFT, an LTPS (low temperature poly silicon) TFT, an OTFT (organic TFT), etc. Meanwhile, the display (130) may be implemented as a touch screen combined with a touch sensor, a flexible display, a 3D display, etc.
[0092] The communication interface (140) is a configuration that performs communication with various types of external devices according to various types of communication methods. For example, the electronic device (100) can perform communication with a server through the communication interface (140).
[0093] The communication interface (140) may include a Wi-Fi module, a Bluetooth module, an infrared communication module, a wireless communication module, etc. Here, each communication module may be implemented in the form of at least one hardware chip.
[0094] Wi-Fi and Bluetooth modules communicate via Wi-Fi and Bluetooth, respectively. When using a Wi-Fi or Bluetooth module, connection information, such as the SSID and session key, is first transmitted and received. This information is then used to establish a communication connection before various other information can be transmitted and received. Infrared communication modules communicate using infrared data association (IrDA) technology, which uses infrared, a wavelength between visible light and millimeter waves, to wirelessly transmit data over short distances.
[0095] In addition to the above-described communication method, the wireless communication module may include at least one communication chip that performs communication according to various wireless communication standards such as zigbee, 3G (3rd Generation), 3GPP (3rd Generation Partnership Project), LTE (Long Term Evolution), LTE-A (LTE Advanced), 4G (4th Generation), 5G (5th Generation), etc.
[0096] Alternatively, the communication interface (140) may include a wired communication interface such as HDMI, DP, Thunderbolt, USB, RGB, D-SUB, DVI, etc.
[0097] In addition, the communication interface (140) may include at least one of a LAN (Local Area Network) module, an Ethernet module, or a wired communication module that performs communication using a pair cable, a coaxial cable, or an optical fiber cable.
[0098] The user interface (150) may be implemented with buttons, a touch pad, a mouse, a keyboard, etc., or may be implemented with a touch screen capable of performing both display and operation input functions. Here, the buttons may be various types of buttons, such as mechanical buttons, touch pads, wheels, etc., formed on any area of the front, side, or back of the main body of the electronic device (100).
[0099] The camera (160) is configured to capture still images or moving images. The camera (160) can capture still images at a specific point in time, but can also capture still images continuously.
[0100] The camera (160) includes a lens, a shutter, an aperture, a solid-state image sensor, an AFE (Analog Front End), and a TG (Timing Generator). The shutter controls the time at which light reflected from a subject enters the camera (160), and the aperture mechanically increases or decreases the size of the opening through which light enters to control the amount of light incident on the lens. When the solid-state image sensor accumulates light reflected from a subject as a photocharge, the image generated by the photocharge is output as an electrical signal. The TG outputs a timing signal for reading out pixel data of the solid-state image sensor, and the AFE samples and digitizes the electrical signal output from the solid-state image sensor.
[0101] The microphone (170) is configured to receive sound and convert it into an audio signal. The microphone (170) is electrically connected to the processor (120) and can receive sound under the control of the processor (120).
[0102] For example, the microphone (170) may be formed as an integrated unit integrated into the upper side, front side, side side, etc. of the electronic device (100). Alternatively, the microphone (170) may be provided in a remote control, etc., separate from the electronic device (100). In this case, the remote control may receive sound through the microphone (170) and provide the received sound to the electronic device (100).
[0103] The microphone (170) may include various configurations such as a microphone that collects sound in analog form, an amplifier circuit that amplifies the collected sound, an A / D conversion circuit that samples the amplified sound and converts it into a digital signal, and a filter circuit that removes noise components from the converted digital signal.
[0104] Meanwhile, the microphone (170) may be implemented in the form of a sound sensor, and any method may be used as long as it has a configuration capable of collecting sound.
[0105] The speaker (180) is a component that outputs various audio data processed by the processor (120) as well as various notification sounds and voice messages.
[0106] As described above, since the electronic device (100) identifies the quality of the image by identifying a plurality of first sub-regions from a plurality of regions using the importance map for the image, the quality identification performance of the image can be improved.
[0107] In addition, since the electronic device (100) identifies the quality of the image by a plurality of first sub-regions, which are parts of each of a plurality of regions rather than the entire image, the quality of the image can be identified with reduced processing time.
[0108] Meanwhile, while the first neural network model has been described above as outputting a significance map corresponding to the input image, this is not a limitation. For example, the first neural network model may be trained to output information on multiple first sub-regions corresponding to the input image. In this case, the processor (120) may input the first image into the first neural network model to identify multiple first sub-regions.
[0109] In addition, although the processor (120) has been described above as inputting multiple first sub-regions into the second neural network model, the present invention is not limited thereto. For example, the processor (120) may input a single image including multiple first sub-regions into the second neural network model to obtain a single score, and identify the quality of the first image based on the single score.
[0110] Hereinafter, the operation of the electronic device (100) will be described in more detail with reference to FIGS. 4 to 8. For convenience of explanation, individual embodiments are described in FIGS. 4 to 8. However, the individual embodiments of FIGS. 4 to 8 may be implemented in any combination.
[0111] FIG. 4 is a drawing for explaining the difference in the identification method of a sub-region according to one embodiment of the present disclosure.
[0112] The processor (120) can divide the first image into a plurality of regions (grids) and identify a plurality of first sub-regions (fragments) from each of the plurality of regions based on a significance map corresponding to the first image.
[0113] For example, as illustrated in FIG. 4, the processor (120) can divide the first image (original image) into a plurality of 7×7 regions, and identify a plurality of first sub-regions (rectangular regions with thin solid lines) from each of the plurality of regions based on an importance map corresponding to the first image.
[0114] In the past, the first image was divided into multiple regions and the first sub-region was identified in each region. However, in the past, the first sub-region was randomly selected, such as the rectangular region of the thick solid line in Fig. 4.
[0115] On the other hand, according to the present disclosure, since the first sub-area is selected based on the importance map, such as a rectangular area of a thin solid line, the degree to which the user's gaze is directed is reflected, and accordingly, the quality identification performance of the image can be improved.
[0116] FIG. 5 is a diagram illustrating a method for identifying a sub-region based on an importance map according to an embodiment of the present disclosure.
[0117] The processor (120) can identify a plurality of first sub-regions from a plurality of regions included in the first image based on an importance map corresponding to the first image.
[0118] The processor (120) may obtain a first sum value corresponding to each pixel by adding up the importance values of each pixel and the surrounding pixels of each pixel included in one of the plurality of regions, and may identify a region including a reference pixel and surrounding pixels of the reference pixel corresponding to the largest first sum value among the first sum values corresponding to each pixel as a first sub-region corresponding to one of the plurality of regions.
[0119] For example, as illustrated in FIG. 5, the processor (120) may divide the first image into a plurality of 7x7 regions, and identify a first sub-region starting from the first region (510) among the plurality of regions. For example, the processor (120) may identify a first pixel (520-1) included in the first region (510), and may obtain a first sum value corresponding to the first pixel (520-1) by adding up the importance values of the pixels included in the region (520-2) including the surrounding pixels of the first pixel (520-1). In addition, the processor (120) may identify a second pixel (530-1) included in the first region (510), and may obtain a first sum value corresponding to the second pixel (530-1) by adding up the importance values of the pixels included in the region (530-2) including the surrounding pixels of the second pixel (530-1). The processor (120) can obtain a first sum value for most of the pixels included in the first region (510) by repeating this method. The processor (120) can omit this operation for some of the pixels included in the first region (510). For example, the processor (120) can omit this operation for pixels located above and to the left of the first pixel (520-1). In FIG. 5, for convenience of explanation, the center of the region (520-2) including the surrounding pixels of the first pixel (520-1) is used as a reference, but the present invention is not limited thereto. For example, the pixel at the upper left of the region (520-2) may be the reference.
[0120] The processor (120) can identify a first sub-region corresponding to the first region (510) as an area (540) corresponding to the largest first sum value among a plurality of first sum values.
[0121] Although the method for identifying a first sub-region in a first region (510) is described in FIG. 5, the processor (120) can identify a first sub-region through the same operation in the remaining regions of the first image. That is, the processor (120) can identify a plurality of 7x7 first sub-regions from a plurality of 7x7 regions.
[0122] The size of the first sub-region may be smaller than the region including the first sub-region among the plurality of regions and may be a preset size. However, the present invention is not limited thereto, and the processor (120) may change the size of the first sub-region. For example, the processor (120) may change the size of the first sub-region based on at least one of the resolution of the first image, the hardware performance of the electronic device (100), or the resource status of the electronic device (100). Alternatively, the processor (120) may determine the size of the first sub-region corresponding to each of the plurality of regions based on the size of each of the plurality of regions included in the first image.
[0123] FIG. 6 is a diagram illustrating an overall method for identifying the quality of an image according to one embodiment of the present disclosure.
[0124] The processor (120) can identify a first image from the video. For example, as illustrated in FIG. 6, the processor (120) can identify a temporally intermediate frame among a plurality of frames included in the video (610) as the first image (620).
[0125] However, the present invention is not limited thereto, and the processor (120) may divide the video (610) into scenes, and identify a frame in the temporal middle of each scene as the first image (620). Alternatively, the processor (120) may input a plurality of frames included in the video (610) into a first neural network model, respectively, to obtain a plurality of importance maps corresponding to the plurality of frames, and identify one of the plurality of frames as the first image (620) based on the plurality of importance maps. For example, the processor (120) may input a plurality of frames included in the video (610) into a first neural network model, respectively, to obtain a plurality of importance maps corresponding to the plurality of frames, add up all importance values included in each of the importance maps to obtain a sum value corresponding to each of the plurality of frames, and identify the frame with the largest sum value as the first image (620).
[0126] The processor (120) can input the first image (620) into the first neural network model (630) to obtain an importance map indicating the importance value of each of a plurality of pixels included in the first image (620). The processor (120) can obtain, from the importance map, coordinates (640-1) of a plurality of first sub-regions each corresponding to a plurality of regions included in the first image (620) and a plurality of second sum values (640-2) each corresponding to the plurality of first sub-regions. Here, each of the plurality of second sum values (640-2) can be obtained by summing the importance values of pixels included in each of the corresponding first sub-regions.
[0127] The processor (120) may obtain a single image including a plurality of first sub-regions. Alternatively, the processor (120) may weight the second sum value (640-2) corresponding to each of the plurality of first sub-regions, thereby obtaining a single image including a plurality of first sub-regions to which weights are applied.
[0128] The processor (120) can acquire a plurality of images (650) by acquiring one image from each of the remaining frames among the plurality of frames based on the coordinates (640-1) of the plurality of first sub-regions. At this time, each of the plurality of images may have a weight applied based on the plurality of second sum values (640-2).
[0129] The processor (120) can input each of the plurality of images (650) into a second neural network model to obtain a score for each of the plurality of images (650), and accumulate the scores to identify the quality of the video (610). Here, the processor (120) can input each of the plurality of images (650) into the second neural network model, but is not limited thereto. For example, the processor (120) can input each of the plurality of first sub-regions included in each of the plurality of images (650) into the second neural network model.
[0130] In FIG. 6, for convenience of explanation, it is described that multiple scores corresponding to each of multiple frames are obtained, but it is not limited thereto. For example, the processor (120) may identify the quality of the video (610) using some or all of the multiple frames included in the video. For example, the processor (120) may identify only the quality of the first image (620) and identify the quality of the first image (620) as the quality of the video (610). In addition, the processor (120) may identify the number of frames for identifying the quality of the video (610) among the multiple frames included in the video (610) based on at least one of the resolution of the video (610), the hardware performance of the electronic device (100), or the resource status of the electronic device (100).
[0131] FIG. 7 is a diagram illustrating an operation of identifying multiple sub-regions in one region according to one embodiment of the present disclosure.
[0132] In the above, it has been described that one first sub-region is identified in one of the plurality of regions, but it is not limited thereto. For example, the processor (120) may obtain multiple third sum values corresponding to each of the plurality of regions by adding up the importance values of pixels included in each of the plurality of regions, and may identify an additional first sub-region in an region corresponding to a third sum value greater than or equal to a preset first size among the plurality of third sum values.
[0133] For example, as illustrated in FIG. 7, the processor (120) may obtain a third sum value corresponding to one area (710) by adding up the importance values of pixels included in one area (710) among a plurality of areas, and if the third sum value is greater than or equal to a preset first size, the processor (120) may identify a plurality of first sub-areas in one area (710).
[0134] A method for identifying multiple first sub-regions can be described in FIG. 5 by obtaining a first sum value corresponding to each pixel included in one region (710) and identifying it based on its size.
[0135] FIG. 8 is a drawing for explaining an operation of differentiating the sizes of multiple areas in an image according to one embodiment of the present disclosure.
[0136] In the above, it has been described that each of the multiple regions has the same shape and size, but this is not limited thereto. For example, the processor (120) may obtain multiple third sum values corresponding to each of the multiple regions by adding up the importance values of the pixels included in each of the multiple regions, and update the sizes of the multiple regions based on the multiple third sum values.
[0137] For example, as illustrated in FIG. 8, the processor (120) may further divide an area corresponding to a third sum value greater than or equal to a preset second size among a plurality of third sum values. The processor (120) may further identify a first sub-area in each of the divided areas. Here, the size of the first sub-area of the undivided area and the size of the first sub-area of the additionally divided area may be the same.
[0138] Alternatively, the processor (120) may reduce the size of an area that is greater than or equal to the average value of a plurality of third sum values, and may enlarge the size of an area that is less than or equal to the average value of a plurality of third sum values.
[0139] This behavior can improve image identification performance by allowing the quality of the first image to be identified through an area where the user's gaze is more focused.
[0140] FIG. 9 is a flowchart for explaining a control method of an electronic device according to an embodiment of the present disclosure.
[0141] First, a saliency map corresponding to the saliency value of each of a plurality of pixels included in a first image is obtained through a first neural network model trained to output a saliency map for an image (S910). Then, based on the saliency map, a plurality of first sub-regions (fragments) are identified from a plurality of regions (grids) included in the first image (S920). Then, the quality of the first image is identified based on a plurality of first scores corresponding to the plurality of first sub-regions obtained through a second neural network model trained to output the quality of the image as a score (S930).
[0142] Additionally, the step of identifying the first sub-region (S920) can identify a portion of each of the plurality of regions as the first sub-region based on the importance values of pixels included in each of the plurality of regions.
[0143] And, the step of identifying the first sub-region (S920) may be performed by adding up the importance values of each pixel and the surrounding pixels of each pixel included in one of the plurality of regions to obtain a first sum value corresponding to each pixel, and identifying a region including a reference pixel and surrounding pixels of the reference pixel corresponding to the largest first sum value among the first sum values corresponding to each pixel as a first sub-region corresponding to one of the plurality of regions.
[0144] In addition, the step of obtaining a plurality of second sum values corresponding to each of the plurality of first sub-regions by adding up the importance values of pixels included in each of the plurality of first sub-regions is further included, and the step (S930) of identifying the quality of the first image can obtain a plurality of first scores by inputting the plurality of first sub-regions and the plurality of second sum values into a second neural network model.
[0145] In addition, the method may further include a step of obtaining a plurality of importance maps corresponding to each of the plurality of frames through a first neural network model, and a step of identifying one of the plurality of frames as a first image based on the plurality of importance maps.
[0146] In addition, the step of obtaining a plurality of third sum values corresponding to each of the plurality of regions by adding up the importance values of pixels included in each of the plurality of regions is further included, and the step (S920) of identifying the plurality of first sub-regions can identify an additional first sub-region in an area corresponding to a third sum value greater than a preset size among the plurality of third sum values.
[0147] In addition, the method may further include a step of obtaining a plurality of third sum values corresponding to each of the plurality of regions by adding up the importance values of pixels included in each of the plurality of regions, and a step of updating the sizes of the plurality of regions based on the plurality of third sum values.
[0148] In addition, the first image is one of a plurality of frames, and the control method may further include a step of obtaining a motion vector based on the first image and a second image immediately after the first image among the plurality of frames, a step of identifying a plurality of second sub-regions corresponding to the second image based on the plurality of first sub-regions and the motion vector, a step of inputting the plurality of second sub-regions into a second neural network model to obtain a plurality of second scores, and a step of identifying a quality of the second image based on the plurality of second scores.
[0149] And, based on the quality of the first image, the step of performing at least one of upscaling and noise removal on the first image may be further included.
[0150] Additionally, the first neural network model may be a model obtained by learning a plurality of first sample images and a plurality of sample importance maps corresponding to the plurality of first sample images, respectively, and the second neural network model may be a model obtained by learning a plurality of second sample images and a plurality of sample scores corresponding to the plurality of second sample images, respectively.
[0151] And, each of the plurality of sample importance maps may be obtained based on the gazes of the plurality of users for the first sample image corresponding to each of the plurality of sample importance maps, and each of the plurality of sample scores may be obtained based on the scores of the plurality of users for the second sample image corresponding to each of the plurality of sample scores.
[0152] According to various embodiments of the present disclosure as described above, the electronic device can identify the quality of an image by identifying a plurality of first sub-regions from a plurality of regions using an importance map for the image, so that the quality identification performance of the image can be improved.
[0153] Additionally, since the electronic device identifies the quality of the image by multiple first sub-regions, which are each a part of multiple regions rather than the entire image, it is possible to identify the quality of the image with reduced processing time.
[0154] Meanwhile, according to a temporary example of the present disclosure, the various embodiments described above can be implemented as software including instructions stored in a machine-readable storage medium that can be read by a machine (e.g., a computer). The device is a device that can call instructions stored from the storage medium and operate according to the called instructions, and may include an electronic device (e.g., electronic device (A)) according to the disclosed embodiments. When an instruction is executed by a processor, the processor can perform a function corresponding to the instruction directly or by using other components under the control of the processor. The instruction may include code generated or executed by a compiler or interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' means that the storage medium does not contain a signal and is tangible, but does not distinguish between data being stored semi-permanently or temporarily in the storage medium.
[0155] Furthermore, according to one embodiment of the present disclosure, the method according to the various embodiments described above may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or online through an application store (e.g., Play Store™). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.
[0156] Furthermore, according to one embodiment of the present disclosure, the various embodiments described above may be implemented in a computer-readable recording medium or a similar device using software, hardware, or a combination thereof. In some cases, the embodiments described herein may be implemented by the processor itself. In a software implementation, embodiments such as the procedures and functions described herein may be implemented as separate software. Each software may perform one or more functions and operations described herein.
[0157] Meanwhile, computer instructions for performing processing operations of a device according to the various embodiments described above may be stored in a non-transitory computer-readable medium. The computer instructions stored in such a non-transitory computer-readable medium, when executed by a processor of a specific device, cause the specific device to perform processing operations in the device according to the various embodiments described above. A non-transitory computer-readable medium refers to a medium that permanently stores data and can be read by a device, rather than a medium that stores data for a short period of time, such as a register, cache, or memory. Specific examples of non-transitory computer-readable media may include a CD, DVD, hard disk, Blu-ray disk, USB, memory card, or ROM.
[0158] In addition, each of the components (e.g., modules or programs) according to the various embodiments described above may be composed of a single or multiple entities, and some of the corresponding sub-components described above may be omitted, or other sub-components may be further included in various embodiments. Alternatively or additionally, some components (e.g., modules or programs) may be integrated into a single entity, which may perform the same or similar functions as those performed by each of the corresponding components prior to integration. Operations performed by modules, programs or other components according to various embodiments may be executed sequentially, in parallel, iteratively or heuristically, or at least some operations may be executed in a different order, omitted, or other operations may be added.
[0159] Although the preferred embodiments of the present disclosure have been illustrated and described above, the present disclosure is not limited to the specific embodiments described above, and various modifications may be made by a person having ordinary skill in the art to which the present disclosure pertains without departing from the gist of the present disclosure as claimed in the claims, and such modifications should not be understood individually from the technical idea or prospect of the present disclosure.
Claims
1. In electronic devices, At least one memory storing a first neural network model trained to output a saliency map for an image and a second neural network model trained to output a quality score of the image; and At least one processor connected to said at least one memory, Based on the first image, an importance map including the importance values of each of a plurality of pixels included in the first image is obtained through the first neural network model, Identifying a plurality of first sub-regions each corresponding to a plurality of regions included in the first image based on the importance map, An electronic device that obtains a quality score for the first image through the second neural network model based on the plurality of first sub-regions identified above, wherein the quality score is based on a plurality of first quality scores each corresponding to the plurality of first sub-regions identified above.
2. In paragraph 1, At least one processor, An electronic device that identifies a portion of each of the plurality of regions as a first sub-region of the plurality of first sub-regions corresponding to the plurality of regions based on the importance values of pixels included in each of the plurality of regions.
3. In paragraph 2, At least one processor, For each of the above multiple areas, Determine a plurality of first sum values corresponding to each pixel included in each area, and each first sum value of the plurality of first sum values is the sum of the importance value of each pixel and the importance values of the surrounding pixels of each pixel, An electronic device that identifies a sub-region of each region including a reference pixel corresponding to the largest first sum value among the plurality of first sum values as a first sub-region corresponding to each region.
4. In paragraph 1, The second neural network model is further trained to output the plurality of first quality scores based on the plurality of first sub-regions and the plurality of second sum values input to the second neural network model, At least one processor, Determine a plurality of second sum values corresponding to each of the plurality of first sub-regions, and each second sum value of the plurality of second sum values is the sum of importance values of pixels included in the corresponding first sub-region of the plurality of first sub-regions, An electronic device that obtains the quality score for the first image through the second neural network model based on the plurality of identified first sub-regions and the plurality of second summed values.
5. In paragraph 1, At least one processor, Through the first neural network model, multiple importance maps corresponding to each of multiple frames are obtained, An electronic device that identifies one of the plurality of frames as the first image based on the plurality of importance maps output by the first neural network model.
6. In paragraph 1, At least one processor, Determine a plurality of third sum values corresponding to each of the plurality of areas, and each third sum value of the plurality of third sum values is the sum of the importance values of pixels included in the corresponding area of the plurality of areas, An electronic device that identifies an area of the plurality of areas corresponding to a third sum value greater than or equal to a preset size among the plurality of third sum values as an additional first sub-area.
7. In paragraph 1, At least one processor, Determine a plurality of third sum values corresponding to each of the plurality of areas, and each third sum value of the plurality of third sum values is the sum of the importance values of pixels included in the corresponding area of the plurality of areas, An electronic device that updates the size of an area of the plurality of areas based on the plurality of third sum values.
8. In paragraph 1, The first image above is, It is the first frame among multiple frames, The second image above is, The second frame of the plurality of frames immediately following the first frame of the plurality of frames, At least one processor, Determine a motion vector based on the first image and the second image, Identifying a plurality of first sub-regions and a plurality of second sub-regions of the second image corresponding to the motion vector, An electronic device that obtains a quality score of the second image through the second neural network model based on the identified plurality of second sub-regions.
9. In paragraph 1, At least one processor, An electronic device that performs at least one of upscaling and noise removal on the first image based on a quality score for the first image.
10. In paragraph 1, The above first neural network model is, Learning a plurality of first sample images and a plurality of sample importance maps each corresponding to the plurality of first sample images, The above second neural network model is, An electronic device that learns a plurality of second sample images and a plurality of sample scores corresponding to each of the plurality of second sample images.
11. In paragraph 1, Each of the above multiple sample importance maps is: Based on the gazes of multiple users for the first sample image corresponding to each of the plurality of sample importance maps, Each of the above multiple sample scores is: An electronic device based on scores of the plurality of users for second sample images corresponding to each of the plurality of sample scores.
12. A control method of an electronic device including at least one processor, wherein the first neural network model trained to output a saliency map for an image and the second neural network model trained to output a quality score of the image are stored, By at least one processor, Based on the first image, an importance map including the importance values of each of a plurality of pixels included in the first image is obtained through the first neural network model; Identifying a plurality of first sub-regions each corresponding to a plurality of regions included in the first image based on the importance map; and A control method for obtaining a quality score for the first image through the second neural network model based on the plurality of first sub-regions identified above, wherein the quality score is based on a plurality of first quality scores each corresponding to the plurality of first sub-regions identified above.
13. In paragraph 12, The step of identifying the plurality of first sub-regions is: By at least one processor, A control method for identifying a portion of each of the plurality of regions as a first sub-region of the plurality of first sub-regions corresponding to the plurality of regions based on the importance values of pixels included in each of the plurality of regions.
14. In paragraph 13, The step of identifying the first sub-area is: By at least one processor, For each of the above multiple areas, Determine a plurality of first sum values corresponding to each pixel included in each area, and each first sum value of the plurality of first sum values is the sum of the importance value of each pixel and the importance values of the surrounding pixels of each pixel, A control method for identifying a sub-region of each region including a reference pixel corresponding to the largest first sum value among the plurality of first sum values as a first sub-region corresponding to each region.
15. In paragraph 12, The second neural network model is further trained to output the plurality of first quality scores based on the plurality of first sub-regions and the plurality of second sum values input to the second neural network model, By at least one processor, Determine a plurality of second sum values corresponding to each of the plurality of first sub-regions, and each second sum value of the plurality of second sum values is the sum of importance values of pixels included in the corresponding first sub-region of the plurality of first sub-regions, The step of obtaining a quality score for the above first image is: A control method for obtaining the quality score for the first image through the second neural network model based on the identified plurality of first sub-regions and the plurality of second summed values.
Citation Information
Patent Citations
Image evaluation apparatus, image evaluation method, and program
JP2016149010A
Device for encoding and decoding for 2d image for 3D image generation
KR1020210126940A
A system for managing and expressing sensing information of identification tags and electronic transportation related information and method thereof
KR1020220068182A
Electronic apparatus providing summarized content and operating method thereof
KR1020250112099A
Image processing device, image display system, image processing method, and recording medium
WO2022180684A1