Content-specific fidelity metrics for image compression based on semantic segmentation model

By using a machine learning model to extract specific types of content from digital images, generate segmentation masks and calculate visual fidelity metrics, the problem of the inability to optimize lossy compression in existing technologies is solved, and visual fidelity assessment of specific types of content and improvement of image compression quality are achieved.

CN120823271APending Publication Date: 2025-10-21SYNAPTICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510446894.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-04-12
Filing Date
2025-04-10
Publication Date
2025-10-21

AI Technical Summary

Technical Problem

Existing visual fidelity metrics cannot effectively evaluate the visual fidelity of specific types of content in digital images, resulting in the inability to optimize lossy compression techniques for different types of image content.

Method used

A machine learning model is used to extract specific types of content from digital images, generate segmentation masks, and calculate visual fidelity metrics based on the segmentation masks to selectively transmit encoded images to optimize the compression scheme.

Benefits of technology

It realizes the visual fidelity evaluation of specific types of content, dynamically selects the optimal image compression scheme, and improves the visual fidelity and quality of image compression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120823271A_ABST
    Figure CN120823271A_ABST
Patent Text Reader

Abstract

The invention provides a method, a device and a system for image compression. The present implementations more particularly relate to systems and techniques for selecting an image compression scheme for a given type of content or application. An image encoder may encode an image based on an image compression scheme. In some aspects, an image encoder may infer first and second segmentation masks from an original image and an encoded image, respectively, based on a machine learning model. The machine learning model may be trained to extract one or more types of content from the input image such that the segmentation mask includes only the extracted content from the image (and excludes any other type of content). The image encoder may further calculate a visual fidelity metric for the encoded image based on the mask, and selectively transmit the encoded image over the communication channel based at least in part on the visual fidelity metric.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present implementations generally relate to image compression, and in particular to a content-specific fidelity metric for image compression based on a semantic segmentation model. Background Art

[0002] A digital image can be represented by an array of pixel values ​​(or multiple arrays of pixel values ​​associated with different channels) that can be displayed or otherwise rendered on an electronic display device (such as a computer, smartphone, or television, among other examples). Digital video is a sequence of digital images (or "frames") that can be displayed or otherwise rendered in succession. Some electronic display devices can receive (one or more) digital images from a source device (such as an image capture device or data repository) via a communication channel (such as a wired or wireless medium). Due to bandwidth limitations of the communication channel, digital image data is typically encoded or compressed before transmission from the source device. Data compression is a technique for encoding information into smaller units of data. The encoded image data is then decoded by a display device to recover the corresponding digital image. As such, data compression can reduce the bandwidth or overhead required to store or transmit digital images over a communication channel.

[0003] Data compression techniques can generally be categorized as either "lossy" or "lossless." Lossless data compression does not result in any loss of information between the encoding step and the decoding step, as long as the communication channel does not introduce errors into the encoded data. As a result, the decoded image is the same (or substantially the same) as the original image before encoding. Example lossless compression techniques include entropy coding (such as arithmetic coding, Huffman coding, or Golomb coding) and run-length encoding (RLE), among other examples. In contrast, lossy data compression may result in some loss of information between the encoding step and the decoding step. As a result, the decoded image may have lower image quality than the original image before encoding. Example lossy compression techniques include transform coding (such as through the application of spatial frequency transforms) and quantization (such as through the application of quantization matrices), among other examples.

[0004] Different lossy compression techniques may be more suitable for encoding different types of image content. For example, some compression techniques may maintain more detail or visual fidelity in text or geometric shapes (also referred to as "screen content") than in other content in a digital image. To determine the suitability of any lossy compression technique for a given application, various visual fidelity metrics can be used to compare the image quality of the compressed image and the original image. Examples of suitable visual fidelity metrics include peak signal-to-noise ratio (PSNR), PSNR based on properties of the human visual system (PSNR-HVS), PSNR-HVS with visual masking (PSNR-HVS-M), video multi-method assessment fusion (VMAF), and learned perceptual image patch similarity (LPIPS), among others. However, when applied to digital images, such visual fidelity metrics only indicate the overall visual fidelity of the image (as a whole). Therefore, new image analysis techniques are needed to assess the visual fidelity of specific types of content in digital images (excluding other types of content). Summary of the Invention

[0005] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter.

[0006] One innovative aspect of the subject matter of the present disclosure can be implemented in a method for image compression. The method includes the steps of receiving an image for transmission over a communication channel; encoding the image into a first encoded image based on a first image compression scheme; inferring a first segmentation mask from the image based on a first machine learning model; inferring a second segmentation mask from the first encoded image based on the first machine learning model; computing a first visual fidelity metric for the first encoded image based on the first segmentation mask and the second segmentation mask; and selectively transmitting the first encoded image over the communication channel based at least in part on the first visual fidelity metric.

[0007] Another innovative aspect of the subject matter of the present disclosure can be implemented in an image encoder comprising a processing system and a memory storing instructions that, when executed by the processing system, cause the image encoder to receive an image for transmission over a communication channel; encode the image into a first encoded image based on a first image compression scheme; infer a first segmentation mask from the image based on a first machine learning model; infer a second segmentation mask from the first encoded image based on the first machine learning model; compute a first visual fidelity metric for the first encoded image based on the first segmentation mask and the second segmentation mask; and selectively transmit the first encoded image over the communication channel based at least in part on the first visual fidelity metric.

[0008] Another innovative aspect of the disclosed subject matter can be implemented in a method for image compression, comprising the steps of generating an input image including content overlaid with other media, generating a segmentation mask based on the content included in the input image, and training a neural network to regenerate the segmentation mask based on the input image. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The present implementations are illustrated by way of example and are not intended to be limited by the figures of the accompanying drawings.

[0010] Figure 1 An example communication system for encoding and decoding data is shown.

[0011] Figure 2 A block diagram illustrating an example image encoding system according to some implementations is shown.

[0012] Figure 3 A block diagram illustrating an example content extractor for digital images, according to some implementations.

[0013] Figure 4 A block diagram illustrating an example machine learning system according to some implementations is shown.

[0014] Figure 5 A block diagram illustrating an example image encoder according to some implementations is shown.

[0015] Figure 6 An illustrative flow diagram depicting example operations for image compression is shown, according to some implementations.

[0016] Figure 7 An illustrative flow chart depicting example operations for training a neural network is shown, according to some implementations. DETAILED DESCRIPTION

[0017] In the following description, many specific details (such as examples of specific components, circuits and processes) are set forth to provide a thorough understanding of the present disclosure. As used herein, the term "coupling" means being directly connected to one or more intermediate components or circuits or being connected through one or more intermediate components or circuits. The terms "electronic system" and "electronic device" can be used interchangeably to refer to any system capable of electronically processing information. Moreover, in the following description, and for the purpose of explanation, specific terms are set forth to provide a thorough understanding of aspects of the present disclosure. However, it will be apparent to those skilled in the art that these specific details may not be required to practice the example embodiments. In other examples, well-known circuits and devices are shown in block diagram form to avoid making the present disclosure difficult to understand. Some parts of the subsequent detailed description are presented based on other symbolic representations of procedures, logic blocks, processing and the operation of data bits in computer memory.

[0018] These descriptions and representations are the means by which those skilled in the art of data processing are used to most effectively convey the essence of their work to other skilled in the art. In the present disclosure, procedures, logic blocks, processes, etc. are considered to be the self-consistent sequences of steps or instructions that lead to a desired result. The steps are those steps that require the physical manipulation of physical quantities. Usually, although not necessarily, these quantities take the form of electrical or magnetic signals that can be stored, transferred, combined, compared, and otherwise manipulated in a computer system. However, it should be remembered that all terms in these and similar terms are associated with appropriate physical quantities and are merely convenient labels that are applied to these quantities.

[0019] Unless otherwise specifically stated, as will be apparent from the following discussion, it is appreciated that throughout this application, discussions utilizing terms such as "access," "receive," "send," "use," "select," "determine," "normalize," "multiply," "average," "monitor," "compare," "apply," "update," "measure," "infer," etc., refer to the actions and processes of a computer system or similar electronic computing device that manipulates and transforms data represented as physical (electronic) quantities within a computer system's registers and memories into other data similarly represented as physical quantities within the computer system's memories or registers or other such information storage, transmission, or display devices.

[0020] In the figures, a single block may be described as performing one or more functions; however, in actual practice, the one or more functions performed by the block may be performed in a single component or across multiple components, and / or may be performed using hardware, using software, or using a combination of hardware and software. In order to clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described below generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. Technicians may implement the described functionality in different ways for each specific application, but such implementation decisions should not be interpreted as causing a departure from the scope of this disclosure. Moreover, the example input device may include components other than those shown, including well-known components such as processors, memories, etc.

[0021] The techniques described herein, unless specifically described as being implemented in a particular manner, may be implemented in hardware, software, firmware, or any combination thereof. Any features described as modules or components may also be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the techniques may be implemented at least in part by a non-transitory processor-readable storage medium comprising instructions that, when executed, perform one or more of the methods described above. The non-transitory processor-readable data storage medium may form part of a computer program product, which may include packaging materials.

[0022] Non-transitory processor-readable storage media may include random access memory (RAM) such as synchronous dynamic random access memory (SDRAM), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), flash memory, other known storage media, etc. Additionally or alternatively, the technology may be implemented at least in part by a processor-readable communication medium that carries or communicates code in the form of instructions or data structures and can be accessed, read, and / or executed by a computer or other processor.

[0023] The various illustrative logical blocks, modules, circuits, and instructions described in connection with the embodiments disclosed herein may be executed by one or more processors (or processing systems). As used herein, the term "processor" may refer to any general-purpose processor, special-purpose processor, conventional processor, controller, microcontroller, and / or state machine capable of executing scripts or instructions of one or more software programs stored in memory.

[0024] As described above, different lossy compression techniques may be more suitable for encoding different types of image content. For example, some compression techniques may maintain more detail or visual fidelity in text or geometric shapes (also referred to as "screen content") than in other content in a digital image. In order to determine the suitability of any lossy compression technique for a given application, various visual fidelity metrics can be used to compare the image quality of the compressed image and the original image. Example suitable visual fidelity metrics include peak signal-to-noise ratio (PSNR), PSNR based on properties of the human visual system (PSNR-HVS), PSNR-HVS with visual masking (PSNR-HVS-M), video multi-method assessment fusion (VMAF), and learned perceptual image patch similarity (LPIPS), among other examples. However, when applied to digital images, such visual fidelity metrics only indicate the overall visual fidelity of the image (as a whole). Aspects of the present disclosure recognize that machine learning models can be trained to extract specific types of content from digital images (excluding other types of content).

[0025] Machine learning is a technique used to improve the ability of a computer system or application to perform a specific task. During the training phase, a machine learning system is provided with multiple "answers" and a large amount of raw input data. The machine learning system analyzes the input data to learn a set of rules (also referred to as a "machine learning model") that can be used to map the input data to answers. During the inference phase, the machine learning system uses the trained machine learning model to infer the answers from new input data. By training the machine learning model to infer (or "extract") only one or more types of content from a digital image, aspects of the present disclosure can use existing visual fidelity metrics to compare the content extracted from the compressed image with the content extracted from the original image. Therefore, the resulting visual fidelity metric can indicate the extent to which the image compression scheme maintains the visual fidelity of a particular type of image content (such as screen content).

[0026] Various aspects generally relate to image compression, and more particularly, to systems and techniques for selecting an image compression scheme for a given type of content or application. An image encoder may receive an image for transmission over a communication channel and encode the image based on an image compression scheme. In some aspects, the image encoder may infer first and second segmentation masks from the original image and the encoded image, respectively, based on a machine learning model. In some implementations, the machine learning model may be trained to extract one or more types of content from the input image such that the segmentation mask only includes the extracted content from the image (and excludes any other type of content). The image encoder may further calculate a visual fidelity metric for the encoded image based on the first and second segmentation masks, and selectively transmit the encoded image over the communication channel based at least in part on the visual fidelity metric. In some implementations, the image encoder may repeat the process using different image compression techniques and transmit the encoded image with the highest visual fidelity metric.

[0027] Particular implementations of the subject matter described in this disclosure can be implemented to achieve one or more of the following potential advantages. By training a machine learning model to generate a segmentation mask that includes only a specific type of content from an input image, aspects of the present disclosure can assess how well an image compression scheme maintains the visual fidelity of such content in a digital image. For example, by computing a visual fidelity metric based on the segmentation mask (rather than the digital image), the resulting visual fidelity metric can indicate the image quality of the desired type of content in the compressed image (to the exclusion of any other type of content) (rather than the overall image quality of the compressed image as a whole). Thus, an image encoder can dynamically select an optimal image compression scheme (among any available image compression schemes) for any given application or image content type.

[0028] Figure 1An example communication system 100 for encoding and decoding data is shown. Communication system 100 includes an encoder 110 and a decoder 120. In some implementations, encoder 110 and decoder 120 may be provided in corresponding communication devices (such as, for example, computers, switches, routers, hubs, gateways, cameras, displays, or other devices capable of transmitting or receiving communication signals). In some other implementations, encoder 110 and decoder 120 may be included in the same device or system.

[0029] The encoder 110 receives input data 102 to be transmitted or stored via a channel 130. For example, the channel 130 may comprise a wired or wireless communication medium that facilitates communication between the encoder 110 and the decoder 120. Alternatively, or in addition, the channel 130 may comprise a data storage medium. In some aspects, the encoder 110 may be configured to compress the size of the input data 102 to accommodate bandwidth, storage, or other resource limitations associated with the channel 130. For example, the encoder 110 may encode each unit of the input data 102 into a corresponding "codeword" (as encoded data 104) that can be transmitted or stored via the channel 130. The decoder 120 is configured to receive the encoded data 104 via the channel 130 and decode the encoded data 104 into output data 106. For example, the decoder 120 may decompress or otherwise reverse the compression performed by the encoder 110 so that the output data 106 is substantially similar to (if not identical to) the original input data 102.

[0030] Data compression techniques can generally be categorized as "lossy" or "lossless." Lossless data compression does not result in any loss of information between the encoding and decoding steps, as long as the channel 130 does not introduce errors into the encoded data 104. As a result, the output data 106 is the same (or substantially the same) as the input data 102. Example lossless compression techniques include entropy coding (such as arithmetic coding, Huffman coding, or Grunb coding) and run-length encoding (RLE), among other examples. In contrast, lossy data compression may result in some loss of information between the encoding and decoding steps. As a result, the output data 106 may be different from the input data 102. Example lossy compression techniques include transform coding (such as through the application of a spatial frequency transform) and quantization (such as through the application of a quantization matrix), among other examples.

[0031] Different lossy compression techniques may be more suitable for encoding different types of input data 102. For example, digital images are often encoded using lossy compression techniques that maintain the visual fidelity of certain aspects of the image (such as text or other "screen content" representing important information) while sacrificing the visual fidelity of other aspects of the image (such as background content or visual material intended to fill empty space). Therefore, the optimal encoding or compression scheme for any given application may depend on the type of content to be prioritized for that application. In some aspects, the encoder 110 may select a lossy compression scheme to be used to encode the input data 102 based at least in part on the type of content to be prioritized in the input data 102. For example, the encoder 110 may compare the performance of various lossy compression schemes with respect to maintaining a particular type of content in the input data 102 and select the compression scheme that produces the greatest performance.

[0032] Figure 2 A block diagram of an example image encoding system 200 is shown according to some implementations. The image encoding system 200 is configured to encode image data 201 into encoded image data 209. The image data 201 may include an array of pixel values ​​(or multiple arrays of pixel values ​​associated with different color channels) representing a frame of a digital image or video captured or acquired by an image source (such as a camera or other image output device). In some implementations, the image encoding system 200 may be Figure 1 An example of the encoder 110. Figure 1 , image data 201 may be an example of input data 102 , and encoded image data 209 may be an example of encoded data 104 .

[0033] The image encoding system 200 includes a plurality (N) of image compression components 210(1)-210(N), a content extraction component 220, an image quality estimation component 230, and an image quality comparison component 240. The image compression components 210(1)-210(N) are configured to encode image data 201 into encoded image data 202(1)-202(N), respectively, according to one or more image compression schemes. In some implementations, each of the image compression components 210(1)-210(N) may implement a corresponding lossy compression scheme. As shown in FIG. Figure 1 As described, different lossy compression techniques may be better at maintaining visual fidelity for different types of content associated with input image data 201. For example, screen content (such as text, geometric shapes, or icons) may have a different level of detail or image quality in encoded image data 202(1) than in encoded image data 202(N).

[0034] The content extraction component 220 is configured to extract one or more types of content from the image data 201 and the encoded image data 202(1)-202(N). In some implementations, the content extraction component 220 may generate a segmentation mask 204(0) (also referred to as a "reference mask") that includes only the specific type(s) of content extracted from the image data 201 (excluding all other types of content from the image data 201), and may generate segmentation masks 204(1)-204(N) that include only the specific type(s) of content extracted from the encoded image data 202(1)-202(N), respectively. In some implementations, each of the segmentation masks 204(0)-204(N) may include only screen content from the image data 201 and the encoded image data 202(1)-202(N). Because different lossy compression techniques are used to generate the encoded image data 202(1)-202(N), each of the segmentation masks 204(1)-204(N) may have a different level of visual fidelity or image quality.

[0035] In some implementations, each of the segmentation masks 204(0)-204(N) may indicate the opacity of content extracted from the corresponding image data (such as image data 201 and encoded image data 202(1)-202(N)). This may ensure that the edges of the content can be more accurately reproduced to provide a more satisfactory viewing experience (particularly for text-based content). For example, the reference mask 204(0) may be an 8-bit (floating point or integer) mask that indicates the extent or amount of a particular content type included in each pixel of the image data 201 (such as on a scale of 256 values). Similarly, each of the segmentation masks 204(1)-204(N) may also be an 8-bit mask that indicates the extent or amount of a particular content type included in each pixel of the encoded image data 202(1)-202(N), respectively.

[0036] The image quality estimation component 230 is configured to compare each of the segmentation masks 204(1)-204(N) with the reference mask 204(0) and calculate visual fidelity metrics 206(1)-206(N) that indicate the visual fidelity of the segmentation masks 204(1)-204(N), respectively. In some implementations, the image quality estimation component 230 may calculate the visual fidelity metrics 206(1)-206(N) using any known visual fidelity or image quality estimation technique. Example suitable visual fidelity metrics include PSNR, PSNR-HVS, PSNR-HVS-M, VMAF, and LPIPS, among others. In some other implementations, the image quality estimation component 230 may use the segmentation masks 204(0)-204(N) as weights to be applied to other types of metrics. For example, the image quality estimation component 230 may calculate weighted differences or convolutions of intermediate weights between the segmentation masks 204(0)-204(N). Thus, visual fidelity metrics 206 ( 1 )- 206 (N) may indicate the extent to which each of image compression components 210 ( 1 )- 210 (N) maintains the visual fidelity of particular type(s) of content in image data 201 .

[0037] The image quality comparison component 240 is configured to compare the visual fidelity metrics 206(1)-206(N) and select one of the image compression components 210(1)-210(N) to be used for a given application based at least in part on the comparison. For example, the image quality comparison component 240 may generate an encoding selection signal 208 indicating the selected image compression component (or scheme). In some implementations, the image quality comparison component 240 may select the image compression component (or scheme) associated with the highest visual fidelity metric (or the visual fidelity metric indicating the highest image quality) among the visual fidelity metrics 206(1)-206(N). For example, if the visual fidelity metric 206(1) is associated with the highest image quality among the visual fidelity metrics 206(1)-206(N), the encoding selection signal 208 may indicate the image compression component 210(1).

[0038] In some implementations, the encoding selection signal 208 may be provided as a selection input to a multiplexer 250, which is configured to output one of the set of encoded image data 202(1)-202(N) as the encoded image data 209. For example, if the encoding selection signal 208 indicates the image compression component 210(1), the multiplexer 250 may output the encoded image data 202(1) as the encoded image data 209. Thus, the image encoding system 200 may output the encoded image data 209 optimized for any given application or content type.

[0039] In some aspects, the image encoding system 200 may transmit the encoded image data 209 to an image decoder (such as decoder 120) via a communication channel (such as channel 130). The image decoder may decode the encoded image data 209 to reproduce a digital image on a display device (such as a television, a computer monitor, a smartphone, or any other device including an electronic display). Figure 1 As described above, the image decoder can reverse the encoding performed by the image encoding system 200 to recover the digital image represented by the original image data 201. In some implementations, the image encoding system 200 can transmit a series of frames of encoded image data 209 (each frame representing a corresponding image or frame of digital video) so that the image decoder can display or render the digital video on a display device.

[0040] Figure 3 A block diagram of an example content extractor 300 for digital images is shown, according to some implementations. The content extractor 300 is configured to receive a reference image 302 and an encoded image 304, and to generate a reference mask 306 and an encoded mask 308 based on the images 302 and 304, respectively.

[0041] In some implementations, the content extractor 300 may be Figure 2 An example of a content extraction component 220. Figure 2 , reference image 302 may be an example of image data 201, and reference mask 306 may be an example of segmentation mask 204(0), while coded image 304 may be an example of any of coded image data 202(1)-202(N), and coded mask 308 may be an example of any of coded image data 204(1)-204(N). Although only one coded image is shown (for simplicity), in actual implementations content extractor 300 may receive as input any number (N) of coded images (such as reference image 304(1)-204(N)). Figure 2 and described).

[0042] The content extractor 300 includes a first mask generation component 310 and a second mask generation component 320. The first mask generation component 310 is configured to extract one or more types of content from the reference image 302 to generate a reference mask 306. The second mask generation component 320 is configured to extract one or more types of content from the encoding image 304 to generate an encoding mask 308. Figure 3 In the example of FIG, each of the mask generation components 310 and 320 is configured to extract screen content from the reference image 302 and the encoded image 304, respectively. In some implementations, each of the masks 306 and 308 can be an 8-bit (floating point or integer) mask that indicates the opacity of the content extracted from the images 302 and 304, respectively (such as the reference image 302 and the encoded image 304). Figure 2 and described).

[0043] like Figure 3 As shown in , reference mask 306 includes only text, geometric shapes, and icons that are overlaid on other media in reference image 302 (such as an image of a building). More specifically, each pixel of reference mask 306 maps to a corresponding pixel of reference image 302 (but every pixel of reference image 302 does not map to a corresponding pixel of reference mask 306). Similarly, encoding mask 308 includes only text, geometric shapes, and icons that are overlaid on other media in encoding image 304 (such as an image of a building). More specifically, each pixel of encoding mask 308 maps to a corresponding pixel of encoding image 304 (but every pixel of encoding image 304 does not map to a corresponding pixel of encoding mask 308).

[0044] Aspects of the present disclosure recognize that a machine learning model can be trained to extract specific types of content from digital images (to the exclusion of other types of content). Machine learning is a technique used to improve the ability of a computer system or application to perform specific tasks. During the training phase, a machine learning system is provided with multiple "answers" and a large amount of raw input data. The machine learning system analyzes the input data to learn a set of rules (also referred to as a "machine learning model") that can be used to map the input data to answers. During the inference phase, the machine learning system uses the trained machine learning model to infer answers from new input data.

[0045] In some aspects, mask generation components 310 and 320 can extract content for segmentation masks 306 and 308 from images 302 and 304, respectively, based on a machine learning (ML) model 301. By using the same ML model 301 to infer (or "extract") one or more types of content (such as screen content) from each of images 302 and 304, aspects of the present disclosure can use existing visual fidelity metrics to compare the content extracted from the encoded image 304 with the content extracted from the reference image 302 (such as a reference image). Figure 2 Thus, the resulting visual fidelity metric may indicate how well the image compression scheme maintains the visual fidelity of a particular type of image content, such as screen content.

[0046] In some aspects, the ML model 301 may extract multiple types of content from the images 302 and 304. In some implementations, a single ML model 301 may be trained to infer multiple segmentation masks from a single input image. For example, each of the segmentation masks may include different types of content (such as text, geometric shapes, or icons) extracted from the same input image. In some other implementations, multiple ML models may be used to extract different types of content from the images 302 and 304. For example, a first ML model may be trained to infer a segmentation mask that includes only text, a second ML model may be trained to infer a segmentation mask that includes only geometric shapes, and a third ML model may be trained to infer a segmentation mask that includes only icons. Generating multiple segmentation masks for each of the images 302 and 304 allows for more fine-grained visual fidelity estimation.

[0047] As reference Figure 2 As described above, an image quality estimation component (such as image quality estimation component 230) may compare the encoding mask 308 to the reference mask 306 and calculate a visual fidelity metric (such as any of visual fidelity metrics 206(1)-206(N)) that indicates the visual fidelity of the screen content in the encoded image 304. Figure 3 As shown in , the screen content in the encoded mask 308 appears grainy, blurry, broken, and faded compared to the screen content in the reference mask 306. Therefore, the lossy compression scheme used to generate the encoded image 304 may not be well suited for the current application or content type.

[0048] Figure 4 A block diagram of an example machine learning system 400 is shown according to some implementations. The machine learning system 400 is configured to generate a neural network model 407 based at least in part on a plurality of input images 401 and screen content 402 to be extracted from the input images 401. In some implementations, the neural network model 407 may be Figure 3 Thus, neural network model 407 may include a set of rules that can be used to infer or extract screen content from an input image (such as any of images 302 or 304).

[0049] Machine learning system 400 includes an image compositor 410, a neural network 420, and a loss calculator 430. Image compositor 410 is configured to combine screen content 402 with input image 401 to generate a composite image 403 (similar to reference image 302). Screen content 402 may include pre-generated text, geometric shapes, icons, or other content for which visual fidelity is to be measured. In some implementations, image compositor 410 may overlay screen content 402 on input image 401 using any known image synthesis technique. Image compositor 410 also generates a ground truth mask 404 based on screen content 402 and input image 401. Ground truth mask 404 includes only screen content 402 from composite image 403 (similar to reference mask 306). In some implementations, image compositor 410 may infer ground truth mask 404 from the alpha channel of composite image 403. For example, the ground truth mask 404 may be a single-channel floating point mask having the same size or dimensions as the synthesized image 403 .

[0050] In some implementations, the machine learning system 300 may train a neural network 420 to reproduce a ground truth mask 404 based on the synthesized image 403. Deep learning is a specific form of machine learning in which the inference and training phases are performed over multiple layers. Deep learning architectures are often referred to as "artificial neural networks" due to the way in which information is processed (similar to biological neural systems). For example, each layer of an artificial neural network may be composed of one or more "neurons." Each layer of neurons may perform different transformations on the output data from the previous layer so that the final output of the neural network produces the desired inference. The collection of transformations associated with the various layers of the network is referred to as a "neural network model." Example suitable neural networks include convolutional neural networks (CNNs) and recurrent neural networks (RNNs), among other examples.

[0051] Neural network 420 receives synthetic image 403 and attempts to reconstruct ground truth mask 404. For example, neural network 420 may form a network of connections across multiple layers of artificial neurons that start from synthetic image 403 and lead to output mask 405. The connections are weighted to produce an output mask 405 that approximates ground truth mask 404. The training operation may be performed over multiple iterations. In each iteration, neural network 420 produces output mask 405 based on the weighted connections across the layers of artificial neurons, and loss calculator 430 updates weights 406 associated with the connections based on the amount of loss (or error) between output mask 405 and ground truth mask 404. When certain convergence criteria are met (such as when the loss is below a threshold level or after a predetermined number of training iterations), neural network 420 may output the weighted connections as neural network model 407.

[0052] In some implementations, the neural network model 407 may be trained to generate multiple output masks 405 for multiple types of screen content 402 (or multiple color channels of each type of screen content 402). In other words, the neural network model 407 may be trained to segment different types of content simultaneously. For example, the neural network model 407 may generate a different output mask 405 for each of text, geometric shapes, and icons. In such implementations, the image synthesizer 410 may generate a corresponding ground truth mask 404 for each type of screen content 402 to be represented in a different output mask 405. As shown in FIG. Figure 3 As described, generating multiple segmentation masks for an input image allows for finer-grained estimation of visual fidelity.

[0053] In some other implementations, the machine learning system 400 may be configured to train multiple neural network models 407 for multiple types of screen content 402 (or multiple color channels of each type of screen content 402). For example, the machine learning system 400 may repeat the training operations described above for different types of screen content 402, such that a different neural network model 407 is generated for each type of screen content 402. Training the neural network model 407 to distinguish between various types of screen content 402 improves the accuracy of the segmentation masks inferred by the neural network model 407. For example, training the neural network model 407 to extract text from the composite image 403 while excluding geometric shapes and icons improves the accuracy of the neural network model 407 in extracting text.

[0054] Figure 5 A block diagram of an example image encoder 500 is shown according to some implementations. In some implementations, the image encoder 500 may be Figure 2 The image encoder 500 is an example of the image encoding system 200. More specifically, the image encoder 500 may be configured to encode image data for transmission over a communication channel.

[0055] The image encoder 500 includes a communication interface 510, a processing system 520, and a memory 530. The communication interface 510 is configured to receive image data from an image source and transmit the encoded image data via a communication channel. In some aspects, the communication interface 510 may include an image source interface (I / F) 512 for communicating with the image source and a channel interface 514 for communicating via the communication channel. In some implementations, the image source interface 512 may receive an image for transmission (such as to be transmitted) via the communication channel.

[0056] The memory 530 may include a non-transitory computer-readable medium (including one or more non-volatile memory elements such as EPROM, EEPROM, flash memory, hard drive, etc.) that may store at least the following software (SW) modules: Image encoding SW module 532, configured to encode the image into a first encoding based on a first image compression scheme image; A mask generation SW module 534 for inferring a first segmentation mask from an image based on a first machine learning model. code, and inferring a second segmentation mask from the first encoded image based on a first machine learning model; Image quality determination SW module 536 for calculating the first segmentation mask and the second segmentation mask a first visual fidelity metric of the encoded image; and An image quality comparison SW module 538 for comparing the image quality of the image based at least in part on the first visual fidelity metric. The communication channel selectively transmits the first coded image. Each software module includes instructions that, when executed by the processing system 520 , cause the image encoder 500 to perform a corresponding function.

[0057] The processing system 520 may include any suitable one or more processors capable of executing scripts or instructions of one or more software programs stored in the image encoder 500 (such as in the memory 530). For example, the processing system 520 may execute the image encoding SW module 532 to encode the image into a first encoded image based on a first image compression scheme. The processing system 520 may execute the mask generation SW module 534 to infer a first segmentation mask from the image based on a first machine learning model, and infer a second segmentation mask from the first encoded image based on the first machine learning model. The processing system 520 may execute the image quality determination SW module 536 to calculate a first visual fidelity metric for the first encoded image based on the first segmentation mask and the second segmentation mask. The processing system 520 may further execute the image quality comparison SW module 538 to selectively transmit the first encoded image over a communication channel based at least in part on the first visual fidelity metric.

[0058] Figure 6 An illustrative flow chart depicting example operations 600 for image compression according to some implementations is shown. In some implementations, the example operations 600 may be performed by an image encoder such as Figure 2 Image coding system 200 or Figure 5 The image encoder 500) is executed.

[0059] An image encoder receives an image for transmission over a communication channel (610). The image encoder encodes the image into a first encoded image based on a first image compression scheme (620). The image encoder infers a first segmentation mask from the image based on a first machine learning model (630). The image encoder also infers a second segmentation mask from the first encoded image based on the first machine learning model (640). The image encoder calculates a first visual fidelity metric for the first encoded image based on the first segmentation mask and the second segmentation mask (650). In some implementations, the first visual fidelity metric may include PSNR, PSNR-HVS, PSNR-HVS-M, a VMAF metric, or a LPIPS metric. The image encoder selectively transmits the first encoded image over the communication channel based at least in part on the first visual fidelity metric (660).

[0060] In some aspects, the image encoder may further encode the image into a second encoded image based on a second image compression scheme different from the first image compression scheme; infer a third segmentation mask from the second encoded image based on the first machine learning model; calculate a second visual fidelity metric for the second encoded image based on the first segmentation mask and the third segmentation mask; and determine whether the first visual fidelity metric or the second visual fidelity metric indicates a higher image quality, wherein the first encoded image is selectively transmitted over the communication channel based at least in part on whether the first visual fidelity metric or the second visual fidelity metric indicates a higher image quality.

[0061] In some implementations, the selective transmission of the first encoded image may include refraining from transmitting the first encoded image over the communication channel in response to determining that the second visual fidelity metric indicates a higher image quality. In some implementations, the image encoder may further transmit the second encoded image over the communication channel instead of the first encoded image in response to determining that the second visual fidelity metric indicates a higher image quality.

[0062] In some aspects, the first machine learning model can be trained to extract a first type of content from one or more input images. In some implementations, the first type of content can include screen content. In some implementations, the first type of content can include text, geometric shapes, or icons.

[0063] In some aspects, the image encoder may further infer a third segmentation mask from the image based on a second machine learning model different from the first machine learning model; infer a fourth segmentation mask from the first encoded image based on the second machine learning model; and calculate a second visual fidelity metric for the first encoded image based on the third segmentation mask and the fourth segmentation mask, wherein the first encoded image is selectively transmitted over the communication channel based on the first visual fidelity metric and the second visual fidelity metric. In some implementations, the second machine learning model may be trained to extract a second type of content different from the first type of content from the one or more input images.

[0064] Figure 7 An illustrative flow chart depicting example operations 700 for training a neural network, according to some implementations, is shown. In some implementations, the example operations 700 may be performed by a machine learning system such as Figure 4 400) executed by the machine learning system.

[0065] The machine learning system generates an input image that includes content overlaid with other media (710). In some implementations, the content may include screen content. In some implementations, the content may include text, geometric shapes, or icons. The machine learning system generates a segmentation mask based on the content included in the input image (720). In some implementations, the segmentation mask may include the content and exclude the other media. In some implementations, the segmentation mask may be associated with an alpha channel of the input image. The machine learning system trains a neural network to regenerate the segmentation mask based on the input image (730).

[0066] Those skilled in the art will appreciate that information and signals may be represented using any of a variety of different technologies and techniques. For example, data, instructions, commands, information, signals, bits, symbols, and chips that may be referenced throughout the above description may be represented by voltages, currents, electromagnetic waves, magnetic fields or particles, optical fields or particles, or any combination thereof.

[0067] In addition, those skilled in the art will appreciate that the various illustrative logical blocks, modules, circuits, and algorithmic steps described in conjunction with the aspects disclosed herein can be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. A skilled person may implement the described functionality in different ways for each specific application, but such implementation decisions should not be interpreted as causing a departure from the scope of this disclosure.

[0068] The methods, sequences, or algorithms described in conjunction with the aspects disclosed herein may be implemented directly in hardware, in a software module executed by a processor, or in a combination of the two. The software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. In an alternative embodiment, the storage medium may be integrated with the processor.

[0069] In the foregoing description, the embodiments have been described with reference to specific examples thereof. However, it will be apparent that various modifications and changes may be made thereto without departing from the broader scope of the present disclosure as set forth in the appended claims. Accordingly, the specification and drawings are to be regarded in an illustrative rather than a restrictive sense.

Claims

1. A method for image compression, comprising: receiving an image for transmission via a communication channel; encoding the image into a first coded image based on a first image compression scheme; inferring a first segmentation mask from the image based on a first machine learning model; inferring a second segmentation mask from the first encoded image based on the first machine learning model; computing a first visual fidelity metric of the first encoded image based on the first segmentation mask and the second segmentation mask; as well as The first encoded image is selectively transmitted over the communication channel based at least in part on the first visual fidelity metric.

2. The method of claim 1 , wherein the first visual fidelity metric comprises a peak signal-to-noise ratio (PSNR), a PSNR based on properties of the human visual system (PSNR-HVS), a PSNR-HVS with visual masking (PSNR-HVS-M), a video multi-method assessment fusion (VMAF) metric, or a learned perceptual image patch similarity (LPIPS) metric.

3. The method of claim 1, further comprising: encoding the image into a second encoded image based on a second image compression scheme different from the first image compression scheme; inferring a third segmentation mask from the second encoded image based on the first machine learning model; computing a second visual fidelity metric for the second encoded image based on the first segmentation mask and the third segmentation mask; as well as A determination is made as to whether the first visual fidelity metric or the second visual fidelity metric indicates a higher image quality, the first encoded image being selectively transmitted over the communication channel based at least in part on whether the first visual fidelity metric or the second visual fidelity metric indicates a higher image quality.

4. The method of claim 3 , wherein the selective transmission of the first coded image comprises: Transmitting the first encoded image over the communication channel is avoided in response to determining that the second visual fidelity metric indicates a higher image quality.

5. The method of claim 3, further comprising: The second encoded image is transmitted over the communication channel instead of the first encoded image in response to determining that the second visual fidelity metric indicates a higher image quality.

6. The method of claim 1, wherein the first machine learning model is trained to extract a first type of content from one or more input images. The method of claim 6 , wherein the first type of content comprises screen content. The method of claim 6 , wherein the first type of content comprises text, a geometric shape, or an icon.

9. The method of claim 6, further comprising: inferring a third segmentation mask from the image based on a second machine learning model different from the first machine learning model; inferring a fourth segmentation mask from the first encoded image based on the second machine learning model; as well as A second visual fidelity metric of the first encoded image is calculated based on the third segmentation mask and the fourth segmentation mask, and the first encoded image is selectively transmitted over the communication channel based on the first visual fidelity metric and the second visual fidelity metric.

10. The method of claim 9, wherein the second machine learning model is trained to extract a second type of content from one or more input images that is different from the first type of content.

11. An image encoder, comprising: processing systems; as well as a memory storing instructions that, when executed by the processing system, cause the image encoder to: receive an image for transmission over a communication channel; encoding the image into a first coded image based on a first image compression scheme; inferring a first segmentation mask from the image based on a first machine learning model; inferring a second segmentation mask from the first encoded image based on the first machine learning model; computing a first visual fidelity metric of the first encoded image based on the first segmentation mask and the second segmentation mask; as well as The first encoded image is selectively transmitted over the communication channel based at least in part on the first visual fidelity metric.

12. The image encoder of claim 11 , wherein execution of the instructions further causes the image encoder to: encoding the image into a second encoded image based on a second image compression scheme different from the first image compression scheme; inferring a third segmentation mask from the second encoded image based on the first machine learning model; computing a second visual fidelity metric for the second encoded image based on the first segmentation mask and the third segmentation mask; and A determination is made as to whether the first visual fidelity metric or the second visual fidelity metric indicates a higher image quality, the first encoded image being selectively transmitted over the communication channel based at least in part on whether the first visual fidelity metric or the second visual fidelity metric indicates a higher image quality.

13. The image encoder of claim 11, wherein the first machine learning model is trained to extract a first type of content from one or more input images.

14. The image encoder of claim 13, wherein execution of the instructions further causes the image encoder to: inferring a third segmentation mask from the image based on a second machine learning model different from the first machine learning model; inferring a fourth segmentation mask from the first encoded image based on the second machine learning model; and A second visual fidelity metric of the first encoded image is calculated based on the third segmentation mask and the fourth segmentation mask, and the first encoded image is selectively transmitted over the communication channel based on the first visual fidelity metric and the second visual fidelity metric.

15. The image encoder of claim 14, wherein the second machine learning model is trained to extract a second type of content from one or more input images that is different from the first type of content.

16. A method for training a neural network, comprising: generating an input image including content overlaid with other media; generating a segmentation mask based on the content included in the input image; as well as The neural network is trained to regenerate the segmentation mask based on the input image. The method of claim 16 , wherein the content comprises text, a geometric shape, or an icon. The method of claim 16 , wherein the content comprises screen content.

19. The method of claim 16, wherein the segmentation mask includes the content and excludes the other media.

20. The method of claim 16, wherein the segmentation mask is associated with an alpha channel of the input image.