Method for Compressing Image Data for Network Transmission
By calculating the distance between the central fovea point and optimizing image compression and filtering, the problem of inefficiency in the central fovea point transmission is solved, and efficient image transmission and visual effect are improved.
Patent Information
- Application Number
- CN202210337715.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-03-31
- Filing Date
- 2022-03-31
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-03-31
AI Technical Summary
The existing fovea transmission technology has problems with inefficiency and imaging artifacts in compression and imaging filtering processing, especially in the process of image transmission, resulting in inefficiency in processing and discontinuity of visibility when the observer's eye moves.
By calculating the distance from the central fovea point to the edge of the image, the full resolution peripheral area data set is determined, and image compression and filtering are optimized using compression parameters and Gaussian filters to ensure compression uniformity in adjacent peripheral areas and removal of visual artifacts.
Improve image processing efficiency, reduce image transmission bandwidth requirements, shorten delays, and minimize visibility discontinuity when the observer's eyes move, improving image quality.
Smart Images

Figure CN115222802B_ABST
Abstract
Description
Technical Field
[0001] The following disclosure relates to the field of digital imaging and processing and more particularly to the field of foveation rendering of images and foveated transmission of images over a network. Background of the Invention
[0002] In the field of digital image rendering and transmission, it is known that when a viewer observes an image, processing efficiency can be obtained by dividing a full-resolution image into a foveal region and a peripheral region. Thus, such processing results in a reduction in the image transmission bandwidth and time requirements and a reduction in the latency from the movement of the viewer's eye to the presentation of an updated image to the viewer. During the so-called foveated transmission process, the data of the peripheral region of the image is compressed or the resolution is reduced, thereby reducing the image data size. The image is then transmitted to the client and displayed to the viewer. Further discussion of foveated transmission can be found in U.S. Patent No. 6,917,715 to Berstis and is incorporated herein by reference in its entirety.
[0003] Although foveated transmission does result in processing efficiency, there is still room for improvement by maximizing the efficiency of the compression and imaging filtering processes. Therefore, it is desirable to improve the foveated transmission technique with respect to using the correct amount of filtering and not wasting processing due to over-filtering. Improved processing can also be made to minimize imaging artifacts from the compression process and to minimize visibility discontinuities when the viewer's eye moves across the display device.
[0004] Accordingly, the present disclosure includes improvements to the compression and filtering techniques used in foveated transmission that will improve image processing efficiency and image quality. Summary of the Invention
[0005] According to a first aspect of the present invention, there is provided a system for providing an image in an image transmission system, which includes an image source server and an image display client. The image source server has a full-resolution image data set, and the image display client has a user display. The method includes the following steps:
[0006] The client determines the points of interest of the viewer in the display device;
[0007] The server divides the full-resolution image data set into a full-resolution foveal region data set and a full-resolution peripheral region data set by the following method:
[0008] Using the points of interest to determine the position of the foveal point on the display device and the full-resolution foveal region data set;
[0009] Calculating the distance from the foveal point to the edge of the full-resolution image data set along at least one axis of the full-resolution image data set;
[0010] Calculate the distance from each edge of the full-resolution dataset to the nearest point on the foveal region surrounding the foveal point; and
[0011] Determine full-resolution peripheral datasets, which are composed of adjacent peripheral regions of the full-resolution dataset;
[0012] Compress the full-resolution peripheral datasets by reducing the image resolution, which is expressed as:
[0013] Use compression parameters to calculate the distribution of available space in the adjacent peripheral regions of the image,
[0014] Calculate a compression curve, where the compression of two adjacent peripheral regions is such that no side of the adjacent peripheral region is compressed more than the other side;
[0015] Use a mapping function to map the compressed image to the uncompressed image;
[0016] Use the compression curve and its derivative to filter the full-resolution peripheral datasets using a Gaussian filter to remove potential visual artifacts that may occur due to image compression;
[0017] Send the full-resolution foveal region dataset and the reduced-resolution peripheral dataset to the client; and
[0018] Display an image to the user through the client based on the transmitted datasets, the image including a full-resolution foveal region at the current eye position surrounded by a reduced-resolution peripheral region.
[0019] A system according to a second embodiment of the present invention calculates the distance from the foveal point to the edge of the full-resolution image dataset along two or more axes of the full-resolution image dataset.
[0020] According to a third embodiment of the present invention, the length of the foveal region is input as a parameter.
[0021] According to a fourth embodiment of the present invention, if the foveal point is closer to the edge of the image than half the length of the foveal region, the foveal point is adjusted so that the foveal region leaving the image edge has half the length.
[0022] According to a fifth embodiment of the present invention, calculating the distribution of available space within the adjacent peripheral regions of the image includes: calculating the length of the output image by dividing the size of the full-resolution dataset by the square root of the compression level, where the compression level is input as a parameter, and subtracting the length of the foveal region from the total length of the output image to obtain the available space into which the peripheral region is compressed.
[0023] According to a sixth embodiment of the present invention, the mapping function includes mapping the center of an individual pixel of the compressed image to a point on the uncompressed image, and wherein the mapping function has a computable derivative that describes any point on the output image, and one or more points on the original image from which it has been compressed.
[0024] According to a seventh embodiment of the present invention, the mapping function has a computable derivative that describes any point on the output image, and one or more points on the original image from which it has been compressed.
[0025] According to an eighth embodiment of the present invention, the core size is controlled by the derivative of the compression curve.
[0026] According to a ninth embodiment of the present invention, the system includes filtering using a Gaussian filter with a core size determined by the derivative of the mapping function to further reduce visual artifacts caused by compression.
[0027] According to a tenth embodiment of the present invention, the step of determining a point of interest of an observer in a display device includes providing an eye movement tracking system associated with an image display client. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] The following detailed description provides a complete disclosure of the present invention when taken in conjunction with the accompanying drawings presented herein.
[0029] Figure 1 Shows an arrangement of system steps according to a preferred embodiment;
[0030] Figure 2 A diagram depicting the foveal point on a display;
[0031] Figure 3 A diagram depicting the foveal region and an adjacent peripheral region along an axis;
[0032] Figure 4 A diagram depicting the mapping function of the compressed image and the original image;
[0033] Figure 5 A diagram depicting the compressed image from which points not intended to be filtered have been removed;
[0034] Figure 6 Is a diagram of the inverse mapping function;
[0035] Figure 7 Depicts the original image before compression; and
[0036] Figure 8 Depicts the transformed image that has been reconstructed from the compressed image. DETAILED DESCRIPTION
[0037] Foveal rendering and foveal transmission allow for improved rendering performance by reducing the resolution of portions of an image within the peripheral region of an observer's vision. As a result, the observer will not notice the resolution degradation within their peripheral vision. The resolution reduction is achieved through compression and filtering algorithms that "discard" "unnecessary resolution, such as by simply not transmitting every other sample value of two sample values to reduce sharpness by 50%.
[0038] In one embodiment, to compress a given original full-resolution image, the system will first determine the location 100 of the fovea on the device's display. Fovea detection can be performed by an eye-tracking system associated with the image display client, which is well-known and will not be further discussed herein. After determining the fovea 210 on the display 200, the system will begin the process of splitting the full-resolution image into a foveal region that remains unchanged and a peripheral region that is to be compressed. In one embodiment, this splitting process is performed separately for the X-axis and Y-axis of the image.
[0039] Splitting the foveal region from the peripheral region can be achieved by first calculating the distance 101 from the fovea 210 to the edge of the image 220. In Figure 3 this, the distances are respectively denoted as L 左 and L 右 . The foveal region L 中央凹区 is then defined as a given length around the fovea 300. The length of the foveal region can be input as a parameter. If the fovea is closer to the edge of the image than half the length of the foveal region, the fovea is adjusted so that the fovea is half the length away from the image edge. Then the distance 103 from each edge of the image to the nearest point of the foveal region is calculated.
[0040] Then the length 105 of the output image is determined by dividing the original image size by a function of the input compression level. The compression level is an input parameter of the system and in one embodiment, the function of the compression level is the square root. Then the length of the foveal region is subtracted from the total length of the output image to obtain the available space 106 into which the peripheral region is to be compressed. The distribution of the available space within the region adjacent to the foveal region is determined using compression parameters 107.
[0041] Then the compression curve is calculated in such a way that regardless of the distribution, the visual distribution on both sides adjacent to the foveal region is similar, i.e., it is not compressed more on either the smaller side or the larger side 108. This effect can be seen in Figure 8 wherein, compared to the unchanged full image in Figure 7 , the foveal region remains unchanged while the peripheral region has been compressed.
[0042] In one embodiment, the mapping function then maps the center of each individual pixel of the compressed image to a point on the uncompressed image, which may or may not be centered on a pixel 109. The mapping function may have a computable derivative that describes any point on the output image, one or several points on the original image from which it has been compressed, as Figure 4 shown.
[0043] Using the compression curve and its derivative, a Gaussian filter is then used to filter the original image 110, where the kernel size is controlled by the derivative. This step is to remove potential visual artifacts that may occur due to the compression of the image. In addition, this step only outputs a plurality of pixels that will be used in the compression step to form the output image. As Figure 5 shown, the compressed peripheral region has "discarded" the image data represented by the black line 500, which means that the filtering step does not need to filter the lost data.
[0044] For each pixel of the compressed output image, the system will find the position on the original image to which the pixel maps and sample that point 111. In general, this will mean interpolating between 4 adjacent pixels. This can be seen in Figure 6 After that, the compressed image is transmitted to the client application 112. The original image can be reconstructed from the compressed image using the inverse mapping function. In one embodiment, a further Gaussian filter can be applied using the kernel size specified by the derivative of the mapping function. This can be used as a final step to further reduce visual artifacts caused by compression.
[0045] Those skilled in the art will recognize that there are many advantages to the methods and techniques described herein, including but not limited to the following advantages.
[0046] Using separate functions in the system: The use of separate functions allows the system to significantly optimize the processing of the required filtering. In addition, separate functions allow the system to produce a rectangular image with a guaranteed size, which can then be video compressed.
[0047] Scaling the filter kernel by the magnitude of the derivative of the remapping curve: This ensures that the system uses the correct amount of filtering at each point just right. As a result, no artifacts are left before video compression, and the system does not waste processing on over-filtering.
[0048] The system uses a remapping curve that has a derivative of 1 at the minimum of the curve function, which ensures as smooth an attenuation of the image quality as possible and minimizes visible discontinuities as the eye moves. Additionally, the system uses a remapping curve based on the tangent function, which allows the system to take into account the tangential arrangement of the display and the tangential attenuation of visual sharpness. This feature is particularly suitable for gaze tracking applications within VR / AR, head-mounted display devices (HMDs), or large-screen displays.
[0049] Although some details of the preferred embodiments have been disclosed with illustrative examples, those skilled in the art will recognize that certain alternative embodiments can be implemented without departing from the spirit and scope of the present invention, including but not limited to using alternative computing platforms and programming methods, introducing other eye movement tracking techniques or systems, and using other computing networks and imaging devices. Accordingly, the scope of the present invention should be determined by the following claims.
Claims
1. A method for providing an image in an image transmission system including an image source server and an image display client, the image source server having a full-resolution image data set and the image display client having a user display, the method comprising the steps of: Determining, by the client, a point of interest to an observer in the display device; Segmenting, by the server, the full-resolution image data set into a full-resolution foveal region data set and a full-resolution peripheral region data set by: Using the point of interest to determine the position of the foveal point on the display device and the full-resolution foveal region data set; Calculating, along at least one axis of the full-resolution image data set, the distance from the foveal point to the edge of the full-resolution image data set; Calculating the distance from each edge of the full-resolution data set to the nearest point on the foveal region surrounding the foveal point; And Determining the full-resolution peripheral region data sets, which consist of adjacent peripheral regions from the full-resolution data set; Compressing the full-resolution peripheral data set by reducing the image resolution, expressed as: Using compression parameters to calculate the distribution of available space in the adjacent peripheral regions of the image; Calculating a compression curve, wherein the compression of two adjacent peripheral regions is such that neither side of the adjacent peripheral regions is compressed more than the other; Using a mapping function to map the compressed image to the uncompressed image; Using the compression curve and its derivative, and using a Gaussian filter to filter the full-resolution peripheral data set to remove potential visual artifacts that may be caused by image compression; Sending the full-resolution foveal region data set and the reduced-resolution peripheral data set to the client; And Displaying, by the client, an image to the user according to the transmitted data sets, the image including a full-resolution foveal region at the current eye position surrounded by a reduced-resolution peripheral region.
2. The method according to claim 1, wherein the distance from the foveal point to the edge of the full-resolution image data set is calculated along two or more axes of the full-resolution image data set.
3. The method according to claim 1, wherein the length of the foveal region is input as a parameter.
4. The method according to claim 1, wherein if the foveal point is closer to the edge of the image than half the length of the foveal region, the foveal point is adjusted so that the foveal region away from the image edge has a length of half.
5. The method according to claim 1, wherein calculating the distribution of available space in the adjacent peripheral region of the image comprises: Calculating the length of the output image by dividing the size of the full-resolution data set by a function of the compression level, wherein the compression level is input as a parameter, and subtracting the length of the foveal region from the total length of the output image to obtain the available space into which the peripheral region is compressed.
6. The method according to claim 5, wherein the length of the output image is calculated by dividing the full-resolution data set by the square root of the compression level.
7. The method according to claim 1, wherein the mapping function includes mapping the center of an individual pixel of the compressed image to a point on the uncompressed image, and wherein the mapping function has a computable derivative that describes one or more points on the original image from which any point on the output image has been compressed.
8. The method according to claim 1, wherein the mapping function has a computable derivative that describes any point on the output image and one or more points on the original image from which it has been compressed.
9. The method according to claim 1, wherein the core size is controlled by the derivative of the compression curve.
10. The method according to claim 1, further comprising filtering using the applied Gaussian filter with a core size determined by the derivative of the mapping function to further reduce visual artifacts caused by compression.
11. The method according to claim 1, wherein the step of determining the point of interest of the observer in the display device comprises providing an eye movement tracking system associated with the image display client.
12. The method according to claim 1, wherein the display device is an HMD.
13. A software - encoded computer - readable medium for providing an image in an image transmission system comprising an image source server and an image display client, the image source server having a full - resolution image data set and the image display client having a user display, the software causing the server and the client to perform the following steps: The client determines the points of interest of the observer in the display device; The server partitions the full - resolution image data set into a full - resolution foveal region data set and a full - resolution peripheral region data set by: Using the point of interest to determine the position of the foveal point on the display device and the full - resolution foveal region data set; Calculating the distance from the foveal point to the edge of the full - resolution image data set along at least one axis of the full - resolution image data set; Calculating the distance from each edge of the full - resolution data set to the nearest point on the foveal region around the foveal point; And Determining the full - resolution peripheral region data sets, which are composed of adjacent peripheral regions from the full - resolution data set; Compressing the full - resolution peripheral data sets by reducing the image resolution, expressed as: Using compression parameters to calculate the distribution of available space in the adjacent peripheral regions of the image; Calculating a compression curve such that the compression of two adjacent peripheral regions does not cause any side of an adjacent peripheral region to be compressed more than the other side; Using a mapping function to map the compressed image to the uncompressed image; Using the compression curve and its derivative, using a Gaussian filter to filter the full - resolution peripheral data sets to remove potential visual artifacts that may be caused by image compression; Sending the full - resolution foveal region data set and the reduced - resolution peripheral data sets to the client; And Displaying an image to the user by the client according to the transmitted data sets, the image comprising a full - resolution foveal region at the current eye position surrounded by a reduced - resolution peripheral region.
14. The computer - readable medium according to claim 13, wherein the distance from the foveal point to the edge of the full - resolution image data set is calculated along two or more axes of the full - resolution image data set.
15. The computer - readable medium according to claim 13, wherein the length of the foveal region is input as a parameter.
16. The computer-readable medium according to claim 13, wherein if the foveal point is closer to the edge of the image than half the length of the foveal region, the foveal point is adjusted so that the foveal region away from the image edge has a length of half.
17. The computer-readable medium according to claim 13, wherein calculating the distribution of available space within an adjacent peripheral region of the image comprises: The length of the output image is calculated by dividing the size of the full-resolution data set by a function of the compression level, where the compression level is input as a parameter, and the length of the foveal region is subtracted from the total length of the output image to obtain the available space into which the peripheral region is compressed.
18. The computer-readable medium according to claim 17, wherein the length of the output image is calculated by dividing the full-resolution data set by the square root of the compression level.
19. The computer-readable medium according to claim 13, wherein the mapping function includes mapping the center of an individual pixel of the compressed image to a point on the uncompressed image, and wherein the mapping function has a computable derivative that describes any point on the output image and one or more points on the original image from which it has been compressed.
20. A system for remotely viewing an image, the system comprising: A local device for determining a point of interest of an observer on a display device; A remote full-resolution image data set splitter for splitting a remote full-resolution image into a full-resolution foveal region data set and a full-resolution peripheral region data set by a server by the following method: Using the point of interest to determine the position of the foveal point on the display device and the full-resolution foveal region data set; Calculating the distance from the foveal point to the edge of the full-resolution image data set along at least one axis of the full-resolution image data set; Calculating the distance from each edge of the full-resolution data set to the nearest point on the foveal region around the foveal point and determining the full-resolution peripheral region data sets, which are composed of adjacent peripheral regions from the full-resolution data set; A compressor for reducing the resolution of the image represented by the full-resolution peripheral data set by the following method to produce a reduced-resolution peripheral data set: Using the compression parameter to calculate the distribution of the available space in the adjacent peripheral regions of the image; Calculating a compression curve, wherein the compression of two adjacent peripheral regions is such that neither side of the adjacent peripheral regions is compressed more than the other side; Using a mapping function to map the compressed image to the uncompressed image; Using the compression curve and its derivative, using a Gaussian filter to filter the full-resolution peripheral data set to remove potential visual artifacts that may be caused by image compression; A transmission network for receiving the full-resolution foveal region data set and transmitting the reduced-resolution peripheral data set to a local client; And A local client image renderer for displaying an image to a user according to the transmitted data sets, the image including a full-resolution foveal region at the current eye position surrounded by a reduced-resolution peripheral region.
Citation Information
Patent Citations
Foveal priority in stereoscopic remote viewing system
US6917715B2
System and method for displaying a stream of images
CN107852521A
Wearable image manipulation and control system with correction for vision defects and augmentation of vision and sensing
WO2018200717A1