Image data transfer device and image compression method
The image data transfer device addresses latency issues in high-definition image transmission by estimating attention levels and varying compression ratios, ensuring reduced latency and high image quality.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2019-09-30
- Publication Date
- 2026-03-30
AI Technical Summary
The delay in image display due to communication between client terminals and servers, particularly in high-definition moving image transmission, can cause impaired user experience, including sense of presence and motion sickness in head-mounted displays.
An image data transfer device and method that estimates attention levels in image frames, varies compression ratios based on these levels, and transmits compressed data, enabling reduced latency and high image quality.
Achieves both high image quality and reduced latency in image display by optimizing data transmission through adaptive compression and encoding.
Smart Images

Figure 0007837134000002 
Figure 0007837134000003 
Figure 0007837134000004
Abstract
Description
Technical Field
[0001] This invention relates to an image data transfer device and an image compression method for processing moving image data of a display target.
Background Art
[0002] With the recent improvement of information processing technology and image display technology, it has become possible to experience the video world in various forms. For example, by displaying a panoramic video on a head-mounted display and displaying an image corresponding to the user's line of sight, the immersion in the video world can be enhanced, or the operability of applications such as games can be improved. In addition, by displaying image data streamed from a server having abundant resources, users can enjoy high-definition moving images and game screens regardless of location and scale.
Summary of the Invention
Problems to be Solved by the Invention
[0003] In the technology of instantaneously displaying image data of an image transmitted via a network on a client terminal, the delay time due to communication between the client terminal and the server can be a problem. For example, when reflecting a user operation on the client terminal side in the displayed image, data transfer such as transmission of the user operation to the server and transmission of image data from the server to the client terminal is required, which may cause an unacceptable delay time. When the head-mounted display is the display destination, it is also conceivable that the sense of presence is impaired or motion sickness is caused due to the delay in display with respect to the movement of the user's head. This problem is more likely to become apparent as higher image quality is pursued.
[0004] This invention has been made in view of such problems, and an object thereof is to provide a technology capable of achieving both reduction of image quality and delay time in image display involving data transmission by communication.
Means for Solving the Problems
[0005] To solve the above problems, one aspect of the present invention relates to an image data transfer device. This image data transfer device is characterized by comprising: a drawing unit that draws a moving image to be displayed; an image content acquisition unit that acquires information relating to the content represented by the moving image; an attention level estimation unit that estimates the level of attention according to the content represented by the moving image for each unit region formed by dividing the frame plane of the moving image; a compression encoding unit that compresses and encodes the data of the moving image by varying the compression ratio in the frame plane based on the distribution of attention levels; and a communication unit that transmits the compressed and encoded data of the moving image.
[0006] Another aspect of the present invention relates to an image compression method. This image compression method is characterized in that an image data transfer device includes the steps of: drawing a moving image to be displayed; acquiring information relating to the content represented by the moving image; estimating the degree of attention for each unit region obtained by dividing the frame plane of the moving image according to the content represented by the moving image; compressing and encoding the data of the moving image by varying the compression ratio in the frame plane based on the distribution of the degree of attention; and transmitting the compressed and encoded data of the moving image.
[0007] Furthermore, any combination of the above components, as well as conversions of the expression of the present invention between methods, apparatus, systems, computer programs, data structures, recording media, etc., are also valid embodiments of the present invention. [Effects of the Invention]
[0008] According to the present invention, it is possible to achieve both high image quality and reduced latency in image display involving data transmission via communication. [Brief explanation of the drawing]
[0009] [Figure 1] This figure shows an example of the configuration of the image processing system in this embodiment. [Figure 2] This figure shows an example of the appearance of the head-mounted display according to this embodiment. [Figure 3]This figure shows the basic configuration of the server and image processing device in this embodiment. [Figure 4] This diagram conceptually illustrates the process from image drawing to display in this embodiment. [Figure 5] This diagram shows the functional blocks of the server and image processing device according to this embodiment. [Figure 6] This diagram illustrates the effect of the server and image processing device performing pipeline processing on a partial image basis in this embodiment. [Figure 7] This figure illustrates the data transmission status of a partial image between the server and the image processing device in this embodiment. [Figure 8] This flowchart shows an example of a processing procedure in which the display control unit, in this embodiment, outputs partial image data to the display panel while adjusting the output target and output timing. [Figure 9] This diagram shows the configuration of the functional block of an image processing apparatus having a reprojection function in this embodiment. [Figure 10] This figure illustrates the reprojection and distortion correction for the eyepiece performed by the first correction unit in this embodiment. [Figure 11] This figure illustrates an example of the procedure for the correction process performed by the first correction unit in this embodiment. [Figure 12] This flowchart shows the processing procedure in which the output target determination unit of the display control unit adjusts the output target when the image processing device performs reprojection in this embodiment. [Figure 13] Figure 12, S94, is a diagram illustrating a method for quantifying the degree of data loss based on the user's perspective. [Figure 14] This figure illustrates the data required for reprojection, which is evaluated in S96 of Figure 12. [Figure 15] This flowchart shows the procedure for processing performed by the server and the image processing device when reprojection is performed in the image processing device of this embodiment. [Figure 16] This is a diagram showing the configuration of functional blocks of a server and an image processing apparatus that can support display on a plurality of display devices with different forms in this embodiment. [Figure 17] This is a diagram illustrating the transition of image formats that can be realized in this embodiment. [Figure 18] This is a diagram illustrating variations in the connection method of a display device on the user side (client side) in this embodiment. [Figure 19] This is a flowchart showing the processing procedure in which the server transmits image data in a format determined according to the display form in this embodiment. [Figure 20] This is a diagram showing the configuration of functional blocks of a server having a function of highly compressing images of multiple viewpoints and an image processing apparatus that processes the same in this embodiment. [Figure 21] This is a diagram illustrating the relationship between the image blocks to be compression-encoded by the first encoding unit and the second encoding unit and the image blocks that refer to the compression-encoding results for that purpose in this embodiment. [Figure 22] This is a diagram for explaining the effect of the server in this embodiment performing compression encoding in units of image blocks by utilizing the similarity of images of multiple viewpoints. [Figure 23] This is a diagram showing the configuration of functional blocks of a compression encoding unit when scalable video encoding is performed in this embodiment. [Figure 24] This is a diagram illustrating the relationship between the image blocks to be compression-encoded by the first encoding unit and the second encoding unit and the image blocks that refer to the compression-encoding results for that purpose when scalable video encoding is performed in this embodiment. [Figure 25] This is a diagram illustrating the relationship between the image blocks to be compression-encoded by the first encoding unit and the second encoding unit and the image blocks that refer to the compression-encoding results for that purpose when scalable video encoding is performed in this embodiment. [Figure 26] This is a diagram showing the configuration of functional blocks of a server having a function of optimizing data size reduction means in this embodiment. [Figure 27] This is a diagram showing an example of determining a score according to the content of a moving image in the present embodiment. [Figure 28] This is a diagram showing an example of determining a score according to the content of a moving image in the present embodiment. [Figure 29] This is a diagram showing an example of determining a score according to the content of a moving image in the present embodiment. [Figure 30] This is a diagram showing an example of determining a score according to the content of a moving image in the present embodiment. [Figure 31] This is a diagram showing an example of determining a score according to the content of a moving image in the present embodiment. [Figure 32] This is a flowchart showing the processing procedure for a server to adjust the data size according to the communication situation in the present embodiment. [Figure 33] This is a diagram showing the configuration of the functional blocks of a server having a function of changing the compression rate according to the region on the frame based on the content represented by the moving image in the present embodiment. [Figure 34] This is a diagram for explaining the process of the attention degree estimation unit estimating the distribution of the attention degree in the image plane in the present embodiment. [Figure 35] This is a flowchart showing the processing procedure for a server to control the compression rate for each region of the image plane in the present embodiment. [Figure 36] This is a diagram showing the configuration of the functional blocks of a server having a function of changing the compression rate according to the region on the frame based on the user's gaze point in the present embodiment. [Figure 37] This is a diagram for explaining the process of the attention degree estimation unit estimating the distribution of the attention degree in the image plane in the present embodiment. [Figure 38] This is a diagram for explaining the method by which the compression encoding processing unit determines the distribution of the compression rate based on the gaze point in the present embodiment.
Embodiments for Carrying Out the Invention
[0010] 1. Overall Configuration of the System Figure 1 shows an example configuration of the image processing system in this embodiment. The image display system 1 includes an image processing device 200, a head-mounted display 100, a flat panel display 302, and a server 400. The image processing device 200 is connected to the head-mounted display 100 and the flat panel display 302 by wireless communication or an interface 300 such as USB Type-C or HDMI®. The image processing device 200 is further connected to the server 400 via a network 306 such as the Internet or a LAN (Local Area Network).
[0011] Server 400, acting as an image data transfer device, generates at least a portion of the image to be displayed and transmits it to the image processing device 200. Here, Server 400 may be a server of a company that provides various distribution services such as cloud games, or it may be a home server that transmits data to any terminal. Therefore, the network 306 is not limited in scale and can be a public network such as the internet or a LAN (Local Area Network). For example, Network 306 may be a mobile phone carrier network, a Wi-Fi hotspot in a city, or a home Wi-Fi access point. Alternatively, the image processing device 200 and Server 400 may be directly connected by a video interface.
[0012] The image processing device 200 processes the image data transmitted from the server 400 as needed and outputs it to at least one of the head-mounted display 100 and the flat panel display 302. For example, the server 400 receives head movements and user operations from multiple image processing devices 200 connected to each head-mounted display 100, each of which is wearing a head-mounted display 100. The server 400 then renders a virtual world that has been changed according to the user operations, using the field of view corresponding to each user's head movement, and transmits it to each image processing device 200.
[0013] The image processing device 200 converts the transmitted image data into a format suitable for the head-mounted display 100 or the flat panel display 302 as needed, and outputs it to the head-mounted display 100 or the flat panel display 302 at the appropriate time. By repeating this process for each frame of the video, a cloud game system with multiple users can be realized. In this case, the image processing device 200 may also combine the image transmitted from the server 400 with a separately prepared UI (User Interface) plane image (also called an OSD (On Screen Display) plane image) or an image captured by the camera of the head-mounted display 100, before outputting it to the head-mounted display 100 or the flat panel display 302.
[0014] The image processing device 200 may also improve the display tracking ability with respect to head movements by correcting the image transmitted from the server 400 based on the position and orientation of the head-mounted display 100 immediately before display. The image processing device 200 may also display the same field of view on the flat panel display 302 so that others can see what image the user wearing the head-mounted display 100 is looking at.
[0015] However, the content of the moving images to be displayed and the destination of their display are not particularly limited in this embodiment. For example, the server 400 may display images captured by a camera (not shown) and live stream them to the image processing device 200. In this case, the server 400 may acquire multi-viewpoint images taken by multiple cameras at an event venue such as a sports competition or concert, and use these to create images in the field of view corresponding to the movement of the head-mounted display 100, thereby generating free-viewpoint live video and distributing it to each image processing device 200.
[0016] Furthermore, the system configuration to which this embodiment can be applied is not limited to that shown in the illustration. For example, the display device connected to the image processing device 200 may be either a head-mounted display 100 or a flat panel display 302, or it may be multiple head-mounted displays 100. Also, the image processing device 200 may be built into the head-mounted display 100 or the flat panel display 302. For example, the flat panel display and the image processing device may be integrated into a personal computer or mobile terminal (portable game console, high-function mobile phone, tablet terminal).
[0017] These devices may be further configured to connect at least one of a head-mounted display 100 and a flat panel display 302 as needed. The image processing device 200 and these terminals may have built-in or connected input devices (not shown). The number of image processing devices 200 connected to the server 400 is also not limited. Furthermore, the server 400 may receive the operations of multiple users viewing their respective flat panel displays 302 from multiple image processing devices 200 connected to each flat panel display 302, generate corresponding images, and transmit them to each image processing device 200.
[0018] Figure 2 shows an example of the appearance of the head-mounted display 100. In this example, the head-mounted display 100 consists of an output mechanism 102 and a mounting mechanism 104. The mounting mechanism 104 includes a mounting band 106 that wraps around the user's head to secure the device. The output mechanism 102 includes a housing 108 shaped to cover the left and right eyes when the user wears the head-mounted display 100, and has a display panel inside that faces the eyes when worn.
[0019] The housing 108 further includes an eyepiece positioned between the display panel and the user's eyes when the head-mounted display 100 is worn, which magnifies the image. The head-mounted display 100 may also be equipped with speakers or earphones positioned to correspond to the user's ears when worn. The head-mounted display 100 may also incorporate motion sensors to detect the translational and rotational movements of the user's head, as well as its position and orientation at each moment in time.
[0020] The head-mounted display 100 further includes a stereo camera 110 on the front of the housing 108, a wide-angle monocular camera 111 in the center, and four wide-angle cameras 112 in the four corners: upper left, upper right, lower left, and lower right, to capture video of the real space in the direction corresponding to the user's face orientation. In one embodiment, the head-mounted display 100 provides a see-through mode that shows the real space in the direction the user is facing directly by immediately displaying the video images captured by the stereo camera 110.
[0021] Furthermore, at least one of the images captured by the stereo camera 110, the monocular camera 111, or the four cameras 112 may be used to generate the display image. For example, the position and orientation of the head-mounted display 100, and by extension the user's head, relative to the surrounding space may be acquired at a predetermined rate using SLAM (Simultaneous Localization and Mapping), and the field of view of the image to be generated in the server 400 may be determined, or the image processing device 200 may correct the image. Alternatively, the image processing device 200 may combine the captured image with the image transmitted from the server 400 to create the display image.
[0022] The head-mounted display 100 may also be equipped with motion sensors such as an accelerometer, gyroscope, or geomagnetic sensor to determine the position, orientation, and movement of the head-mounted display 100. In this case, the image processing device 200 acquires information on the position and orientation of the user's head at a predetermined rate based on the measurements of the motion sensors. This information can be used to determine the field of view of the image generated by the server 400 or to correct the image in the image processing device 200.
[0023] Figure 3 shows the basic configuration of the server 400 and image processing device 200 in this embodiment. The server 400 and image processing device 200 in this embodiment are equipped with local memory in key locations to store partial images smaller than one frame of the displayed image. The compression encoding and transmission of image data in the server 400, and the reception, decoding and decompression of data, various image processing, and output to the display device in the image processing device 200 are pipelined in units of these partial images. This reduces the delay time from image rendering in the server 400 to display on the display device connected to the image processing device 200.
[0024] In the server 400, the rendering control unit 402 is implemented by the CPU (Central Processing Unit) and controls the rendering of images in the image rendering unit 404. As described above, the content of the images to be displayed in this embodiment is not particularly limited, but the rendering control unit 402 may, for example, advance a cloud game and cause the image rendering unit 404 to render frames of moving images representing the results. In this case, the rendering control unit 402 may acquire information related to the position and posture of the user's head from the image processing device 200 and control the rendering of each frame in the field of view corresponding to that information.
[0025] The image rendering unit 404 is implemented by a GPU (Graphics Processing Unit) and, under the control of the rendering control unit 402, renders frames of the moving image at a predetermined or variable rate and stores the results in the frame buffer 406. The frame buffer 406 is implemented by RAM (Random Access Memory). The video encoder 408, under the control of the rendering control unit 402, compresses and encodes the image data stored in the frame buffer 406 in units of partial images smaller than one frame. A partial image is an image of each region obtained by dividing the image plane of a frame into predetermined sizes. That is, a partial image is an image of each region obtained by dividing the image plane by boundaries set horizontally, vertically, bidirectionally and vertically, or diagonally.
[0026] In this case, the video encoder 408 may start the compression encoding of a frame without waiting for the server's vertical synchronization signal as soon as the image drawing unit 404 has finished drawing one frame. According to conventional technology, which synchronizes various processes such as frame drawing and compression encoding based on the vertical synchronization signal, frame order management is easy by aligning the time allocated to each process from image drawing to display on a frame-by-frame basis. However, in this case, even if the drawing process finishes early depending on the content of the frame, the compression encoding process must wait until the next vertical synchronization signal. In this embodiment, as will be described later, unnecessary waiting time is avoided by managing the generation time on a partial image basis.
[0027] The video encoder 408 can use any common encoding scheme for compression coding, such as H.264 / AVC or H.265 / HEVC. The video encoder 408 stores the compressed and coded partial image data in the partial image storage unit 410. The partial image storage unit 410 is a local memory implemented as SRAM (Static Random Access Memory) or the like, and has a storage area corresponding to the data size of a partial image smaller than one frame. The "partial image storage unit" described later is similar. The video stream control unit 414 reads the compressed and coded partial image data each time it is stored in the partial image storage unit 410, and packets it after including audio data and control information as necessary.
[0028] The control unit 412 constantly monitors the data writing status of the video encoder 408 to the partial image storage unit 410 and the data reading status of the video stream control unit 414, and appropriately controls the operation of both. For example, the control unit 412 controls the partial image storage unit 410 to prevent data shortage, i.e., buffer underrun, and data overflow, i.e., buffer overrun.
[0029] The input / output interface 416 establishes communication with the image processing device 200, and the video stream control unit 414 sequentially transmits the packetized data via the network 306. In addition to image data, the input / output interface 416 may also transmit audio data as appropriate. Furthermore, as described above, the input / output interface 416 may acquire information related to user operation and the position and posture of the user's head from the image processing device 200 and supply it to the drawing control unit 402. In the image processing device 200, the input / output interface 202 sequentially acquires image and audio data transmitted from the server 400.
[0030] The input / output interface 202 may also appropriately acquire information related to user operation and the position and posture of the user's head from the head-mounted display 100 or an input device (not shown) and transmit it to the server 400. The input / output interface 202 decodes the packets acquired from the server 400 and stores the extracted image data in the partial image storage unit 204. The partial image storage unit 204 is a local memory located between the input / output interface 202 and the video decoder 208, and constitutes a compressed data storage unit. The control unit 206 constantly monitors the data writing status of the input / output interface 202 to the partial image storage unit 204 and the data reading status of the video decoder 208, and appropriately controls the operation of both.
[0031] The video decoder 208, acting as a decoding and decompression unit, reads the partial image data each time it is stored in the partial image storage unit 204, decodes and decompresses it according to the encoding scheme, and then sequentially stores it in the partial image storage unit 210. The partial image storage unit 210 is a local memory located between the video decoder 208 and the image processing unit 214, and constitutes the decoded data storage unit. The control unit 212 constantly monitors the data writing status of the video decoder 208 to the partial image storage unit 210 and the data reading status of the image processing unit 214, and appropriately controls the operation of both.
[0032] The image processing unit 214 reads the decoded and decompressed partial image data each time it is stored in the partial image storage unit 210 and performs the necessary processing for display. For example, in the head-mounted display 100, in order to allow the viewer to see a distortion-free image when viewed through the eyepiece, a correction process is performed that applies a distortion opposite to the distortion caused by the eyepiece.
[0033] Alternatively, the image processing unit 214 may refer to a separately prepared UI plane image and superimpose it onto the image transmitted from the server 400. The image processing unit 214 may also superimpose an image captured by the camera of the head-mounted display 100 onto the image transmitted from the server 400. The image processing unit 214 may also correct the image transmitted from the server 400 so that it corresponds to the field of view of the user's head position and posture at the time of processing. The image processing unit 214 may also perform image processing suitable for output to the flat panel display 302, such as super-resolution processing.
[0034] In any case, the image processing unit 214 processes the partial images stored in the partial image storage unit 210 and sequentially stores them in the partial image storage unit 216. The partial image storage unit 216 is a local memory located between the image processing unit 214 and the display controller 220. The control unit 218 constantly monitors the data writing status of the image processing unit 214 to the partial image storage unit 216 and the data reading status of the display controller 220, and appropriately controls the operation of both.
[0035] The display controller 220 reads the data of a processed partial image from the partial image storage unit 216 each time the data is stored and outputs it to the head-mounted display 100 or the flat panel display 302 at the appropriate timing. Specifically, it outputs the data of the uppermost partial image of each frame at a timing that matches the vertical synchronization signal of those displays, and then sequentially outputs the partial image data downwards.
[0036] 2. Pipeline processing of each partial image Next, we will describe the partial image pipeline processing implemented in the server 400 and image processing device 200 from image drawing to display. Figure 4 conceptually shows the processing from image drawing to display in this embodiment. As described above, the server 400 generates moving image frames 90 at a predetermined or variable rate. In the illustrated example, frame 90 has a configuration in which images for the left eye and right eye are represented in areas divided into left and right halves, but this does not mean that the configuration of images generated by the server 400 is limited to this.
[0037] As described above, the server 400 compresses and encodes frame 90 into individual sub-images. In the diagram, the image plane is divided horizontally into five sections, which are sub-images 92a, 92b, 92c, 92d, and 92e. As a result, the sub-images are compressed and encoded one after another in this order and transmitted to the image processing device 200 for display, as shown by the arrows. That is, while the uppermost sub-image 92a is being processed—compression, transmission, decoding / decompression, and output to the display panel 94—the sub-images below it, 92b, and then 92c, are transmitted and displayed sequentially. This allows various processes necessary from image rendering to display to be performed in parallel, enabling display to progress with minimal delay even with transmission time involved.
[0038] Figure 5 shows the functional blocks of the server 400 and image processing device 200 in this embodiment. Each functional block shown in the figure can be implemented in hardware terms using a CPU, GPU, encoder, decoder, arithmetic unit, and various types of memory, and in software terms using a program that performs various functions such as information processing, image rendering, data input / output, and communication, which is loaded from the recording medium into memory. Therefore, it will be understood by those skilled in the art that these functional blocks can be implemented in various ways using hardware alone, software alone, or a combination thereof, and are not limited to any one of these. The same applies to the functional blocks described later.
[0039] The server 400 comprises an image generation unit 420, a compression encoding unit 422, a packetization unit 424, and a communication unit 426. The image generation unit 420 consists of a drawing control unit 402, an image drawing unit 404, and a frame buffer 406 as shown in Figure 3, and generates frames of moving images to be transmitted to the image processing device 200, such as game images, at a predetermined or variable rate. Alternatively, the image generation unit 420 may acquire moving image data from a camera or storage device (not shown). In this case, the image generation unit 420 can be read as an image acquisition unit. The same applies in the following description.
[0040] The compression encoding unit 422 consists of the video encoder 408, partial image storage unit 410, and control unit 412 shown in Figure 3, and compresses and encodes the image data generated by the image generation unit 420 in units of partial images. Here, the compression encoding unit 422 performs motion compensation and encoding in units of a predetermined number of rows, such as one row or two rows, or rectangular areas of a predetermined size, such as 16x16 pixels or 64x64 pixels. Therefore, the compression encoding unit 422 may start compression encoding once the image generation unit 420 has generated the data for the smallest unit area required for compression encoding.
[0041] The partial image, which is the unit of pipeline processing in compression encoding and transmission, may be the same as or larger than the area of the smallest unit. The packetization unit 424 consists of the video stream control unit 414 and control unit 412 shown in Figure 3, and packets the compressed and encoded partial image data in a format according to the communication protocol used. At this time, the time when the partial image was drawn (hereinafter referred to as the "generation time") is obtained from the image generation unit 420 or the compression encoding unit 422 and associated with the partial image data.
[0042] The communication unit 426 is configured as the input / output interface 416 shown in Figure 3 and transmits a packet containing the compressed and encoded partial image data and its generation time to the image processing unit 200. With this configuration, the server 400 performs compression encoding, packetization, and transmission in parallel by pipeline processing in units of partial images smaller than one frame. The image processing unit 200 includes an image data acquisition unit 240, a decoding and decompression unit 242, an image processing unit 244, and a display control unit 246.
[0043] The decoding and decompression unit 242 and the image processing unit 244 share a common function in that they perform predetermined processing on partial image data to generate partial image data for display, and at least one of them can be collectively referred to as the "image processing unit." The image data acquisition unit 240 consists of the input / output interface 202, partial image storage unit 204, and control unit 206 shown in Figure 3, and acquires compressed and encoded partial image data from the server 400 along with its generation time.
[0044] The decoding / decompression unit 242 consists of the video decoder 208, partial image storage unit 210, control unit 206, and control unit 212 shown in Figure 3, and decodes and decompresses the compressed and encoded partial image data. The decoding / decompression unit 242 may start the decoding / decompression process once the image data acquisition unit 240 has acquired the data for the smallest unit area necessary for compression encoding, such as motion compensation and encoding. The image processing unit 244 consists of the image processing unit 214, partial image storage unit 216, control unit 212, and control unit 218 shown in Figure 3, and performs predetermined processing on the partial image data to generate partial image data for display. For example, as described above, the image processing unit 244 takes into account the distortion of the eyepiece lens of the head-mounted display 100 and applies a correction that gives the opposite distortion.
[0045] Alternatively, the image processing unit 244 synthesizes images to be displayed together with the moving image, such as the UI plane image, on a partial image basis. Alternatively, the image processing unit 244 obtains the position and orientation of the user's head at that moment and corrects the image generated by the server 400 so that it fits correctly into the field of view during display. This minimizes the time lag between the user's head movement and the displayed image caused by the transfer time from the server 400.
[0046] The image processing unit 244 may also perform any or a combination of commonly performed image processing. For example, the image processing unit 244 may perform Gamma curve correction, tone curve correction, contrast enhancement, etc. That is, based on the characteristics of the display device and user specifications, it may perform necessary offset corrections on the pixel values and luminance values of the decoded and decompressed image data. The image processing unit 244 may also perform noise reduction processing, such as superimposing, weighted averaging, and smoothing, by referring to neighboring pixels.
[0047] The image processing unit 244 may also match the resolution of the image data with the resolution of the display panel, or refer to neighboring pixels and perform weighted averaging, oversampling, bilinear or trilinear processing, etc. The image processing unit 244 may also refer to neighboring pixels to determine the type of image texture and selectively perform denoising, edge enhancement, smoothing, and tone / gamma / contrast correction accordingly. In this case, the image processing unit 244 may also process the image size in conjunction with an upscaler or downscaler.
[0048] The image processing unit 244 may also perform format conversion if the pixel format of the image data differs from the pixel format of the display panel. For example, it may perform conversions from YUV to RGB, from RGB to YUV, between 444, 422, and 420 in YUV, and between 8, 10, and 12-bit color in RGB. Furthermore, if the decoded image data is in an HDR (High Dynamic Range) brightness range compatible format, but the HDR brightness range compatible range of the display display is narrow (for example, the displayable brightness dynamic range is narrower than specified in the HDR format), the image processing unit 244 may perform pseudo-HDR processing (color space change) to convert the image to an HDR brightness range format compatible with the display panel while preserving the characteristics of the HDR image as much as possible.
[0049] Furthermore, if the decoded image data is in an HDR-compatible format but the display only supports SDR (Standard Dynamic Range), the image processing unit 244 may perform a color space conversion to the SDR format while preserving the characteristics of the HDR image as much as possible. If the decoded image data is in an SDR-compatible format but the display supports HDR, the image processing unit 244 may perform an enhancement conversion to the HDR format, matching the characteristics of the HDR panel as much as possible.
[0050] Furthermore, if the display has low grayscale expression capability, the image processing unit 244 may add error diffusion or perform dithering processing in conjunction with pixel format conversion. Also, if there are partial defects or abnormalities in the decoded image data due to loss or bit corruption of network transmission data, the image processing unit 244 may perform correction processing on those areas. In addition, the image processing unit 244 may perform correction by monochrome filling, correction by duplicating neighboring pixels, correction by neighboring pixels of the previous frame, or correction by using pixels estimated from past frames or the surrounding area of the current frame through adaptive defect correction.
[0051] Furthermore, the image processing unit 244 may perform image compression in order to reduce the required bandwidth of the interface output from the image processing unit 200 to the display device. In this case, the image processing unit 244 may perform lightweight entropy coding by neighboring pixel referencing, index value reference coding, Huffman coding, etc. Also, if the display device uses a liquid crystal panel, high resolution is possible, but the response speed is slow. If the display device uses an organic EL panel, the response speed is fast, but high resolution is difficult, and a phenomenon called black smearing, in which color bleeding occurs in and around black areas, may occur.
[0052] Therefore, the image processing unit 244 may perform corrections to eliminate the various adverse effects caused by such display panels. For example, in the case of an LCD panel, the image processing unit 244 resets the liquid crystal by inserting a black image between frames to improve the response speed. In the case of an organic EL panel, the image processing unit 244 applies an offset to the brightness value and the gamma value in gamma correction to make color bleeding due to black smearing less noticeable.
[0053] The display control unit 246 consists of the display controller 220 and control unit 218 shown in Figure 3, and sequentially displays partial image data for display on the display panel of the head-mounted display 100 or the flat panel display 302. However, in this embodiment, since compressed encoded data for partial images is acquired individually from the server 400, the acquisition order may be changed depending on the communication status, or partial image data itself may not be acquired due to packet loss.
[0054] The display control unit 246 then derives the elapsed time since the partial image was drawn from the generation time of each partial image, and adjusts the output timing of the partial image to the display panel to reproduce the drawing timing on the server 400. Specifically, the display control unit 246 includes a data acquisition status identification unit 248, an output target determination unit 250, and an output unit 252. The data acquisition status identification unit 248 identifies the data acquisition status, such as the original display order and display timing of the partial image data and the amount of missing data in the partial image, based on the generation time of the partial image data and / or the elapsed time since the generation time.
[0055] The output target determination unit 250 changes the output target to the display panel or appropriately adjusts the output order and timing depending on the data acquisition status. For example, depending on the data acquisition status, the output target determination unit 250 decides whether to output the data of the original partial image contained in the next frame or to output the data of the partial image contained in an earlier frame again. The output target determination unit 250 makes such an output target by the timing of the vertical synchronization signal, which is the display start time of the next frame.
[0056] For example, the output target determination unit 250 may change the output target according to the amount (percentage) of partial images acquired, such as replacing the output target with data from the previous frame if a predetermined percentage or more of partial images are missing in a frame. The output target determination unit 250 may also change the output target for the next frame display period according to past output performance of frames or the elapsed time since generation. The output unit 252 outputs the data of the partial images determined as output targets to the display panel in the order and timing determined by the output target determination unit 250.
[0057] Figure 6 is a diagram illustrating the effect of pipeline processing performed by the server 400 and the image processing device 200 on a partial image basis in this embodiment. The horizontal direction of the figure represents the passage of time, and each processing time is indicated by an arrow along with the processing name. Processing on the server 400 side is shown by a thin line, and processing on the image processing device 200 side is shown by a thick line. The notation in parentheses next to the processing name indicates that (m) is the processing of one frame with frame number m, and (m / n) is the processing of the nth partial image of frame number m.
[0058] Furthermore, the vertical synchronization signal on the server 400 side is denoted as vsync(server), and the vertical synchronization signals on the image processing device 200 and display device side are denoted as vsync(client). First, (a) shows a conventional configuration in which processing progresses one frame at a time for comparison. In this example, the server 400 controls the processing for each frame using the vertical synchronization signal. Therefore, the server 400 starts the compression encoding of the data for the first frame stored in the frame buffer in accordance with the vertical synchronization signal.
[0059] The server 400 then starts compressing and encoding the data for the second frame in accordance with the next vertical synchronization signal, and simultaneously packets the compressed and encoded data from the first frame into predetermined units and sends them out. The image processing device 200 performs decoding and decompression processing in the order in which the frames arrive. However, even after decoding and decompression is complete and the image is ready for display, the display is delayed until the timing of the next vertical synchronization signal arrives. As a result, in the example shown, a delay of more than two display cycles occurs between the completion of rendering one frame and the start of compression encoding processing by the server 400 and the start of display.
[0060] Depending on the communication time between the server 400 and the image processing device 200, and the difference in the timing of their vertical synchronization signals, further delays may occur. According to the embodiment shown in (b), the server 400 starts sending the data of the first partial image of the first frame when the compression encoding of that partial image is completed. While that data is being transmitted across the network 306, the transmission process progresses on a partial image basis, with the compression encoding and transmission of the second partial image, the compression encoding and transmission of the third partial image, and so on.
[0061] The image processing device 200 sequentially decodes and decompresses the acquired partial image data. As a result, the data for the first partial image reaches a displayable state much faster than in case (a). The data for the first partial image waits to be displayed until the timing of the next vertical synchronization signal. Subsequent partial images are output sequentially following the output of the first partial image. Since the display time for one frame is the same as in case (a), the display of the nth partial image will be completed by the time of the next vertical synchronization signal.
[0062] In this way, by proceeding in parallel from image data compression encoding to display in units smaller than one frame, the illustrated example can achieve display one frame earlier than in case (a). In the illustrated example, in case (b), the compression encoding by server 400 is also started in response to the vertical synchronization signal, but as mentioned above, the delay time can be further reduced by performing compression encoding without waiting for the vertical synchronization signal.
[0063] Furthermore, the display control unit 246 of the image processing device 200 may shift the timing of at least one of the vertical synchronization signal and the horizontal synchronization signal within a range permitted by the display device, depending on when the data of a partial image becomes ready to be displayed. For example, the display control unit 246 may shorten the waiting time from when the display panel is ready to output until it is actually output by changing the operating frequency of the pixel clock, which is the basis for all display timings, for a predetermined minute, or by changing the horizontal retrace period and the vertical retrace period for a predetermined minute. The display control unit 246 may repeat this change every frame so that a single large change does not cause the display to break down.
[0064] The example shown in Figure 6 illustrates an ideal case where partial image data sent independently from the server 400 arrives at the image processing device 200 without interruption and in approximately the same transmission time. On the other hand, by making the processing unit granularity finer, changes in the data acquisition status at the image processing device 200 are more likely to occur. Figure 7 illustrates the transmission status of partial image data between the server 400 and the image processing device 200.
[0065] The vertical axis of the diagram represents the passage of time, and the arrows indicate the time it takes for the data of the first to seventh partial images, shown on the axis of the server 400, to reach the image processing device 200. The communication status between the server 400 and the image processing device 200 can constantly change, and the data size of the compressed and encoded partial images can also change each time, so the transmission time of the data for each partial image will vary. Therefore, even if the server 400 sends the partial image data sequentially at roughly the same interval, that state may not be maintained when the image processing device 200 acquires it.
[0066] In the illustrated example, the data acquisition interval t1 for the first and second subimages is significantly different from the data acquisition interval t2 for the second and third subimages. Furthermore, the data acquisition order for the fourth and fifth subimages is reversed. In addition, data may not reach the image processing device 200 due to packet loss, as is the case with the data for the sixth subimage. In such a situation, if the image processing device 200 outputs the acquired subimage data to the display panel in the same order, it is possible that the original image will not be displayed or the display cycle will break down.
[0067] The data acquisition status identification unit 248 then refers to the generation time of each partial image on the server 400 side and grasps the various situations shown in Figure 7. Based on this, the output target determination unit 250 optimizes the output targets and the output timing of partial images during the display period of each frame. For this reason, the communication unit 426 of the server 400 may transmit to the image processing device 200 the data of the partial image and its generation time, along with a history of the generation times of a predetermined number of transmitted partial images.
[0068] For example, the communication unit 426 transmits the generation times of the 64 most recently transmitted partial images along with the data of the next partial image. The data acquisition status identification unit 248 of the image processing device 200 can identify data loss or reversal of the acquisition order by comparing the history of these generation times with the generation times of the actually acquired partial images. In other words, when data loss occurs, the data acquisition status identification unit 248 can obtain the generation time of that partial image from the history of generation times transmitted along with the subsequent partial images.
[0069] Figure 8 is a flowchart illustrating an example of a processing procedure in this embodiment in which the display control unit 246 outputs partial image data to the display panel while adjusting the output target and output timing. This flowchart shows the procedure for processing that should be performed for frames in which display should start in accordance with the vertical synchronization signal, at a predetermined timing prior to the timing of the vertical synchronization signal of the display panel. In other words, the illustrated processing is repeated for each frame.
[0070] First, the data acquisition status identification unit 248 of the display control unit 246 identifies the acquisition status of the partial image included in the target frame (S10). The data acquisition status identification unit 248 may also record the output record of the partial image in previous frames and refer to it. Here, the output record is, for example, at least one of the following types of data. 1. History of the classification selected in a specified period of time from the three classifications described below. 2. History, occurrence rate, and area percentage of missing partial images in the first category described below for a specified period in the past. 3. Time elapsed since the last display image was updated.
[0071] The output target determination unit 250 then determines which of the pre-prepared classifications the identified situations fall into (S12). The output target determination unit 250 basically makes a comprehensive determination from various perspectives and determines the output targets in order to provide the best user experience. For this reason, in S10, the data acquisition status identification unit 248 acquires at least one of the following parameters.
[0072] 1. The number of partial images that have been acquired from the partial images that make up the target frame. 2. Missing area of the partial image in the target frame. 3. Display duration of the same image frame 4. Duration of the blackout, as described below. 5. Elapsed time since the generation of the partial image Set one or more thresholds for each of the above parameters.
[0073] The output target determination unit 250 then classifies the situation by assigning a score to each parameter acquired for the target frame, depending on the range it falls into. For example, each parameter acquired in S10 is assigned a score based on a predetermined table, and the situation is classified based on the distribution of scores for all parameters. In the illustrated example, three classifications are prepared. If the data falls into the first classification, the output target determination unit 250 determines the most recent partial image data acquired so far as the output target and has the output unit 252 output it (S14).
[0074] For example, if the number of constituent partial images has been acquired to a predetermined value (predetermined percentage) or more, and the elapsed time from the generation time of the partial images is within an acceptable range, the output target determination unit 250 classifies the target frame as the first category based on the score determination. At this time, the output target determination unit 250 adjusts the timing so that the partial images are output in the order corresponding to each generation time. Ideally, as shown in Figure 6, the partial images are output sequentially from the upper part of the frame.
[0075] However, if there are missing parts of the image, you may further score each missing section using the parameters below to determine whether to reuse a partial image from the same position in the previous frame or to black out that section. 1. Display duration of the same image frame 2. Duration of the blackout, as described below. 3. Time elapsed since the generation of the partial image
[0076] Furthermore, if the situation in which the original image cannot be displayed has continued for a predetermined time or longer, the output target determination unit 250 may classify the target frame as first category even if the partial image of the target frame has not been acquired to a predetermined value (predetermined percentage) or more. The table for determining the score mentioned above may be set up in this manner. This makes it possible to represent the movement of the image to the extent possible, even if only partially. In addition, the image processing unit 244 may estimate and repair the missing images.
[0077] If the frame falls into the second category, the output target determination unit 250 determines that the image data of the frame preceding the target frame is the output target and causes the output unit 252 to output it (S16). In this case, the same frame will continue to be displayed on the display panel. For example, if the partial image of the target frame cannot be acquired by a predetermined value (predetermined percentage) or more, but the partial image of the frame within the previous predetermined time has been acquired by a predetermined value (predetermined percentage) or more, the output target determination unit 250 classifies the target frame into the second category. The table for determining the score mentioned above may be set up in this manner.
[0078] If the frame falls into category 3, the output target determination unit 250 decides not to output anything for the period during which the data of the target frame should be output (S18). In this case, a blackout period of one frame occurs on the display panel. For example, if a portion of the target frame cannot be acquired by a predetermined value (predetermined percentage) or more, and too much time has passed since the generation time to continue displaying the already displayed image, the output target determination unit 250 classifies the target frame into category 3. The table for determining the score mentioned above may be set in this manner. Note that the blackout basically involves displaying a black filled image, but another color set in advance may be used.
[0079] If the display in the second or third category continues for a predetermined time, the system can be branched to the first category regardless of the partial image acquisition status, as described above, and some image updates, even if only partial, may be performed. Note that the user experience tends to deteriorate as the system progresses from the first to the third category. Therefore, a table is established for each parameter acquired in S10 so that a high score is given when the user experience should be improved.
[0080] Then, in S12, the parameters obtained in S10 are assigned scores based on the table, and the multiple scores obtained are summed up. The display method classification from the first to the third classification is selected based on how large the sum is. Here, a threshold is predetermined where a larger sum corresponds to the first classification and a smaller sum corresponds to the third classification.
[0081] The illustrated example is only one embodiment, and the type of information acquired by the data acquisition status identification unit 248, the criteria for classification by the output target determination unit 250, and the output targets in each classification are determined appropriately based on the content of the moving image to be displayed, the acceptable degree and duration of loss, the acceptable display delay and display stop time, etc. The display control unit 246 may also store images accumulated while waiting for the next vertical synchronization signal, a certain range of displayed images, the generation time, and the judgment results and scores from S10 and S12 in a memory not shown.
[0082] Furthermore, the output target determination unit 250, separate from the determination in S12, determines whether the situation related to the target frame meets the conditions for warning the user (S20). For example, the output target determination unit 250 determines that a warning to the user is necessary if the amount of blackout time per unit time or the amount of partial image loss exceeds a threshold (Y in S20). At this time, the output target determination unit 250 displays a message indicating that the communication status is affecting the image display (S22).
[0083] The message may be displayed by the image processing unit 244 by superimposing it onto a portion of the image. This allows the user to understand the cause of the problem with the displayed image. If the warning conditions are not met, the message is not displayed (N in S20). Processing for the target frame is terminated by the above procedure, and processing for the next frame is started.
[0084] The data acquisition status identification unit 248 may also derive a trend in data transmission delay time based on the elapsed time from the generation time of the partial image acquired in S10 to the processing time. For example, the data acquisition status identification unit 248 generates a histogram of the elapsed time from the generation time of a predetermined number of partial images that have been acquired in the past. The data acquisition status identification unit 248 then detects an increasing trend in elapsed time when the histogram is biased in the direction of increasing elapsed time by more than a certain threshold value.
[0085] At this time, the data acquisition status identification unit 248 may request the server 400, via the image data acquisition unit 240 or the like, to reduce the size of the image data to be transmitted. For example, the data acquisition status identification unit 248 may request that the transmission of one frame's worth of image data be skipped, or that the compression ratio be increased by a predetermined amount. It may also request that the screen resolution be reduced by a predetermined amount. Alternatively, the data acquisition status identification unit 248 may request the output target determination unit 250 to skip the output of one frame's worth of image data.
[0086] Alternatively, the data acquisition status identification unit 248 may transmit to the server 400 the elapsed time from the generation time of the partial image to its acquisition. The communication unit 426 of the server 400 receives this data, and the compression encoding unit 422 generates a histogram. When the histogram is biased toward increasing elapsed time, the unit may detect the increasing trend in elapsed time and suppress the size of the image data to be transmitted. Alternatively, the data acquisition status identification unit 248 may notify the server 400 of the occurrence amounts of the first to third classifications described above. These measures can prevent the delay time from increasing, which could lead to significant delays in the display of subsequent frames or data loss.
[0087] In the pipeline processing for each partial image described above, in a system where the image processing device 200, which is the client, receives and displays image data generated by the server 400, the server 400 compresses and encodes the image data in units of partial images smaller than one frame and transmits it. The image processing device 200 also decodes and decompresses the data in units of partial images, performs the necessary processing, and outputs it sequentially to the display panel. This makes it possible for both the server 400 and the image processing device 200 to perform pipeline processing at a granularity finer than one frame.
[0088] As a result, the waiting time can be reduced compared to processing each frame individually, and the delay time from drawing to display can be reduced. Furthermore, by having the server 400 transmit the generation time along with the data of the partial image, the image processing device 200 can advance the display by reproducing the order in which the server 400 generated the images, even if independent data is transmitted one after another in short intervals.
[0089] Furthermore, because missing image segments can be identified based on the generation time, it is possible to appropriately select whether to display the latest data, reuse data from the previous frame, or output nothing, depending on the display content and display policy. This increases the range of possible countermeasures for various data transmission situations, enabling low-latency display while minimizing disruptions by taking measures to avoid degrading the user experience.
[0090] 3. Reprojection processing in the image processing unit Next, the image processing device 200, which has a reprojection function, will be described. Figure 9 shows the configuration of the image processing device 200 with a reprojection function and the functional block of the server 400 which has the corresponding function. Reprojection refers to the process of correcting an image that has been drawn to correspond to the position and posture of the head immediately before display, in an embodiment in which an image is displayed in a field of view that corresponds to the movement of the user's head using a head-mounted display 100.
[0091] For example, in this embodiment, when image data is transmitted from the server 400, even if the server acquires the position and orientation of the head-mounted display 100 and generates a corresponding image, changes in position and orientation during the data transmission period are not reflected in the image. Even if the image processing device 200 performs corrections according to the position and orientation of the head, if the corrections are performed on a frame-by-frame basis, changes in position and orientation that occur during the correction process are still not reflected in the image.
[0092] As a result, a significant delay occurs in the displayed image in response to head movements, diminishing immersion in virtual reality and potentially causing motion sickness, thus degrading the quality of the user experience. As mentioned above, by performing pipeline processing on a segment-by-segment basis and then applying corrections to reflect the head position and posture for each segment image just before display, a more responsive display can be achieved.
[0093] A rendering device and head-mounted display for performing reprojection on a frame-by-frame basis are disclosed, for example, in International Publication No. 2019 / 026765. In this technology, the rendering device predicts the position and orientation of the head-mounted display at the time of frame display and draws the image, and the head-mounted display further corrects the image based on the difference between the predicted values and the latest position and orientation before displaying it. However, if the rendering device is used as a server, the transmission path to the head-mounted display and consequently the transmission time will vary, and it is conceivable that a discrepancy that is difficult to correct may occur between the predicted position and orientation and the actual position and orientation at the time of display.
[0094] Therefore, in this embodiment, the time from image generation to display is shortened by pipeline processing on a partial image basis, and the display target is changed according to the degree of deviation, etc. In Figure 9, the server 400 includes an image generation unit 420, a compression encoding unit 422, a packetization unit 424, and a communication unit 426. The image processing device 200 includes an image data acquisition unit 240, a decoding and decompression unit 242, an image processing unit 244, and a display control unit 246.
[0095] The compression encoding unit 422, packetization unit 424, communication unit 426 of the server 400, and the image data acquisition unit 240 and decoding / decompression unit 242 have the same functions as described in Figure 5. However, this is not intended to limit their functions; for example, the image data acquisition unit 240 may acquire uncompressed partial image data from the server 400. In this case, the functions of the compression encoding unit 422 in the server 400 and the decoding / decompression unit 242 in the image processing device 200 can be omitted. On the other hand, when compression encoding processing is performed in the server 400 and decoding / decompression processing is performed in the image processing device 200, the time required from drawing to display becomes apparent, making it easier to demonstrate the effects of reprojection.
[0096] The server 400 may also implement multiplayer games or live streaming of sports competitions involving multiple players. In this case, the server 400 continuously acquires the head movements of each of the multiple users, generates images with the corresponding field of view, and streams them to each user's image processing device 200. In the case of a multiplayer game, the server 400 renders the virtual world from each player's viewpoint based on their footprint in three-dimensional space. In the case of a sports broadcast, the server 400 generates images corresponding to each user's viewpoint based on images of the competition captured by multiple cameras placed in a distributed manner.
[0097] The image generation unit 420 of the server 400 includes a position and orientation acquisition unit 284, a position and orientation prediction unit 286, and a drawing unit 288. The position and orientation acquisition unit 284 acquires information on the position and orientation of the user's head wearing the head-mounted display 100 from the image processing device 200 at a predetermined rate. The position and orientation prediction unit 286 predicts the user's position and orientation at the time the generated image frame is displayed. That is, the position and orientation prediction unit 286 determines the delay time from the generation of the image frame until it is displayed on the head-mounted display 100, and predicts how the position and orientation acquired by the position and orientation acquisition unit 284 will change after the delay time has elapsed.
[0098] The delay time is derived based on the processing performance of the server 400 and the image processing device 200, as well as the delay time in the transmission path. The change in position and orientation is then calculated by multiplying the translational speed and angular velocity of the head-mounted display 100 by the delay time, and this is added to the position and orientation acquired by the position and orientation acquisition unit 284. The drawing unit 288 sets the view screen based on the predicted position and orientation information and draws the image frame. In this case, the packetization unit 424 obtains the predicted values of the user's head position and orientation, which were used as assumptions when the image frame was drawn, from the image generation unit 420 along with the generation time of the image frame, and associates them with the partial image data.
[0099] As a result, the image data acquisition unit 240 of the image processing device 200 acquires the generation time and predicted values of the user's head position and orientation along with the data of the partial image. In case of partial image data being lost during transmission or the order in which the data arrives at the image processing device 200 being reversed, the communication unit 426 of the server 400 may transmit a predetermined number of recently transmitted predicted position and orientation values along with the data of the next partial image. When focusing on the reprojection process, the image processing device 200 may acquire the generation time for each frame from the server 400 instead of the generation time for each partial image. In the latter case, the delay time described below will be in frames.
[0100] The image processing unit 244 of the image processing device 200 includes a position and orientation tracking unit 260, a first correction unit 262, a synthesis unit 264, and a second correction unit 266. The position and orientation tracking unit 260 acquires images captured by at least one of the cameras of the head-mounted display, or measured values from a motion sensor built into the head-mounted display 100, and derives the position and orientation of the user's head at a predetermined rate.
[0101] As described above, any of the various methods that have been put into practical use can be used to derive the position and orientation of the head. Alternatively, the head-mounted display 100 may derive this information internally, and the position and orientation tracking unit 260 may simply acquire this information from the head-mounted display 100 at a predetermined rate. This information is transmitted to the server 400 via the image data acquisition unit 240. At this time, the time at which the captured image or motion sensor measurement values that formed the basis of the head position and orientation information of the target are obtained is associated with the transmission.
[0102] The first correction unit 262 applies reprojection processing to the partial image transmitted from the server 400 based on the difference between the head position and orientation most recently acquired by the position and orientation tracking unit 260 and their predicted values at the time the partial image was generated in the server 400. The difference used as the basis for reprojection may be at least one of the user's head position and the user's head orientation, but these are collectively referred to as "position and orientation".
[0103] The first correction unit 262 may precisely determine the latest position and orientation information at the time of correction by interpolating position and orientation information acquired at predetermined time intervals. The first correction unit 262 then performs the correction by deriving the difference between the position and orientation at the time of correction and its predicted value. More specifically, the first correction unit 262 creates a displacement vector map in which displacement vectors indicating where pixels in the image before correction are displaced by the correction are represented on the image plane.
[0104] Then, for each pixel constituting the partial image, the corrected partial image is generated by obtaining the pixel position of the displacement target by referring to the displacement vector map. At this time, the area of the corrected image that can be generated from the uncorrected partial image may change. The first correction unit 262 starts the correction process by referring to the displacement vector map when the data of the uncorrected partial image, which is necessary to generate the data of the corrected partial image, is stored in the local memory of the preceding stage. This makes it possible to process the corrected image on a partial image basis as well.
[0105] If the image transmitted from server 400 does not have distortion due to the eyepiece, the first correction unit 262 may also simultaneously perform a correction to introduce such distortion. In this case, the displacement vector represented for each pixel in the displacement vector map is a vector that combines the displacement vector for reprojection and the displacement vector for distortion correction. Of these, the displacement vector for distortion correction is eyepiece-specific data and is not dependent on user movement, etc., so it can be created in advance.
[0106] The first correction unit 262 updates the displacement vector map by combining the displacement vectors necessary for reprojection with the displacement vectors prepared in this manner for distortion correction, and then performs the correction. This allows for simultaneous reprojection and distortion correction for the eyepiece for each pixel of the partial image with a single displacement. Furthermore, the displacement vector map is updated for each region corresponding to the partial image, reflecting the previous head position and orientation, thereby enabling the display of an image with minimal delay from head movement throughout the entire frame.
[0107] The compositing unit 264 composites UI plane images onto partial images that have been corrected, such as through reprojection, on a partial image basis. However, the compositing unit 264 does not only composite UI plane images onto partial images, but can also composite any image, such as an image captured by the camera of the head-mounted display 100. In any case, when displaying on the head-mounted display 100, the images to be composited are also pre-distorted for the eyepiece.
[0108] The compositing process by the compositing unit 264 may be performed before the reprojection and distortion correction by the first correction unit 262. In this case, distortion should not be applied to the images to be composited, and distortion and other corrections can be applied to the composited image all at once. Also, if there are no images to be composited, the processing of the compositing unit 264 can be omitted. The second correction unit 266 performs the remaining correction processing of the corrections that should be applied to the display image. For example, when correcting chromatic aberration, the first correction unit 262 applies a common distortion corresponding to the eyepiece, regardless of the primary color of the display panel.
[0109] For example, considering the characteristics of the human eye when viewing a display panel, a correction is first made to the green color. Then, the second correction unit 266 generates a partial image of the red component by correcting only the difference between the red displacement vector and the green displacement vector. Similarly, a partial image of the blue component is generated by correcting only the difference between the blue displacement vector and the green displacement vector. For this reason, the second correction unit 266 prepares a displacement vector map in which the displacement vectors representing the differences for generating the red and blue images are represented on the image plane.
[0110] This allows for the generation of display-ready partial images with different corrections applied to each of the primary colors: red, green, and blue. The first correction unit 262, the synthesis unit 264, and the second correction unit 266 may perform their respective processes on a partial image basis to implement pipeline processing. In this case, the partial image storage unit and control unit shown in Figure 3 may be provided in each functional block. Furthermore, the content and order of the processes performed by the first correction unit 262, the synthesis unit 264, and the second correction unit 266 are not limited. For example, the first correction unit 262 may refer to a displacement vector map for each primary color and correct chromatic aberration simultaneously with other corrections.
[0111] The display control unit 246 basically has the same functions as shown in Figure 5, but in addition to the elapsed time since the partial image was drawn, it also changes the output target to the display panel based on the difference between the predicted position and orientation and the actual position and orientation. In other words, the data acquisition status identification unit 248a identifies the data acquisition status, such as the original display order and timing of the partial image data and the amount of data missing from the partial image, based on the elapsed time since the generation time and the difference between the predicted position and orientation and the actual position and orientation.
[0112] The output target determination unit 250b changes the output target to the display panel or appropriately adjusts the output order and timing according to these results. In this case, the output target determination unit 250b adds a criterion related to the user's head position and orientation, such as whether the display will not break down when reprojection is performed, to the various classifications shown in Figure 8. If it is determined that the display will not break down, the image processing unit 244 performs reprojection, and the output unit 252 outputs it to the display panel. In other words, in this case, after the output target is determined by the display control unit 246, various processes including reprojection are performed in the image processing unit 244. Alternatively, both are executed in parallel.
[0113] Figure 10 is a diagram illustrating the reprojection and distortion correction for the eyepiece performed by the first correction unit 262. As schematically shown in (a), the server 400 predicts the position and orientation of the user's 120 head and sets the view screen 122 to the corresponding position and orientation. The server 400 then projects objects within the viewing frustum 124 of the space to be displayed onto the view screen 122.
[0114] The server 400 then compresses and encodes each partial image as appropriate and transmits it to the image processing device 200. The image processing device 200 then decodes and decompresses the image as appropriate and outputs it sequentially to the display panel. However, if there is a large difference between the predicted position and orientation values on the server 400 and the actual position and orientation, the head movement and the display may not be synchronized, which could cause discomfort to the user or lead to motion sickness. Therefore, the first correction unit 262 shifts the image on the image by the amount of the difference so that the most recent head position and orientation is reflected in the display.
[0115] In the illustrated example, it is assumed that the user 120's head (face) has moved slightly downward and to the right compared to when the image was drawn. At this time, the first correction unit 262 sets a new view screen 126 to correspond to the latest position and orientation. View screen 126 is the original view screen 122 shifted downward and to the right. As shown in (b), when view screen 122 is shifted downward and to the right to become view screen 126, the image moves in the opposite direction, i.e., upward and to the left.
[0116] Therefore, the first correction unit 262 performs a correction that displaces the image by the amount of the view screen's displacement in the opposite direction to the view screen's displacement direction. Note that the view screen is not limited to two-dimensional translation; its orientation in three-dimensional space may be changed depending on the movement of the head. In this case, the amount of image displacement may change depending on the position in the image plane, but the displacement vector can be calculated using a general transformation formula used in computer graphics.
[0117] As described above, the first correction unit 262 may also simultaneously perform distortion correction for the eyepiece. That is, as shown in the lower part of (b), distortion is introduced into the original image so that the original image can be viewed without distortion when viewed through the eyepiece. A general formula for correcting lens distortion can be used for this process.
[0118] However, in this embodiment, as described above, the required correction amount and correction direction are calculated for each pixel and prepared as a displacement vector map. The first correction unit 262 generates a displacement vector map by combining the displacement vector for distortion correction with the displacement vector for reprojection obtained in real time, and by referring to this map, it realizes both corrections at once.
[0119] Figure 11 is a diagram illustrating an example of the correction process performed by the first correction unit 262. (a) shows the image plane before correction, and (b) shows the image plane after correction. S00, S01, S02... in the image plane before correction represent the positions where displacement vectors are set in the displacement vector map. For example, displacement vectors are set discretely in the horizontal and vertical directions of the image plane (for example, at equal intervals such as every 8 pixels or 16 pixels).
[0120] In the corrected image plane, D00, D01, D02, ... represent the displacement destinations of S00, S01, S02, ... respectively. In the figure, as an example, the displacement vector (Δx, Δy) from S00 to D00 is shown by a white arrow. The first correction unit 262 maps the image before correction to the corrected image in units of the smallest triangle with the pixels for which the displacement vector is set as vertices. For example, a triangle with vertices S00, S01, and S10 in the image before correction is mapped to a triangle with vertices D00, D01, and D10 in the corrected image.
[0121] Here, the pixels inside the triangle are shifted to positions interpolated linearly, bilinearly, or trilinearly, depending on their distance from D00, D01, and D10. The first correction unit 262 then determines the pixel values of the corrected image by reading the values of the corresponding pixels in the pre-correction partial image stored in the connected local memory. In this process, the pixel values of the corrected image are derived by interpolating the values of multiple pixels within a predetermined range from the reading location in the pre-correction image using bilinear, trilinear, or other methods.
[0122] This allows the first correction unit 262 to draw the corrected image in pixel sequence order, using units of triangles that are the destinations of the displacement of the triangles in the image before correction. As mentioned above, the first correction unit 262 may update the displacement vector map in units of regions corresponding to partial images in order to achieve correction that reflects the position and orientation of the head in real time. Similarly, the second correction unit 266 can refer to a different displacement vector map than the first correction unit 262 and map pixels to each smallest triangle. For example, when correcting chromatic aberration, images of each primary color component can be generated by using different displacement vector maps for each primary color.
[0123] Figure 12 is a flowchart showing the processing procedure in which the output target determination unit 250b of the display control unit 246 adjusts the output target when the image processing device 200 performs reprojection. This flowchart is performed after the determination process in S12 in the flowchart shown in Figure 8. More specifically, even if the frame was determined to fall under the first or second category in the determination process in S12, the output target determination unit 250b performs an additional determination to change it to the third category if necessary. Therefore, in S12 of Figure 8, if the target frame is determined to fall under the third category based on the acquisition status of the partial image, the output target determination unit 250b terminates the process (N in S90).
[0124] On the other hand, if it is determined that the user falls into either the first or second category (Y in S90), the output target determination unit 250b first determines whether the difference between the user's head position and orientation predicted by the server 400 when generating the image frame and the latest position and orientation is within an acceptable range (S92). As mentioned above, transmitting data from a cloud server or the like often takes more time than transmitting image data from a rendering device located near the head-mounted display 100. Therefore, the likelihood of a widening discrepancy between the predicted and actual position and orientation values increases.
[0125] Therefore, if the difference between the predicted position and orientation and the actual position is too large to be covered by reprojection, the data for that frame is not output to the display panel. In other words, if the difference exceeds the acceptable range, the output target determination unit 250b changes the classification of the target frame, which was classified as the first or second category, to the third category (N in S92, S98). In this case, the display blacks out. Alternatively, if the original was the first category, it may remain as the first category, and in S14 of Figure 8, past frames may not be used to display the missing portion.
[0126] The criteria for determining whether the difference in position and orientation is within an acceptable range may differ depending on whether the entire frame area can be covered by the most recent partial image (Category 1), or whether a partial image from a past frame is used for any part of it (Category 1 or Category 2). Specifically, the smaller the difference in position and orientation, the more permissible it is to use a partial image from a past frame. Whether or not it is within an acceptable range may be determined by its relationship to a threshold set for the difference in position and orientation, or a function may be used as the score value for the acceptable range, such that the score decreases as the difference in position and orientation increases, and this score may be added together with the score value used in the determination in Figure 8 to make a comprehensive determination.
[0127] If the difference in position and orientation is determined to be within an acceptable range (Y in S92), the output target determination unit 250b then determines whether the degree of data loss in the target frame, evaluated from the user's viewpoint, is within an acceptable range (S94). Specifically, the degree of data loss is quantified with a greater weight given to points closer to the user's point of focus, and if this value exceeds a threshold, it is determined that the degree of data loss is not within an acceptable range. If the degree of loss is not within an acceptable range (N in S94), the output target determination unit 250b changes the target frame, which was classified as the first category, to the second or third category.
[0128] Alternatively, the frame that was classified as Category 2 is changed to Category 3 (S98). This adjusts the output so that if there is a lot of data loss in areas that are easily visible to the user, the frame will not be output and a past frame will be reused or the frame will be blacked out. The determination process in S94 may be performed simultaneously with the determination process in S12 in Figure 8. If the degree of data loss is determined to be within an acceptable range (Y in S94), the output target determination unit 250b then determines whether there is enough data necessary for reprojection (S96).
[0129] In other words, when the field of view is shifted by reprojection, it is determined whether sufficient data of the partial image included in that field of view has been obtained. If sufficient data has not been obtained (N in S96), the output target determination unit 250b changes the target frame that was classified as the first category to the second or third category. Alternatively, it changes the target frame that was classified as the second category to the third category (S98). If sufficient data has been obtained, the output target determination unit 250b terminates processing with the original classification (Y in S96).
[0130] Note that S92, S94, and S96 are not limited to independent judgments; scores may be calculated based on each judgment criterion and then added together to determine comprehensively and simultaneously whether a change in classification or display content is necessary. As a result of these additional judgments, one of the processes S14, S16, or S18 in Figure 8 is performed. However, during the processing of S14 and S16, the image processing unit 244 performs correction processing including reprojection as described above. That is, in the case of the first classification, the image processing unit 244 performs reprojection on the latest image frame or on an image frame using past frames for missing parts.
[0131] In the case of the second classification, reprojection is performed on past frames that were selected for output. In these processes, the image processing unit 244 starts the correction process only after all the image data for the range to be used for reprojection of the partial image to be processed has arrived at the preceding partial image storage unit. When using past frames, the image processing unit 244 reads the relevant data from the memory (not shown) of the display control unit 246. The range of the image used for reprojection is determined by the difference between the predicted position and orientation and the actual position and orientation.
[0132] Figure 13 is a diagram illustrating a method for quantifying the degree of data loss based on the user's viewpoint in S94 of Figure 12. In this example, the user's gaze point 292 is assumed to be located near the center of the display screen 290. Since a user wearing a head-mounted display 100 usually turns their face in the direction they want to see, the center of the display screen 290 can also be considered as the gaze point 292.
[0133] In typical human vision, the area 294, corresponding within 5° of the line of sight from the pupil to the point of fixation, is called the discriminative field of vision and exhibits superior visual function such as acuity. Additionally, the area 296, corresponding within approximately 30° horizontally and 20° vertically, is called the effective field of vision, where information can be received instantaneously through eye movements alone. Furthermore, the area 298, corresponding within 60-90° horizontally and 45-70° vertically, is the stable fixation field of vision, and the area 299, corresponding within 100-200° horizontally and 85-130° vertically, is the auxiliary field of vision. Thus, the ability to discriminate information decreases as you move away from the point of fixation 292.
[0134] Therefore, as shown in the top and left of the figure, weighting functions 320a and 320b are set so that the closer the point of focus 292 is to the point of focus on the display screen 290, the greater the weighting. In this figure, the weighting functions 320a and 320b are shown for the one-dimensional horizontal and vertical positions on the plane of the display screen 290, but in reality, they are functions or tables for the two-dimensional position coordinates on the plane. The output target determination unit 250b derives the degree of loss as a numerical value by multiplying the missing area of the partial image by a weight based on the position coordinates where the loss occurs, and summing this over the entire area of the target frame.
[0135] This allows for a higher estimation of the degree of loss if a more easily visible area is missing, even if the missing area is the same, and enables a determination of whether it is within an acceptable range, taking into account the visual impression. Note that the shapes of the weighting functions 320a and 320b shown are merely examples, and their shapes may be optimized or discontinuous based on the visual characteristics of each range described above. Furthermore, if a gaze point detector is provided on the head-mounted display 100, the gaze point 292 can be determined more precisely, not limited to the center of the display screen 290. In this case, the output target determination unit 250b should move the position where the weighting functions 320a and 320b are maximized in accordance with the movement of the gaze point 292.
[0136] Figure 14 is a diagram illustrating the data required for reprojection, which is evaluated in S96 of Figure 12. First, (a) shows the view screen 340a set to correspond to the position and orientation predicted by the server 400. The server 400 draws the image 344 included in the viewing frustum 342a determined by the view screen 340a onto the view screen 340a. Here, let's assume that the position and orientation at the time of display was slightly to the left of the predicted position and orientation, as indicated by the arrow.
[0137] In this case, the image processing unit 244 of the image processing device 200 rotates the view screen 340b slightly to the left, as shown in (b), and corrects the image to produce a corresponding image. However, the larger the area 346 of the image included in the newly set viewing frustum 342b that has not been transmitted from the server 400, the more difficult reprojection becomes. Therefore, if the area 346 of the display area after reprojection where data has not been acquired exceeds a predetermined proportion, it is determined that the data is insufficient.
[0138] On the other hand, the image transmitted from the server 400 also contains a region 348 that is not included in the newly set viewing frustum 342b. In other words, the data in region 348 becomes an unnecessary region after reprojection, so its absence does not affect the display. Therefore, when the output target determination unit 250b evaluates the missing area, etc., in S12 of Figure 8, it may exclude the region 348 from the evaluation target. Even if sufficient image data necessary for reprojection is acquired, if the difference between the position and orientation predicted by the server 400 and the actual position and orientation is too large, the image after reprojection may appear unnatural.
[0139] In other words, reprojection is a two-dimensional correction of the image drawn by the server 400, and therefore cannot accurately represent changes in the line of sight to the three-dimensional object being displayed. For this reason, as shown in Figure 12, the output target determination unit 250b performs the determination in S92 separately from the determination in S96, and cancels output to the display panel if the difference in position and orientation is large. In any case, as shown in Figure 14, it is desirable to minimize the area 346 of the area required for reprojection where data has not been acquired, regardless of the movement of the user's head.
[0140] Therefore, the server 400 may increase the probability that reprojection will be performed successfully by speculatively generating an image of region 346 and sending it to the image processing device 200. Figure 15 is a flowchart showing the procedure of processing performed by the server 400 and the image processing device 200 when reprojection is performed in the image processing device 200. This flowchart is basically performed on a frame-by-frame basis of the moving image. First, the image data acquisition unit 240 of the image processing device 200 acquires the latest position and orientation of the user's head from the position and orientation tracking unit 260 and sends it to the server 400 (S100).
[0141] The image data acquisition unit 240 also acquires from the display control unit 246 the history of delay time from the generation time at the server 400 to processing at the image processing device 200, and the history of the difference between the predicted value and the actual value of the user's head position and orientation, for past frames, and transmits them to the server 400 (S102). The transmission processes in S100 and S102 may be performed at any timing that is not synchronized with the frame. In addition, in S102, the history for a predetermined number of past frames can be transmitted to prepare for transmission failures.
[0142] The position and orientation acquisition unit 284 of the server 400 receives this information, and the position and orientation prediction unit 286 predicts the position and orientation of the user's head (S104). That is, it predicts the position and orientation after a delay time until it is processed by the image processing device 200, using the position and orientation transmitted from the image processing device 200. Next, the drawing unit 288 draws an image corresponding to the predicted position and orientation (S106). At this time, the drawing unit 288 identifies how much the position and orientation may deviate in the most recently predicted delay time, based on the history of the delay time and the history of the position and orientation difference transmitted from the image processing device 200. In other words, as shown by the arrow in Figure 14(a), it predicts a vector representing the most likely amount of deviation and the direction of deviation for the predicted position and orientation value.
[0143] The drawing unit 288 then expands the drawing target in the direction in which the view screen has been displaced by the amount of the vector. That is, it determines an area outside the frame corresponding to the predicted position and orientation value, such as area 346 in Figure 14(b), based on the vector representing the predicted displacement, and additionally draws that image. The packetization unit 424 then packets the drawn image into partial image units as needed and transmits them sequentially from the communication unit 426 to the image processing device 200. At this time, the communication unit 426 transmits the generation time of the partial image and the predicted position and orientation value used for drawing, associating them with the partial image (S108).
[0144] The communication unit 426 further transmits the history of the generation time and the history of the predicted position and orientation values for a predetermined number of already transmitted partial images, in order to prepare for transmission failures to the image processing device 200. When the image data acquisition unit 240 of the image processing device 200 receives this data, the display control unit 246 acquires the delay time since the partial image was generated and the difference between the predicted position and orientation values and the actual values (S110). The display control unit 246 then controls the output target by classifying the frames based on this data (S112).
[0145] Except in cases where blackout is performed according to the third classification, the image processing unit 244 applies reprojection to the current frame or previous frame determined by the display control unit 246, based on the position and orientation at that time, in units of partial images, and the output unit 252 outputs to the display panel (S114). At this time, the image processing unit 244 uses images additionally drawn by the server 400, taking into account the deviation from the prediction, as needed.
[0146] As described above, the image processing unit 244 may perform corrections to remove lens distortion and chromatic aberration as appropriate, along with reprojection. When focusing on reprojection, each process can function effectively even if performed on a frame-by-frame basis rather than on a partial image basis. In this case, any additional image data generated by the server 400 based on deviations from the predicted position and orientation values may be transmitted from the server 400 to the image processing unit 200 on a partial image basis.
[0147] As described above, the reprojection process in the image processing unit allows the server to predict the user's head position and orientation at the time of display and generate an image within the corresponding field of view. The image processing unit then corrects the image transmitted from the server immediately before display to correspond to the head position and orientation at that moment. This improves the display's ability to track head movements, which is often a bottleneck when displaying images streamed from a server on a head-mounted display.
[0148] Furthermore, by using a pipeline processing method that performs the entire process from rendering on the server to displaying on the head-mounted display on a partial image basis, the delay time from rendering to display can be reduced, thus requiring only minimal correction for reprojection. As a result, accurate reprojection can be performed with simpler calculations compared to cases requiring significant correction.
[0149] Furthermore, in reprojection, by using a displacement vector map that represents the distribution of correction amount and direction on the image plane, independent correction becomes possible for each pixel. This enables high-precision correction even at the partial image level. In addition, by including the distortion correction component for the eyepiece in the displacement vector, unified correction becomes possible, minimizing the time required for reprojection.
[0150] Furthermore, as a preliminary step to reprojection, the image processing device controls the output target based on factors such as the difference between the position and orientation predicted by the server and the actual position and orientation, the degree of data loss evaluated from the user's perspective, and the acquisition rate of image data used for reprojection. For example, if the reprojection result using the latest data is predicted to be unsatisfactory, it may prevent the display or use data from past frames as the target for reprojection. The server also further predicts the deviation of the predicted position and orientation and speculatively generates images of the corresponding areas. Through these processes, the reprojection result can be made as good as possible, and high-quality images can be displayed with good responsiveness.
[0151] 4. Data formation process for different display formats As shown in Figure 1, the image display system 1 of this embodiment has the function of simultaneously displaying the same image on the head-mounted display 100 and the flat panel display 302. In the case of games such as virtual reality using the head-mounted display 100, the displayed content is not visible to anyone other than the user wearing the head-mounted display 100. Therefore, multiple users cannot watch the game progress together and share a sense of realism as a group.
[0152] Therefore, as mentioned above, it is conceivable to simultaneously display the content shown on the head-mounted display 100 on the flat-panel display 302 so that it can be shared by multiple users. Furthermore, when multiple users view the same content using the head-mounted display 100, it is necessary to generate and display images from different viewpoints based on the position and orientation of each user's head.
[0153] The head-mounted display 100 and the flat panel display 302 display images in significantly different formats. However, generating and transmitting data for multiple images in different formats from the server 400 can be inefficient in terms of transmission bandwidth and processing load. Therefore, by coordinating the server 400 and the image processing device 200, the server 400 and the image processing device 200 can appropriately select the division of data format conversion tasks and timing, enabling efficient display of multiple image formats.
[0154] Figure 16 shows the configuration of the functional blocks of a server 400 and an image processing device 200 that can support display on multiple display devices of different forms. The server 400 includes an image generation unit 420, a compression encoding unit 422, a packetization unit 424, and a communication unit 426. The image processing device 200 includes an image data acquisition unit 240, a decoding / decompression unit 242, an image processing unit 244, and a display control unit 246. The compression encoding unit 422, packetization unit 424, and communication unit 426 of the server 400, and the image data acquisition unit 240 and decoding / decompression unit 242 of the image processing device 200 have the same functions as described in Figure 5.
[0155] However, this is not intended to limit those functions; for example, the server 400 may transmit data of a partial image that has not been compressed or encoded. In this case, the functions of the compression encoding unit 422 of the server 400 and the decoding / decompression unit 242 of the image processing device 200 can be omitted. Furthermore, the image processing unit 244 of the image processing device 200 may have at least one of the various functions shown in Figure 9.
[0156] The image generation unit 420 of the server 400 includes a drawing unit 430, a formation content switching unit 432, and a data formation unit 434. The drawing unit 430 and the data formation unit 434 may be implemented by a combination of the image drawing unit 404 (GPU) and the drawing control unit 402 (CPU) and software shown in Figure 3. The formation content switching unit 432 may be implemented by a combination of the drawing control unit 402 (CPU) and software shown in Figure 3. The drawing unit 430 generates frames of moving images at a predetermined or variable rate. The images to be drawn here may be general images that do not depend on the form of the display device connected to the image processing device 200. Alternatively, the drawing unit 430 may sequentially acquire frames of captured moving images from a camera (not shown).
[0157] The forming content switching unit 432 switches the content of the forming process that should be performed by the server 400 according to a predetermined rule, among the forming processes necessary to create a format corresponding to the display mode of the display device connected to the image processing device 200. In its simplest form, the forming content switching unit 432 prepares images for the left eye and the right eye, and then decides whether or not to apply distortion corresponding to the eyepiece, depending on whether or not a head-mounted display 100 is connected to the image processing device 200.
[0158] If the image processing is performed on the server 400 side before the image is transmitted to the image processing device 200 to which the head-mounted display 100 is connected, at least part of the correction processing on the image processing device 200 can be omitted. In particular, when using the image processing device 200, which does not have abundant processing resources, being able to display the data transmitted from the server 400 almost as is is advantageous for low-latency display.
[0159] On the other hand, if a flat panel display 302 is also connected to the image processing device 200, and the server 400 transmits images for the left and right eyes that have been distorted for the eyepiece, it becomes necessary to crop one of the images or remove the distortion, which can actually increase the processing load. Transmitting multiple image formats from the server 400 increases the required transmission bandwidth.
[0160] Therefore, the formatting content switching unit 432 appropriately selects the processing to be performed on the server 400 side according to the situation, thereby enabling the transmission of the most efficient image format to the image processing device 200 without increasing the transmission bandwidth. It also enables the transmission of the image format that provides the best image quality and display delay depending on the application. Specifically, the formatting content switching unit 432 determines the processing content based on at least one of the following: the configuration of the display device connected to the image processing device 200, the processing performance of the correction unit that performs further formatting processing in the image processing device 200, the characteristics of the moving image, and the communication status with the image processing device 200.
[0161] Here, the display device configuration refers to, for example, resolution, supported frame rates, supported color spaces, optical parameters when viewed through a lens, and the number of display devices. Communication status refers to, for example, the communication bandwidth and communication delay currently being achieved. This information may be obtained from the image data acquisition unit 240 of the image processing device 200.
[0162] Based on this information, the image formation content switching unit 432 decides whether or not to generate images for the left and right eyes that match the display mode of the head-mounted display 100, whether or not to perform distortion correction for the eyepiece, and so on. However, the necessary processing is not limited to these, and various other processes such as resolution reduction and image deformation may be considered depending on the display mode that each user wants to achieve with their image processing device 200. Therefore, the candidate processing content to be determined by the image formation content switching unit 432 is also set appropriately accordingly.
[0163] The data formation unit 434 performs a portion of the formation process necessary to format each frame of the moving image into a format corresponding to the display mode realized by the image processing device 200, in accordance with the decision of the formation content switching unit 432. The communication unit 426 transmits the appropriately formed image data to the image processing device 200. When multiple image processing devices 200 are to be used as destinations, the data formation unit 434 performs a formation process appropriate for each image processing device 200.
[0164] For a single image processing device 200, even if the connected display device has multiple display modes, an increase in the required communication bandwidth is prevented by transmitting image data in one appropriately selected format. The communication unit 426 transmits the image data along with information indicating what kind of formation process was performed on the server 400 side. For multiple image processing devices 200 that realize different display modes, the image data in a format suitable for each device and the information indicating the content of the formation process are transmitted in association.
[0165] The image processing unit 244 of the image processing device 200 includes a first forming unit 270a (first correction unit) and a second forming unit 270b (second correction unit) as correction units, which convert the image transmitted from the server 400 into a format corresponding to the display mode to be implemented. In the illustrated example, the image processing device 200 is assumed to be connected to a head-mounted display 100 and a flat panel display 302, so two forming units, the first and second, are provided to correspond to them, but of course, as many forming units as there are display modes to be implemented may be provided.
[0166] In other words, the image processing unit 244 has the function of generating multiple frames of different formats from a single frame transmitted from the server 400, and the number of frames depends on the number of display modes to be implemented. The first forming unit 270a and the second forming unit 270b determine the processing necessary to create a format corresponding to each display mode, based on the content of the forming process performed on the server 400 side, which is transmitted from the server 400 along with the image data.
[0167] Depending on the content of the forming process performed on the server 400, the first forming unit 270a or the second forming unit 270b may omit its function. For example, if image data that can be directly displayed on the head-mounted display 100 is transmitted from the server 400, the first forming unit 270a corresponding to the head-mounted display 100 may omit its processing.
[0168] As another variation, either the first forming unit 270a or the second forming unit 270b may perform necessary forming processes on the image formed by the other. For example, the second forming unit 270b may perform forming processes on the image for the head-mounted display 100 formed by the first forming unit 270a to generate an image for the flat panel display 302.
[0169] Furthermore, various image processing operations shown in the image processing unit 244 of Figure 9 may be combined with forming operations that match the display format as appropriate. For example, a synthesis unit 264 may be incorporated immediately before the first forming unit 270a and the second forming unit 270b. Then, images captured by the camera of the head-mounted display 100 and UI plane images are synthesized with the frame transmitted from the server 400, and the first forming unit 270a produces an image in the display format of the head-mounted display 100, while the second forming unit 270b produces an image in the display format of the flat panel display 302.
[0170] The synthesis unit 264 may be incorporated immediately after the first forming unit 270a and the second forming unit 270b. In this case, the images to be synthesized must be in a format corresponding to the images formed by the first forming unit 270a and the second forming unit 270b, respectively. Alternatively, reprojection may be performed by including the function of the first correction unit 262 shown in Figure 9 in the first forming unit 270a.
[0171] In any case, the first forming unit 270a and the second forming unit 270b generate a displacement vector map as described above and perform corrections for each partial image based on it. When the first forming unit 270a performs reprojection or distortion correction, it may perform multiple corrections at once by generating a displacement vector map formed by synthesizing those displacement vectors.
[0172] The display control unit 246 includes a first control unit 272a and a second control unit 272b, which output images formed by the first forming unit 270a and the second forming unit 270b to the display panels of the head-mounted display 100 and the flat panel display 302, respectively. As described above, the transmission from the server 400 and the processing within the image processing device 200, including the processing of the first forming unit 270a and the second forming unit 270b, are carried out sequentially in units of partial images. For this reason, the first control unit 272a and the second control unit 272b each include a partial image storage unit for storing the data after the forming process.
[0173] Figure 17 illustrates the evolution of image formats that can be realized in this embodiment. As shown in Figure 16, the image processing device 200 is assumed to be connected to a head-mounted display 100 and a flat panel display 302. The figure shows four patterns (a), (b), (c), and (d), but the final format of the image to be displayed on the head-mounted display 100 and the flat panel display 302 does not depend on the pattern.
[0174] In other words, the image 132 to be displayed on the head-mounted display 100 consists of a left-eye image and a right-eye image, each having a format (first format) that is distorted for the eyepiece. The image 134 to be displayed on the planar display 302 consists of a single image common to both eyes and has a general image format (second format) that is free from lens distortion.
[0175] In contrast, (a) is a pattern in which the server 400 generates a pair of images 130a for the left eye and the right eye. In this case, the first forming unit 270a of the image processing device 200 applies distortion to the transmitted image for the eyepiece lens of the head-mounted display 100. As described above, this may be combined with the first correction unit 262, and further corrections such as reprojection may be performed. The second forming unit 270b crops either the left eye image or the right eye image and sets it to an appropriate image size.
[0176] (b) is a pattern in which the server 400 transmits an image 130b suitable for a flat panel display. In this case, the data formation unit 434 of the server 400 does not need to perform any formation processing on the image drawn by the drawing unit 430. The first formation unit 270a of the image processing device 200 generates images for the left eye and the right eye from the transmitted image and then applies distortion for the eyepiece. As described above, this may be combined with the first correction unit 262 and further corrections such as reprojection may be performed. In this case, the second formation unit 270b can omit its formation processing.
[0177] (c) is a pattern in which the server 400 generates images for the left and right eyes of the head-mounted display 100, and then generates an image 130c with distortion for the eyepiece. In this case, the first forming unit 270a of the image processing device 200 can omit the forming process. However, the first correction unit 262 may perform corrections such as reprojection. The second forming unit 270b cuts out either the left-eye image or the right-eye image, applies corrections to remove distortion, and then sets it to an appropriate image size.
[0178] (d) is a pattern in which the server 400 transmits the panoramic image 130d. Here, the panoramic image 130d is generally a 360° image of the entire sky represented by an equirectangular projection. However, the image format is not limited, and any of the various formats used for representing images with a fisheye lens, such as polyconic projection, equidistant projection, or others, may be used. Also, when using an image from a fisheye lens, the server 400 may generate a 360° panoramic image using images from two eyes.
[0179] However, the panoramic image is not limited to images captured by the camera; the server 400 may also render images viewed from multiple virtual starting points. In this case, the data formation unit 434 of the server 400 does not need to perform any formation processing on the panoramic image rendered by the rendering unit 430. The first formation unit 270a of the image processing device 200 then extracts the fields of view of the left and right eyes corresponding to the latest position and orientation of the user's head from the transmitted image and generates images for the left eye and the right eye.
[0180] Furthermore, the first forming unit 270a removes distortions from the camera lens and distortions from equirectangular projection, polyconic projection, equidistant projection, etc., in the transmitted image, and then applies distortion for the eyepiece of the head-mounted display 100. In the case of polyconic projection images, the gaps between cones are joined together, taking into account the distortion of the eyepiece. The image processing unit 244 may also perform a batch correction using a displacement vector map in combination with at least one of the various corrections described above.
[0181] The second forming unit 270b extracts the field of view for both eyes from the transmitted image and removes any distortion caused by the camera lens. In this case, the images displayed on the head-mounted display 100 and the flat-panel display 302 may be of different ranges as shown in the figure, or they may be of the same range. For example, the image displayed on the head-mounted display 100 may be of a range corresponding to the position and orientation of the user's head, while the image displayed on the flat-panel display 302 may be of a range separately specified by user operation via a game controller or the like.
[0182] Regarding the image 132 output to the head-mounted display 100, if reprojection is performed immediately before display, pattern (a) allows distortion correction and reprojection to be performed simultaneously in the image processing device 200. On the other hand, in (c), distortion correction is performed separately in the server 400 and reprojection is performed separately in the image processing device 200. Therefore, if the server 400 and the image processing device 200 have the same processing capabilities and perform the same displacement vector and pixel interpolation processing, the time required to output to the head-mounted display 100 is likely to be shorter in (a).
[0183] In pattern (d), the area corresponding to the most recent user's field of view is extracted immediately before display, so whether it is pre-generated panoramic video content or game images generated in real time, the processing equivalent to the aforementioned reprojection is not required. Also, since there is no restriction on the range of the transmitted image, a situation will not occur where there is insufficient data necessary to generate the display image by reprojection.
[0184] In pattern (a), the image 134 output to the flat panel display 302 can be generated by simply cropping a portion of the image transmitted from the server 400, thus enabling output at almost the same speed as output to the flat panel display 302 in pattern (b). Note that in patterns (b) and (d), the image transmitted from the server 400 does not contain parallax or depth information, so the image 132 output to the head-mounted display 100 is a parallax-free image.
[0185] In pattern (c), the image 134 output to the flat panel display 302 is output at a slower speed than in patterns (a) and (b) because the image processing device 200 processes the image that has been distorted in the server 400 back to its original state. Also, the process of restoring the distorted image back to its original state increases the likelihood of a decrease in image quality compared to patterns (a) and (b).
[0186] Pattern (d) allows the server 400 to transmit only one type of image data regardless of the viewing angle, thus streamlining processing. On the other hand, since it also transmits data for areas unnecessary for display, when transmitting data with the same bandwidth as (a), (b), and (c) and generating a display image from a single viewpoint, the image quality may be lower than that of those patterns. As each pattern has its own characteristics, the formation content switching unit 432 determines one of the patterns according to rules set based on these characteristics.
[0187] Figure 18 illustrates variations in the connection method of the display device on the user side (client side). First, (a) shows a configuration in which a head-mounted display 100 and a flat-panel display 302 are connected in parallel to a processor unit 140 that acts as a distributor. Note that the processor unit 140 may be built into an image processing device 200 that is not integrated with the head-mounted display 100 and the flat-panel display 302. (b) shows a configuration in which the head-mounted display 100 and the flat-panel display 302 are connected in series.
[0188] (c) shows a configuration in which the head-mounted display 100 and the flat panel display 302 are located in different places, and each acquires data from the server 400. In this case, the server 400 may send data in different formats to match the display format of the destination. However, in this configuration, since the head-mounted display 100 and the flat panel display 302 are not directly connected, it is difficult to reflect the images captured by the camera of the head-mounted display 100 on the flat panel display 302 when the images are combined and displayed.
[0189] In any case, the four transmission patterns described in Figure 17 can be applied to any of the connection systems shown in Figure 18. For example, if the image processing device 200 is integrally installed on both the head-mounted display 100 and the flat panel display 302, the first forming unit 270a and the second forming unit 270b determine whether or not to operate themselves, and if so, what processing content to perform, depending on the combination of the connection systems (a), (b), and (c) shown in Figure 18 and the format of the data transmitted from the server 400. Then, by operating the first control unit 272a and the second control unit 272b in accordance with the corresponding display mode, it becomes possible to display images that match the form of each display device, regardless of the configuration shown in Figure 18.
[0190] Figure 19 is a flowchart showing the processing procedure for sending image data in a format determined by the display mode of the server 400. This flowchart is initiated when the user selects a game to play or a video to watch from the image processing device 200. In response, the image data acquisition unit 240 of the image processing device 200 requests this from the server 400, and the server 400 establishes communication with the image processing device 200 (S30). Then, the format content switching unit 432 of the server 400 confirms the necessary information through a handshake with the image processing device 200 (Y in S31, S32). Specifically, at least one of the following items is confirmed.
[0191] 1. Among the transmission patterns shown in Figure 17, those that can be handled by the image processing device 200. 2. Whether or not to display an image on the flat panel display 302 from the same viewpoint as the head-mounted display 100. 3. Output formats available on the server side. 4. Required values for delay time, image quality, and presence or absence of left / right parallax. 5. Communication speed (communication bandwidth or transfer bandwidth, communication delay or transfer delay) 6. What can be processed by the image processing unit 244 of the image processing device 200 (processing capability) 7. Resolution, frame rate, color space, and eyepiece optical parameters of the head-mounted display 100 and the planar display 302.
[0192] Regarding the communication bandwidth and communication delay mentioned in item 5 above, actual values after system startup or actual values between the same server 400 / image processing device 200 may be referenced. The formation content switching unit 432 then determines the image formation content to be performed on the server 400 side according to a pre-set rule based on the confirmed information (S34). For example, among the patterns shown in Figure 17, the output delay time to the head-mounted display 100 is (c) ≈ (d) > (b) > (a), and the output delay time to the flat panel display 302 is (c) ≈ (d) > (a) > (b). Also, the image quality displayed on the flat panel display 302 is highest for (b) and lowest for (c).
[0193] Furthermore, when displaying to the head-mounted display 100, if the server 400 and the image processing unit 200 perform the same displacement vector and pixel interpolation processing, (a), which performs lens distortion and reprojection simultaneously, is superior to (c). However, the server 400 generally has more processing power than the image processing unit 200 and can perform image quality improvement processing such as high-density displacement vector maps and multi-tap pixel interpolation processing in a short time. Therefore, even in (c), where the image for the head-mounted display 100 is corrected twice, it may still be superior in terms of image quality. Note that a direct comparison of image quality between (a), (c) and (b) is not possible due to differences such as the presence or absence of parallax processing.
[0194] Rules are set up to allow for the selection of the optimal pattern, depending on the balance between latency and image quality, and the patterns that can be executed by the image processing unit 200 and the server 400. In this case, as mentioned above, communication bandwidth (transfer bandwidth) and communication delay (transfer delay) are also taken into consideration. Here, the total latency is the sum of the processing time at the server 400, the communication delay, and the processing time at the image processing unit 200. As a rule for pattern selection, an acceptable level for the total latency may be predetermined.
[0195] For example, if the communication delay (transfer delay) is greater than the relevant level, a pattern that shortens the total delay time is selected, even at the expense of image quality. In such cases, pattern (a) may be selected, assuming that the image processing device 200 performs lens distortion and reprojection processing simultaneously with standard image quality. If the communication delay (transfer delay) is smaller than that, a pattern with longer processing time on the server 400 and image processing device 200 may be selected to improve image quality.
[0196] For example, if the server 400 has high processing power, the pattern shown in (c) of Figure 17 may be selected instead of (a) for display on the head-mounted display 100. Also, for example, if the second forming unit 270b of the image processing unit 244 of the image processing device 200 does not support distortion correction, the pattern shown in (c) cannot be selected, and either the pattern shown in (a) or (b) must be selected.
[0197] The data formation unit 434 applies the determined content to the frame drawn by the drawing unit 430 (S36). If the data formation unit 434 decides to transmit the frame as drawn by the drawing unit 430, the formation process is omitted. Subsequently, the compression encoding unit 422 compresses and encodes the image data as needed, and the packetization unit 424 associates the image data with the content applied to that data to form a packet, which the communication unit 426 then transmits to the image processing device 200 (S38). In practice, the processes in S36 and S38 are carried out sequentially in units of partial images as described above.
[0198] Furthermore, if the communication bandwidth (transfer bandwidth) is insufficient to transfer the data formed according to the content determined in S34, the data size may be reduced by changing at least one of the quantization parameters, resolution, frame rate, color space, etc., during data compression using the control described later. In this case, the image processing unit 244 of the image processing device 200 may perform upscaling of the image size, etc. If there is no need to stop the transmission of the image due to user operation on the image processing device 200 (N in S40), the drawing, formation processing, and transmission are repeated for subsequent frames (N in S31, S36, S38).
[0199] However, when the image formation content is switched on the server 400 side, the formation content switching unit 432 updates the formation content through processing S32 and S34 (Y in S31). The timing for switching the image formation content can occur when the application being executed on the image processing device 200 is switched, when the mode is switched within the application, or at a timing specified by user operation on the image processing device 200.
[0200] Since the communication status between the server 400 and the image processing device 200 may change, the formation content switching unit 432 may also constantly monitor the communication speed and dynamically switch the formation content as needed. If it becomes necessary to stop transmitting images, the server 400 terminates all processing (Y in S40).
[0201] According to the data formation processes for different display formats described above, the server 400 performs a portion of the data formation process necessary for display, taking into account the display format to be realized in the image processing device 200, which is the destination of the image data. By determining the content of the formation process based on the type of display device connected to the image processing device 200, the processing performance of the image processing device, the communication status, and the content of the video to be displayed, the responsiveness and image quality of the display can be optimized under the given environment.
[0202] Furthermore, when displaying the same video on multiple display devices with different display modes, the server 400 can select a suitable data format, perform formatting processing, and then transmit the data, allowing data to be transmitted with the same communication bandwidth regardless of the number of display modes. This enables the image processing device 200 to format data according to each display mode with minimal processing. It also facilitates the combination of various corrections such as reprojection and image synthesis with other images. As a result, images provided over the network can be displayed with low latency and high image quality, regardless of the type or number of display modes to be implemented.
[0203] 5. Compression encoding of multiple images corresponding to each frame. This section describes compression encoding techniques for video data where each frame is composed of multiple images. First, consider the case where each frame of a video is composed of multiple images obtained from different viewpoints. For example, if display images are generated from viewpoints corresponding to a person's left and right eyes, and these are displayed in the left and right eye regions of a head-mounted display 100, the user can enjoy an immersive visual world. As an example, an event such as a sports competition can be filmed with multiple cameras placed in space, and this footage is used to generate and display left and right eye images corresponding to the head movement of the head-mounted display 100. This allows the user to view the event from various viewpoints, giving them the feeling of being at the venue.
[0204] In this technology, the images for the left eye and right eye to be transmitted from the server 400 are highly similar because they essentially represent images in the same space. Therefore, the compression encoding unit 422 of the server 400 compresses one image and uses the other data as information representing the difference between the two, thereby achieving a higher compression ratio. Alternatively, the compression encoding unit 422 may obtain information representing the difference with the compressed encoding result of the corresponding viewpoint from past frames. In other words, the compression encoding unit 422 performs predictive encoding by referring to the compressed encoding result of at least one of the following: another image representing the image at the same time from the video data, or an image from a past frame.
[0205] Figure 20 shows the functional block configuration of a server 400 having the function of compressing images from multiple viewpoints at high efficiency, and an image processing device 200 that processes them. The server 400 includes an image generation unit 420, a compression encoding unit 422, a packetization unit 424, and a communication unit 426. The image processing device 200 includes an image data acquisition unit 240, a decoding and decompression unit 242, an image processing unit 244, and a display control unit 246. The image generation unit 420, packetization unit 424, and communication unit 426 of the server 400, and the image data acquisition unit 240, image processing unit 244, and display control unit 246 of the image processing device 200 have the same functions as described in Figure 5 or Figure 16.
[0206] However, the image processing unit 244 of the image processing device 200 may have the various functions shown in Figure 9. Also, the image generation unit 420 of the server 400 acquires data of multiple moving images obtained from different viewpoints. For example, the image generation unit 420 may acquire images taken by multiple cameras placed at different locations, or it may render images viewed from multiple virtual viewpoints. In other words, the image generation unit 420 is not limited to generating images itself, but may also acquire image data from external cameras or the like.
[0207] The image generation unit 420 may also generate images for the left eye and the right eye based on the acquired images from multiple viewpoints, corresponding to the field of view of the user wearing the head-mounted display 100. Alternatively, images from multiple viewpoints, not limited to just the left eye and right eye, may be sent to the server 400, and the image processing device 200 at the receiving end may generate images for the left eye and right eye. Alternatively, the image processing device 200 may perform some image analysis using the transmitted images from multiple viewpoints.
[0208] In any case, the image generation unit 420 acquires or generates multiple frames of motion images, each containing at least a portion of the object being represented, at a predetermined or variable rate, and sequentially supplies them to the compression encoding unit 422. The compression encoding unit 422 comprises a division unit 440, a first encoding unit 442, and a second encoding unit 444. The division unit 440 forms image blocks by dividing corresponding frames of multiple motion images along a common boundary of the image plane. The first encoding unit 442 compresses and encodes a frame from one viewpoint among multiple frames with different viewpoints, in units of image blocks.
[0209] The second encoding unit 444 uses the data compressed and encoded by the first encoding unit 442 to compress and encode frames from other viewpoints in image block units. The encoding method is not particularly limited, but it is desirable to adopt a method that minimizes image quality degradation even when compressing individual partial regions independently without using the entire frame. This allows for sequential compression and encoding of frames from multiple viewpoints in units smaller than one frame, and enables transmission of each partial image to the image processing device 200 with minimal delay time using the same pipeline processing as described above.
[0210] Therefore, the boundaries of the image blocks are appropriately set according to the order in which the image generation unit 420 acquires data in the image plane. That is, the boundaries of the image blocks are determined so that areas acquired earlier are compressed, encoded, and transmitted earlier. For example, when the image generation unit 420 acquires the data of the pixel columns of each row in the image plane from top to bottom, the division unit 440 sets boundary lines in the horizontal direction to form an image block consisting of a predetermined number of rows.
[0211] However, the order of data acquired by the image generation unit 420 and the direction of image division are not particularly limited. For example, if the image generation unit 420 sequentially acquires data of vertical pixel rows, the division unit 440 sets a boundary line in the vertical direction. The boundary line is not limited to one direction; it may be divided in both vertical and horizontal directions to create tile-like image blocks. Alternatively, the image generation unit 420 may acquire data using an interlaced method, acquiring data of every other row of pixel rows in one scan and acquiring data of the remaining rows in the next scan, or it may acquire data in a meandering manner across the image plane. In the latter case, the division unit 440 sets a division boundary line in the diagonal direction of the image plane.
[0212] In any case, as described above, the division pattern corresponding to the data acquisition order is appropriately set so that the data of the region acquired first by the image generation unit 420 is compressed and encoded first, and the division unit 440 divides the image plane in a pattern corresponding to the data acquisition order of the image generation unit 420. Here, the division unit 440 uses the smallest division unit as the smallest region necessary for motion compensation and encoding. The first encoding unit 442 and the second encoding unit 444 perform compression encoding in units of image blocks thus divided by the division unit 440.
[0213] Here, as described above, the second encoding unit 444 uses the image data compressed and encoded by the first encoding unit 442 to increase the compression ratio of image data from other viewpoints. Furthermore, the first encoding unit 442 and the second encoding unit 444 may also use the compressed and encoded data of the image block at the same position in a previous frame to compress and encode each image block of the current frame.
[0214] Figure 21 illustrates the relationship between the image blocks compressed and encoded by the first encoding unit 442 and the second encoding unit 444, and the image blocks that reference the compressed encoding results. The figure shows the image plane when processing images for the left eye and right eye that have been distorted for the eyepiece, with N being a natural number, the left side showing the Nth frame and the right side showing the N+1th frame. For example, as shown in the Nth frame, the first encoding unit 442 compresses and encodes the left eye image 350a in units of image blocks (1-1), (1-2), ..., (1-n), and the second encoding unit 444 compresses and encodes the right eye image 350b in units of image blocks (2-1), (2-2), ..., (2-n).
[0215] Here, the second encoding unit 444 uses the data of the image block for the left eye image (e.g., left eye image 350a) that the first encoding unit 442 has compressed and encoded as a reference, as shown by the solid arrow (e.g., solid arrow 352), and compresses and encodes each image block for the right eye image (e.g., right eye image 350b). Alternatively, the second encoding unit 444 may use the compressed and encoded result of the image block in the right eye image 350b of the Nth frame as a reference, as shown by the dashed-dotted arrow (e.g., dashed-dotted arrow 354), and compress and encode the image block at the same position in the right eye image of the N+1th frame.
[0216] Alternatively, the second encoding unit 444 may compress and encode the target image block using both the references indicated by the solid arrow and the dashed arrow simultaneously. Similarly, the first encoding unit 442 may use the compressed encoding result of the image block in the left-eye image 350a of the Nth frame as a reference, as shown by the dashed arrow (e.g., dashed arrow 356), to compress and encode the image block at the same position in the left-eye image of the N+1th frame. In these embodiments, for example, the MVC (Multiview Video Coding) algorithm, which is an encoding method for multiview video, can be used.
[0217] MVC is based on the principle of using decoded images from multiple viewpoints, compressed and encoded using methods such as AVC (H264, MPEG-4), to predict images from other viewpoints, and then using the difference between these predicted images and the actual images as the target for compression and encoding of the other viewpoint. However, the encoding format is not particularly limited; for example, MV-HEVC (Multi-View Video Coding Extension), which is an extension of HEVC (High Efficiency Video Coding) for multi-view encoding, may also be used. In addition, other encoding methods such as VP9, AV1 (AOMedia Video 1), and VVC (Versatile Video Coding) may also be used as the base encoding method.
[0218] In any case, in this embodiment, the compression encoding process is performed on an image block basis to shorten the delay time until transmission. For example, while the first encoding unit is compressing and encoding the nth (n is a natural number) image block in a frame from one viewpoint, the second encoding unit compresses and encodes the (n-1)th image block in frames from other viewpoints. This allows the frame data to be acquired sequentially and in parallel, and the pipeline operation to be performed simultaneously, thereby speeding up the compression process.
[0219] The first encoding unit 442 and the second encoding unit 444 may each target one or more viewpoints for compression encoding. The frames of the video images with multiple viewpoints, compressed and encoded by the first encoding unit 442 and the second encoding unit 444, are sequentially supplied to the packetization unit 424 and transmitted from the communication unit 426 to the image processing device 200. At this time, each data is associated with information that distinguishes whether the data was compressed by the first encoding unit 442 or the second encoding unit 444.
[0220] The image block, which is the compression unit, may be the same as the sub-image, which is the unit of pipeline processing described above, or a single sub-image may contain data from multiple image blocks. The decoding and decompression unit 242 of the image processing device 200 comprises a first decoding unit 280 and a second decoding unit 282. The first decoding unit 280 decodes and decompresses the viewpoint image compressed and encoded by the first encoding unit 442. The second decoding unit 282 decodes and decompresses the viewpoint image compressed and encoded by the second encoding unit 444.
[0221] In other words, the first decoding unit 280 uses only the data compressed and encoded by the first encoding unit 442 to decode frames of the video for some viewpoints using a general decoding process according to the encoding scheme. The second decoding unit 282 uses the data compressed and encoded by the first encoding unit 442 and the data compressed and encoded by the second encoding unit 444 to decode frames of the video for the remaining viewpoints. For example, the former predicts the image of the target viewpoint, and the decoding result of the data compressed and encoded by the second encoding unit 444 is added to decode the image of that viewpoint.
[0222] Although the first decoding unit 280 and the second decoding unit 282 are shown separately as functional blocks, it will be understood by those skilled in the art that in reality, a single circuit or software module can sequentially perform two types of processing. As described above, the decoding and decompression unit 242 acquires compressed and encoded partial image data from the image data acquisition unit 240, decodes and decompresses it in units, and then supplies it to the image processing unit 244. By performing predictive coding using the similarity of images from multiple viewpoints, the data size to be transmitted from the server 400 can be reduced, enabling low-latency image display without straining the communication bandwidth.
[0223] Figure 22 is a diagram illustrating the effect of the server 400 in this embodiment on compressing and encoding images in block units by utilizing the similarity of images from multiple viewpoints. The upper part of the figure illustrates the division of images when processing images for the left eye and right eye that have been distorted for the eyepiece. The horizontal direction in the lower part of the figure represents the passage of time, and the time for compression and encoding processing for each region is indicated by arrows along with the numbers of each region shown in the upper part.
[0224] The procedure in (a) shows, for comparison, a case where the entire left-eye image from (1) is first compressed and encoded, and the difference between the predicted image and the actual image of the right eye using the result is used as the target for compression of the right-eye image in (2). In this procedure, when compressing and encoding the right-eye image in (2), it is necessary to first wait for the compression and encoding of the entire left-eye image to be completed, and then to perform prediction using that data or to calculate the difference with the actual image. As a result, as shown in the lower panel, it takes a relatively long time to complete the compression and encoding of both images.
[0225] Procedure (b) shows, for comparison, a case where the left-eye image and the right-eye image are treated as a single image and compressed and encoded in units of image blocks (1), (2), ..., (n) which are divided horizontally. As shown in the lower section, this procedure does not involve processes such as prediction or difference calculation from the compressed and encoded data, so the time required for compression and encoding can be shortened compared to procedure (a). However, since predictive encoding is not used, the compression ratio is lower than in (a).
[0226] The procedure in (c) is as described above in this embodiment, in which the first encoding unit 442 compresses and encodes the left eye image in units of image blocks (1-1), (1-2), ..., (1-n), and the second encoding unit 444 compresses the difference between the predicted image and the actual image of the right eye image using each of the data, thereby obtaining compressed encoded data of the right eye image in units of image blocks (2-1), (2-2), ..., (2-n). In this procedure, as shown in the lower section (c), the compression encoding of the image blocks (2-1), (2-2), ..., (2-n) of the right eye image can be started immediately after the completion of the compression encoding of the corresponding image blocks (1-1), (1-2), ..., (1-n) of the left eye image.
[0227] Furthermore, because the area processed at once during compression encoding of the left-eye image is small, and the area of the left-eye image that needs to be referenced during compression encoding of the right-eye image is small, the compression encoding time for each image block is shorter than the compression encoding time per image block calculated in (a). As a result, a higher compression ratio can be achieved in a time as short as that of procedure (b). Moreover, as described above, since the compressed image blocks are sequentially packetized and transmitted to the image processing device 200, the delay until display can be significantly reduced compared to procedure (a).
[0228] In procedure (c), the first encoding unit 442 and the second encoding unit 444 are shown separately as functional blocks, but in reality, a single circuit or software module can perform two types of processing sequentially. On the other hand, (c)' shown in the lower part of the figure shows a configuration in which the first encoding unit 442 and the second encoding unit 444 can perform compression encoding in parallel. In this case, during each period in which the first encoding unit 442 is compressing and encoding image blocks (1-2), ..., (1-n) in the left eye image, the second encoding unit 444 compresses and encodes the previous image block in the right eye image, i.e., image blocks (2-1), ..., (2-(n-1))th image block.
[0229] This further reduces the time required for compression encoding compared to case (c). Procedure (c)'' is a variation of procedure (c)', and shows a case where, after a predetermined unit of region in the image block of the left-eye image has been compressed and encoded, compression encoding of the corresponding region of the right-eye image is started using that data. Since the image of the same object appears in the left-eye and right-eye images at positions shifted by the amount of the disparity, in principle, if there is a reference image in the range of a few pixels to a few tens of pixels that takes this shift into account, predictive encoding of the other image is possible.
[0230] As mentioned above, the smallest unit area required for motion compensation and encoding in compression encoding is an area of a predetermined number of rows, such as one or two rows, or a rectangular area of a predetermined size, such as 16x16 pixels or 64x64 pixels. Therefore, the second encoding unit 444 can start compression encoding when the compression encoding of the left-eye image of the area required as a reference is completed in the first encoding unit 442, and its smallest unit is the unit area required for motion compensation and encoding.
[0231] For example, in the case where the first encoding unit 442 compresses and encodes the image blocks of the left eye image sequentially from one end to the other, the second encoding unit 444 can start compressing and encoding the corresponding image blocks of the right eye image before the compression and encoding of the entire image block is completed. In this way, the time required for compression and encoding can be further reduced compared to case (c)'.
[0232] Furthermore, the compression encoding unit 422 as a whole can shorten the time required for compression encoding by starting the compression encoding process as soon as the data for the smallest unit area required for compression encoding is prepared. Up to this point, we have mainly described the compression encoding of images with multiple viewpoints, but even if the parameters other than the viewpoint differ for multiple images corresponding to each frame of a video, low-latency and highly efficient compression encoding and transmission can be achieved with the same processing. For example, in scalable video encoding technology, when compressing a single video, the data in each layer is highly similar by creating layers with different resolutions, image quality, and frame rates to generate redundant data.
[0233] Therefore, by compressing and encoding the image in the base layer, and then using data from other layers as information representing the difference between that and the compressed image, a higher compression ratio is achieved. In other words, predictive encoding is performed by referring to the compressed encoding result of at least one of the following: another image representing the same image at the same time from the video data, or an image from a past frame. Here, the other image representing the same image at the same time from the video data has different resolutions and image quality defined hierarchically.
[0234] For example, an image representing the same picture may have different resolutions, such as 4K (3840 x 2160 pixels) and HD (1920 x 1080 pixels), and image quality levels, such as level 0 and level 1 in quantization parameters (QP). Here, past frames refer to the immediately preceding frame at different frame rates defined hierarchically. For example, in a system with frame rates of 60fps and 30fps, the immediately preceding frame in the 30fps layer refers to the frame from 1 / 30th of a second ago, which corresponds to two frames ago in the 60fps layer. Even in scalable video coding technology with such characteristics, compression coding at each layer may be performed on an image block basis.
[0235] As for scalable video coding, for example, SVC (Scalable Video Coding), an extension of AVC (H264, MPEG-4), and SHVC (Scalable High Efficiency Video Coding), an extension of HEVC, are known (see, for example, Takahiro Kimoto, "Standardization Trends of Scalable Video Coding in MPEG," Information Processing Society of Japan Research Report, 2005, AVM-48(10), pp. 55-60), and either of these may be adopted in this embodiment. In addition, the base coding scheme may also be VP9, AV1, VVC, etc.
[0236] Figure 23 shows the configuration of the functional blocks of the compression encoding unit 422 when scalable video encoding is performed. The other functional blocks are the same as those shown in Figure 20. In this case, the compression encoding unit 422 includes a resolution conversion unit 360, a communication status acquisition unit 452, and a transmission target adjustment unit 362, in addition to the division unit 440, first encoding unit 442, and second encoding unit 444 shown in Figure 20. The resolution conversion unit 360 gradually reduces each frame of the video generated or acquired by the image generation unit 420 to create images of multiple resolutions.
[0237] The division unit 440 divides the images of multiple resolutions along a common boundary of the image plane to form image blocks. The division rule may be the same as described above. The first encoding unit 442 compresses and encodes the image with the lowest resolution in image block units. As described above, this compression encoding is performed by referring to the image in question, or further, to the image block at the same position in a past frame of the same resolution.
[0238] The second encoding unit 444 uses the result of the compression encoding performed by the first encoding unit 442 to compress and encode the high-resolution image in image block units. In the figure, two encoding units, the first encoding unit 442 and the second encoding unit 444, are shown, assuming a case with two resolution levels. However, if there are three or more resolution levels, an encoding unit is provided for each level, and predictive encoding is performed using the compression encoding result of an image with a lower resolution.
[0239] In this case, it is also possible to refer to the image block at the same position in a past frame of the same resolution. Furthermore, in the encoding section, the frame rate may be adjusted or the quantization parameters used may be adjusted as needed to create differences in frame rate and image quality between layers.
[0240] For example, if a frame rate hierarchy is established, the frame rate used for processing in the compression encoding unit 422 will be the rate of the highest frame rate hierarchy. When compressing and encoding using past frames as references, the referenced frames will change depending on the frame rate. Furthermore, if a hierarchical image quality hierarchy is established, the first encoding unit 442 selects QP level 0, which is the base image quality, using quantization parameters, and the second encoding unit 444 and the nth encoding unit above it select a higher quality QP level to achieve this.
[0241] For example, if there is a base layer with HD resolution and QP level 0 image quality, and another layer with 4K resolution and QP level 1 image quality, the first encoding unit 442 processes the former, and the second encoding unit 444 processes the latter. Furthermore, if there is another layer with 4K resolution and QP level 2 image quality, the processing of that layer is handled by a third encoding unit (not shown). In other words, if the QP level is different even with the same resolution, an encoding unit is provided for that layer. Note that the combinations of resolution, image quality, and frame rate layers are not limited to these.
[0242] In any case, in this case as well, while the first encoding unit 442 is compressing and encoding the nth image block (where n is a natural number) in the lowest resolution (lowest layer) image, the second encoding unit 444 compresses and encodes the (n-1)th image block in the image of the layer above it. If there are three or more layers, each encoding unit performs the compression and encoding of the (n-2)th image block in the third layer, the (n-3)th image block in the fourth layer, and so on, in parallel, with each layer shifting one position forward as the layer increases. Examples of the parameters of the scalable video encoding realized in this embodiment are shown below.
[0243] [Table 1]
[0244] The communication status acquisition unit 452 acquires the communication status with the image processing device 200 at a predetermined rate. For example, the communication status acquisition unit 452 acquires the delay time from when the image data is transmitted until it reaches the image processing device 200 based on the response signal from the image processing device 200, or based on the aggregated information of the image processing device 200. Alternatively, the communication status acquisition unit 452 acquires from the image processing device 200 the percentage of transmitted image data packets that have reached the image processing device 200 as the data arrival rate. Furthermore, based on this information, it acquires the amount of data that could be transmitted per unit time, i.e., the available bandwidth. The communication status acquisition unit 452 monitors the communication status by acquiring at least one of these pieces of information at a predetermined rate.
[0245] The transmission target adjustment unit 362 determines, depending on the communication status, which data from the hierarchical data compressed and encoded by scalable video coding should be transmitted to the image processing device 200. Here, the transmission target adjustment unit 362 prioritizes the image data with the lowest resolution compressed and encoded by the first encoding unit 442 among the hierarchical data. Furthermore, the transmission target adjustment unit 362 determines, depending on the communication status, whether or not to expand the transmission target in the direction of higher hierarchical levels. For example, multiple thresholds are set for the available transfer bandwidth, and each time these thresholds are exceeded, the transmission target is gradually expanded to image data of higher hierarchical levels.
[0246] Conversely, if the available transfer bandwidth does not exceed the minimum threshold, only the image data with the lowest resolution will be transmitted. Alternatively, priority control may be applied, and the server 400 may transmit data from all layers regardless of the communication status. In this case, the transmission target adjustment unit 362 adjusts the transfer priority assigned to each layer according to the communication status. For example, the transmission target adjustment unit 362 may lower the transfer priority of higher-level data as needed. Here, the transfer priority may be multi-level.
[0247] At this time, the packetization unit 424 packets the information indicating the transfer priority that has been assigned along with the image data, and transmits it from the communication unit 426. However, if a router or other device on a node of the network 306 experiences a transfer that exceeds its processing capacity, it may discard packets as necessary. In streaming data transfer, a transfer protocol that omits the handshake, such as UDP, is generally used to achieve low latency and high transfer efficiency. In this case, the transmitting server 400 does not know whether a packet was discarded along the path of the network 306. At this time, the router or other device discards packets in order from the highest layer based on the assigned transfer priority. When server 400 simultaneously transmits layered data to multiple image processing devices 200, the processing capacity of the decoding and decompression unit 242 of each image processing device 200 and the communication performance of the network to each image processing device 200 vary.
[0248] As described above, the transmission target adjustment unit 362 may individually determine the transmission target according to the status of each image processing device 200, but in that case, it would be necessary to adjust multiple communication sessions separately, resulting in a very high processing load. On the other hand, with the priority control described above, the server 400 can handle multiple communication sessions under the same conditions, thereby improving processing efficiency. The transmission target adjustment unit 362, etc., may also be implemented in a combination of controlling the transmission target itself and performing priority control, as described above.
[0249] As a variation, the server 400 may transmit compressed and encoded hierarchical data to a relay server (not shown), and the relay server may simultaneously transmit data to multiple image processing devices 200. In this case, the server 400 may transmit data for all hierarchical levels to the relay server using the priority control described above, and the relay server may individually determine and transmit the data to each image processing device 200 according to its status. The transmission target adjustment unit 362 may further (1) change the parameters used by each hierarchical level, and / or (2) change the hierarchical relationships, based on at least one of the communication status and the content represented by the moving image described later.
[0250] In case (1), the transmission target adjustment unit 362, for example, if there is a 4K (3840 x 2160 pixels) layer as a resolution layer, will replace 4K with a WQHD (2560 x 1440 pixels) layer. Note that this does not mean adding a WQHD layer. Such changes may be made when the threshold for the higher bandwidth among the multiple thresholds for the realizable bandwidth is often below the threshold on the higher bandwidth side. Alternatively, if there is a QP level 1 layer as an image quality layer, the transmission target adjustment unit 362 will replace QP level 1 with a QP level 2 layer. In this case as well, this does not mean adding a QP level 2 layer.
[0251] Such changes may be implemented when the threshold for the higher bandwidth among the multiple thresholds for the available bandwidth is frequently exceeded. Alternatively, they may be implemented when the image analysis results indicate that image quality should be prioritized in a given scene. Case (2) will be discussed later. In cases (1) and (2), the transmission target adjustment unit 362 may, as necessary, request the resolution conversion unit 360, the first encoding unit 442, and the second encoding unit 444 to process under the determined conditions. Specific examples of what the moving image represents that serves as the basis for the adjustment will be described later.
[0252] Figure 24 illustrates the relationship between the image blocks compressed and encoded by the first encoding unit 442 and the second encoding unit 444, and the image blocks that reference the compressed encoding results, when scalable video encoding is implemented. This example assumes that there are two levels of resolution hierarchy, but no hierarchy for frame rate and image quality, with N being a natural number, and the left side shows the Nth frame and the right side shows the N+1th frame. In this case, the first encoding unit 442 compresses and encodes the low-resolution images 480a and 480b from each frame in units of image blocks (1-1), (1-2), ..., (1-n).
[0253] The second encoding unit 444 compresses and encodes the high-resolution images 482a and 482b in units of image blocks (2-1), (2-2), ..., (2-n). Here, the second encoding unit 444 uses the data of the image blocks at the corresponding positions of the low-resolution images 480a and 480b in the same frame, which were compressed and encoded by the first encoding unit 442, as a reference, as shown by the solid arrows (for example, solid arrows 484a and 484b), and compresses and encodes each image block of the high-resolution images 482a and 482b.
[0254] Alternatively, the second encoding unit 444 may use the compressed encoding result of the image block in the high-resolution image 482a of the Nth frame as a reference, as shown by the dashed-dotted arrow (for example, dashed-dotted arrow 488), and compress and encode the image block at the same position in the high-resolution image 482b of the N+1th frame. Alternatively, the second encoding unit 444 may use both the reference indicated by the solid arrow and the dashed-dotted arrow simultaneously to compress and encode the target image block.
[0255] Similarly, the first encoding unit 442 may use the compressed encoding result of an image block in the Nth frame's low-resolution image 480a as a reference, as shown by the dashed arrow (for example, dashed arrow 486), to compress and encode the image block at the same position in the N+1th frame's low-resolution image 480b. Figure 25 also illustrates the relationship between the image blocks compressed and encoded by the first encoding unit 442 and the second encoding unit 444, and the image blocks that reference the compressed encoding result for that purpose, when scalable video encoding is implemented.
[0256] This example assumes that both the resolution hierarchy and the frame rate hierarchy have two levels. The horizontal axis of the diagram represents the time axis, and from left to right, it shows the 1st, 2nd, 3rd, and 4th frames. As shown in the diagram, the 1st frame is an I-frame (intra-frame), and the other frames are P-frames (forward-predictive frames). However, this is not intended to limit the specific number of frames.
[0257] The vertical axis of the diagram represents the frame rate hierarchy, indicating that each hierarchy (level 0, level 1) has two levels of resolution images. Each image also has a code a-d in the upper left corner to identify the hierarchy combination. Representing each image using the frame number and the code a-d, the images compressed and encoded by the first encoding unit 442 are in the order 1a, 2c, 3a, 4c. The images compressed and encoded by the second encoding unit 444 are in the order 1b, 2d, 3b, 4d. When the communication bandwidth is smallest, the images displayed on the image processing device 200 are in the order 1a, 3a, ... When the communication bandwidth is largest, the images displayed on the image processing device 200 are in the order 1b, 2d, 3b, 4d.
[0258] Each encoding unit compresses and encodes the image in units of image blocks, as described above. Here, the second encoding unit 444 uses the data of the image block at the corresponding position in the low-resolution image in the same frame, which was compressed and encoded by the first encoding unit 442, as a reference, as shown by the solid arrows (for example, solid arrows 500a and 500b), and compresses and encodes each image block of the high-resolution image. Alternatively, the first encoding unit 442 and the second encoding unit 444 may use the compressed and encoded result of an image block of the same resolution in the same frame rate hierarchy as a reference, as shown by the dashed arrows (for example, dashed arrows 502a and 502b), and compress and encode the image block at the same position in the next frame.
[0259] Furthermore, the first encoding unit 442 and the second encoding unit 444 may use the compressed encoding result of an image block of the same resolution in different frame rate hierarchies as a reference, as indicated by the dashed-dotted arrow (for example, dashed-dotted arrows 504a and 504b), to compress and encode the image block at the same position in the next frame. The first encoding unit 442 and the second encoding unit 444 may use any of these references to compress and encode, or may use multiple references simultaneously to compress and encode.
[0260] Also, the processing procedure for compression encoding of each image block may adopt any of (c), (c)’, and (c)’’ shown in FIG. 22. Specifically, when the first encoding unit 442 and the second encoding unit 444 are actually one circuit or software module and perform processing in order, the procedure of (c) is adopted. When parallel processing is possible in the first encoding unit 442 and the second encoding unit 444, the procedure of (c)’ or (c)’’ is adopted. As described in (2) above, in the case of changing the hierarchical relationship based on at least any one of the communication situation and the content represented by the moving image, the transmission target adjustment unit 362 changes the number of hierarchical levels in at least any one of resolution, image quality, and frame rate. Or change the hierarchical level of the reference destination during compression encoding. For example, the transmission target adjustment unit 362 changes the reference destination as follows.
[0261] In the case of FIG. 25, for example, when there are level 1 and level 0 in the frame rate hierarchy and level 1 refers to level 0 every time, the frame sequence is Frame 1 Level 0 Frame 2 Level 1 (refer to the previous level 0 frame 1) Frame 3 Level 0 (refer to the previous level 0 frame 1) Frame 4 Level 1 (refer to the previous level 0 frame 3) Frame 5 Level 0 (refer to the previous level 0 frame 3) Frame 6 Level 1 (refer to the previous level 0 frame 5) Frame 7 Level 0 (refer to the previous level 0 frame 5) ··· It becomes like this. The transmission target adjustment unit 362 changes this, for example, by increasing the number of references between level 1 as follows.
[0262] Frame 1 Level 0 Frame 2 Level 1 (refer to the previous level 0 frame 1)[[ID=...]] Frame 3 Level 1 (refer to the previous level 1 frame 2) Frame 4 Level 0 (refer to the previous level 0 frame 1) Frame 5 Level 1 (refer to the previous level 0 frame 4) Frame 6, Level 1 (See previous Level 1 Frame 5) Frame 7, Level 0 (See previous Level 0 frame) ... Such changes may be implemented when the data arrival rate to the image processing device 200 frequently exceeds a predetermined threshold.
[0263] The compression encoding of multiple images corresponding to each frame described above divides the multiple images corresponding to each frame of a video into image blocks according to a rule based on the order in which the image data is acquired, and compresses and encodes each block. In this process, the compression ratio is improved by using the compression encoding result of one image to compress and encode another image. This reduces the time required for the compression encoding itself, and since transmission can begin from the image block as soon as the compression encoding is complete, low-latency image display is possible even when distributing images over a network.
[0264] Furthermore, because the size of the data to be transmitted can be reduced, it provides robustness against changes in communication conditions. For example, even when transmitting three or more image data, the required communication bandwidth can be reduced compared to transmitting each image in its entirety, and since the data is self-contained in image block units, recovery is easy even if data is lost during transmission. As a result, images with low latency can be displayed with high image quality. Moreover, it can flexibly adapt to various environments such as the number of image processing units at the destination, communication conditions, and the processing performance of the image processing units. For example, in live game broadcasting (eSports broadcasting) and cloud gaming with many participants, it can achieve high efficiency, low latency, and high image quality with high flexibility.
[0265] 6. Optimization of data size reduction methods To transmit data from server 400 to image processing device 200 reliably with low latency, it is desirable to minimize the size of the data to be transmitted. On the other hand, to provide a high level of user experience, such as a sense of presence and immersion in the displayed world, it is desirable to maintain a certain level of resolution and frame rate, and to avoid increasing the compression ratio, creating a dilemma. To strike a suitable balance between these two, server 400 optimizes the data transfer reduction method depending on the content of the image.
[0266] Figure 26 shows the configuration of a functional block of server 400, which has a function to optimize the data size reduction means. Server 400 includes an image generation unit 420, a compression encoding unit 422, a packetization unit 424, and a communication unit 426. The image generation unit 420, the packetization unit 424, and the communication unit 426 have the same functions as those described in Figures 5 and 12. Note that the image generation unit 420 may draw the moving image to be transmitted in place, that is, it may dynamically draw video that did not exist before.
[0267] Furthermore, when transmitting multiple images corresponding to each frame, the compression coding unit 422 may further include a splitting unit 440, a second coding unit 444, a transmission target adjustment unit 362, etc., as shown in Figures 20 and 23. The compression coding unit 422 includes an image content acquisition unit 450, a communication status acquisition unit 452, and a compression coding processing unit 454. The image content acquisition unit 450 acquires information relating to the content represented by the moving image to be processed.
[0268] The image content acquisition unit 450 acquires, for example, the characteristics of the image being drawn by the image generation unit 420 from the image generation unit 420 at predetermined timings, such as frame by frame, when the scene changes, and when the processing of the moving image begins. In this embodiment, since the moving image being displayed is basically generated and acquired in real time, the image content acquisition unit 450 can acquire accurate information related to the content without increasing the processing load, even at a fine-grained level such as frame by frame.
[0269] For example, the image content acquisition unit 450 can obtain information from the image generation unit 420 such as whether or not it is a scene change timing, the type of image texture represented in the frame, the distribution of feature points, depth information, the amount of objects, the amount of each level of the mipmap texture used for 3D graphics, the amount of LOD (Level Of Detail), the amount of each level of tessellation, the amount of characters and symbols, and the type of scene represented.
[0270] Here, the type of image texture refers to the type of region represented by the texture on the image, such as edge regions, flat regions, dense regions, detailed regions, and crowd regions. An edge is a location on the image where the rate of change in brightness value exceeds a predetermined value. The distribution of feature points refers to the position of feature points and edges on the image plane, as well as the intensity of edges, i.e., the rate of change in brightness value. The various texture regions described above are determined in two dimensions to determine what kind of texture they have in a two-dimensional image, which is the result of rendering as a three-dimensional graphic. Therefore, they are different from the textures used to generate the three-dimensional graphics.
[0271] Depth information is the distance to the object represented by each pixel, and in 3D graphics, it is obtained as the Z value. The quantity of objects refers to the number of objects represented, such as chairs or cars, or the area they occupy on the image plane. The mipmap texture level is the level of resolution selected in the mipmap technique, where texture data representing the surface of an object is prepared at multiple resolutions, and the appropriate resolution texture is used depending on the distance to the object and, consequently, its apparent size.
[0272] LOD and tessellation levels are levels of detail in techniques that represent objects with appropriate detail by adjusting the number of polygons depending on the distance of the object and, consequently, its apparent size. The types of scenes represented refer to the output situation and genre of content such as games that the image generation unit 420 is performing, and include types such as movie screens, menu screens, settings screens, loading screens, first-person perspective rendering, overhead perspective rendering, 2D pixel art games, 3D rendering games, first-person shooter games, racing games, sports games, action games, simulation games, and adventure novel games.
[0273] The image content acquisition unit 450 acquires at least one of the pieces of information from the image generation unit 420. The image content acquisition unit 450 and the score calculation function described later may be provided in the image generation unit 420. Specifically, a game engine or other software framework running on the image generation unit 420 may implement these functions. The image content acquisition unit 450 may also acquire at least one of the following transmitted from the image processing device 200: the timing of user operations on content defining moving images such as games, the interval between such operations, the content of the user operations, the status of the content as understood by the image generation unit 420, and the status of the audio generated by the content.
[0274] Since the timing of user operations may differ from the timing of frame drawing of moving images in the image generation unit 420, this can be used to detect scene changes and to adjust the data size and switch the amount of adjustment. The content status is, for example, information obtained from the image generation unit 420 that determines at least one of the following: a) a scene that requires user operation and affects the processing of the content, b) a movie scene that does not require user operation, c) a scene other than a movie scene that does not require user operation, or d) a scene that requires user operation but is not the main part of the content.
[0275] As a result, it is possible to optimize the data size adjustment means and the adjustment amount from the viewpoints of prioritizing responsiveness to user operations or prioritizing image quality. Also, the audio situation is information for discriminating at least any one of the presence or absence of audio, the number of audio channels, the content of background music, and the content of sound effects (SE), and is acquired from an audio generation device not shown in the figure.
[0276] Alternatively, the image content acquisition unit 450 may acquire at least any one of the above-described information by analyzing the image generated by the image generation unit 420 by itself. Also in this case, the timing of the analysis process is not particularly limited. Further, the image content acquisition unit 450 may acquire the type of the above-described scene or the like by reading the bibliographic information from a storage device not shown or the like when the server 400 starts distributing a moving image.
[0277] The image content acquisition unit 450 may also acquire information related to the content of the moving image by using the information acquired in the compression encoding process performed by the compression encoding unit 454. For example, when motion compensation is performed in the compression encoding process, the amount of optical flow, that is, in which direction the pixels are moving and the speed at which the pixel region is moving can be obtained. Also by motion estimation (ME), it is possible to obtain in which direction a rectangular region on the image is moving and its speed if it is moving.
[0278] Furthermore, in the compression encoding process, the allocation status of the encoding units used by the compression encoding unit 454 for the process, the timing of inserting an intra-frame, etc. can also be obtained. The latter serves as a basis for specifying the timing of scene switching in the moving image. Note that the image content acquisition unit 450 may acquire any of the parameters shown in the "score control rule at the scene switching timing" described later. The image content acquisition unit 450 may acquire at least any one of these information from the compression encoding unit 454 or may acquire the necessary information by analyzing the image by itself. Also in this case, the acquisition timing is not particularly limited.
[0279] As explained in Figure 23, the communication status acquisition unit 452 acquires the communication status with the image processing device 200 at a predetermined rate. The compression encoding processing unit 454 compresses and encodes the video data so that the data size is appropriate, using a method determined based on the content represented by the video, in response to changes in the communication status. Specifically, the compression encoding processing unit 454 changes at least one of the video's frame rate, resolution, and quantization parameters to achieve a data size determined by the communication status. Basically, it evaluates the content of the image from multiple perspectives and determines the optimal combination of numerical values for (frame rate, resolution, and quantization parameters).
[0280] In this process, the compression encoding processing unit 454 determines the means and amount of adjustment to be adjusted based on priorities that should be prioritized to best provide the user experience under the conditions of limited communication bandwidth and processing resources. For example, in scenes such as first-person shooters or fighting action games where the image movement is fast and immediate response to user input is required, or when realizing virtual or augmented reality, it is desirable to prioritize the frame rate. In scenes with many intricate objects where it is necessary to read characters, symbols, signs, or pixel art, it is desirable to prioritize the resolution.
[0281] In scenes where the image movement is slow but high contrast and dynamic range are required, and high image quality with minimal compression breakdown in tonal expression is necessary, it is desirable to prioritize quantization quality. Since the priorities differ depending on the content of the image, the compression encoding processing unit 454 takes this into account and ensures that the data size is appropriate in a way that maintains the user experience as much as possible even if the communication conditions deteriorate.
[0282] The rules for determining the adjustment means and adjustment amount from the content of the video are prepared in advance in the form of a table or calculation model and stored internally within the compression encoding processing unit 454. However, as mentioned above, the information indicating the content of the video is diverse, so it is necessary to optimize multiple adjustment means by looking at them all together. Therefore, the compression encoding processing unit 454 derives a score for determining the adjustment means and adjustment amount using at least one of the following rules.
[0283] 1. Rules for determining balance based on the content context The balance (priority) of weights given to frame rate, resolution, and quantization quality is determined by which of the above a-d conditions the content output is. Specifically, a) for scenes that require user interaction and affect the processing of the content, at least one of the scoring rules described below is followed. b) for movie scenes that do not require user interaction, quantization quality or resolution takes precedence over frame rate. c) for scenes other than movie scenes that do not require user interaction, resolution or frame rate takes precedence over quantization quality. d) for scenes that require user interaction but are not part of the main content, resolution or quantization quality takes precedence over frame rate.
[0284] 2. Rules for determining balance based on content genre In the case of (a) above, that is, a scene that requires user interaction and affects the processing of the content, the genre of the content being processed is read from a storage device or similar, and a separate table is prepared, which refers to the balance of scores representing the weights of resolution, frame rate, and quantization parameters generally recommended for each genre. Furthermore, when applying multiple rules described below in parallel and comprehensively evaluating their results to determine the final score, this rule may also be used as the initial value.
[0285] 3. Scoring rules based on the size of the image on the image. The system references LOD, mipmap textures, tessellation, and object size. The more objects there are that are higher resolution than a predetermined value, and the larger the proportion of the overall image they occupy (i.e., the closer the large, high-resolution objects are to the view screen), the higher the score representing the weight of quantization quality and resolution.
[0286] 4. Scoring rules based on object detail The more objects smaller than a predetermined value, the more detailed characters, symbols, signs, and pixel art there are, and the greater the proportion of the overall image area they occupy (i.e., the fewer large, high-resolution objects there are near the view screen), the higher the score representing the resolution weight.
[0287] 5. Scoring rules based on contrast and dynamic range By referencing contrast based on the distribution of pixel values and dynamic range based on the distribution of brightness, the score representing the weight of quantization quality is increased as the proportion of the total image area occupied by regions with higher contrast or higher dynamic range than a predetermined standard increases. 6. Scoring rules based on the movement of the image The system references the object's image movement, optical flow, and the magnitude and amount of motion estimation. The more objects whose movement is greater than a predetermined value, and the larger the proportion of the image they occupy, the higher the score representing the frame rate weight.
[0288] 7. Score determination rules based on texture type The image plane is divided into units, and the types of image textures in each unit region are tallied. The higher the proportion of dense, detailed, and crowded areas in the overall image, the higher the score representing the resolution weight.
[0289] 8. Score control rules at scene transition timings Information about the timing of the switch is obtained from the image generation unit 420. Alternatively, the score is switched when any of the following change abruptly or are reset in the time series direction: object quantity, feature points, edges, optical flow, motion estimation, pixel contrast, dynamic range of brightness, presence or absence of audio, number of audio channels, or audio mode. In this case, at least two frames are referenced. The compression encoding processing unit 454 can also detect scene changes based on correlation with the previous frame in order to understand the need for frames to be used as intraframes in compression encoding.
[0290] In this case, the accuracy of scene transition detection can be improved by using the score determination rules described above in addition to the determination by the conventional compression encoding processing unit 454. Since the scene transition detected in this way requires an intraframe, and the data size tends to increase rapidly, the score representing the weight of the quantization parameter and resolution is reduced for the target frame and a predetermined number of subsequent frames. Furthermore, prioritizing a rapid scene transition, the score representing the weight of the frame rate may be set high until the scene transition, and this constraint may be removed after the transition.
[0291] 9. Scoring rules based on the timing of user actions In the game and other content generated by the image generation unit 420, the longer the interval between user operations, the easier it is to lower the frame rate, so the score representing the weight of the frame rate is lowered (the shorter the interval between user operations, the lower the score representing the weight of other parameters).
[0292] 10. Scoring rules based on user actions In content such as games generated by the image generation unit 420, if the amount of change in user operations during the preceding predetermined period is large, it is presumed that the user expects high responsiveness from the content. Therefore, the larger the amount of change, the higher the score representing the weight of the frame rate. If operations are performed but the amount of change is small, the level of the score is lowered by one step.
[0293] User operations are acquired from the image processing device 200 by input information obtained from input devices (not shown) such as a game controller, keyboard, and mouse; the position, posture, and movement of the head-mounted display 100; gesture (hand sign) instructions based on the results of analyzing images taken by the head-mounted display 100's camera or an external camera (not shown); and voice instructions acquired by a microphone (not shown).
[0294] 11. Scoring rules for objects manipulated by the user When an object that meets predetermined criteria and is considered to reflect user actions during the preceding predetermined period is present on the screen, the score derived by other decision rules is adjusted in a direction that increases it. Here, the user may be a single person, or multiple people sharing the same image processing device 200, or having different image processing devices 200 connected to the same server 400. The objects on the screen that reflect user actions are typically the objects that the user is most interested in, such as people, avatars, robots, vehicles, and machines that are the target of the user's actions, objects that confront such main objects, and unit areas containing indicators intended to notify the player of information, such as the player's life status, weapon status, and game score in a game.
[0295] Information indicating the relationship between user operations and objects on the screen is, in principle, obtained from the image generation unit 420. The image content acquisition unit 450 may also infer this information. If such an object is detected according to a predetermined criterion, the compression encoding processing unit 454 will, for example, raise the score in the score determination in steps 3-9 and 10 above compared to other cases. In other words, this determination rule has the effect of preferentially applying the above determination rule to the object being operated on by the user. For example, in step 3, if the overall score of the screen based on the size of the image falls slightly short of the criterion score where the weight of quantization quality and resolution is high, the compression encoding processing unit 454 will raise the score to that criterion.
[0296] 6. If the overall score of the screen based on the movement of the image falls slightly short of the threshold score that would give a higher weight to the frame rate, the compression encoding processing unit 454 raises the score to that threshold. 8. Regarding the timing of scene transitions, if there is a sudden change in the area of the object being manipulated by the user, but the overall score of the screen falls slightly short of the threshold that would indicate the presence of an intra-frame, the compression encoding processing unit 454 raises the score to that threshold.
[0297] 12. Controlling the frequency of switching between the adjustment target and the adjustment amount. The system references the history of parameters (frame rate, resolution, quantization parameters) used in the preceding specified period and adjusts the scores representing the weights of each parameter so that the switch is acceptable from a user experience perspective. The "acceptable range" is determined based on a predefined table or model.
[0298] The compression encoding unit 454 comprehensively determines the final weights of the frame rate, resolution, and quantization parameters based on one of the above rules, or the sum of the scores obtained from multiple rules. In addition, the compression encoding unit 454 may refer to the Z value in rules 1 to 8 above to determine the position of the object, the projection order, the size, and the relative distance from the view screen.
[0299] It is desirable to optimize the decision rules described above. Therefore, the decision rules may be optimized using machine learning or deep learning while collecting the results of adjustments in various past cases. When using machine learning, the target of optimization may be either the table defining the decision rules or the computational model. In the case of deep learning, the computational model is optimized.
[0300] These learning techniques use, for example, manually created score databases or user-generated gameplay experiences as training data. Furthermore, subjective rendering cases are used as constraints on the computational model, and metrics such as PSNR (Peak Signal-to-Noise Ratio) and SSIM (Structural Similarity), parameter switching frequency, and time-series smoothness are used for training. In this example, parameter switching frequency is included as a metric because excessively frequent changes in resolution or frame rate can actually worsen the user experience.
[0301] In this way, the compression coding unit 454 derives a final score representing the weights assigned to each of the frame rate, resolution, and quantization quality, and determines a combination of values for (frame rate, resolution, and quantization parameters) such that a data size corresponding to the communication conditions is obtained and the balance represented by the score is followed. More specifically, the compression coding unit 454 predicts and calculates combinations of each value to satisfy the target data size determined according to the communication conditions, but in doing so, it controls the process so as not to reduce the weight of the parameters with high weights as much as possible.
[0302] Once the optimal combination is determined, the compression encoding processing unit 454 compresses and encodes the data accordingly. For example, to reduce the frame rate, the compression encoding processing unit 454 reduces the frames generated by the image generation unit 420 by a predetermined amount per unit time. To reduce the resolution, the compression encoding processing unit 454 applies existing calculations such as the nearest neighbor method, bilinear method, or bicubic method to the images generated by the image generation unit 420.
[0303] Furthermore, the compression encoding processing unit 454 compresses and encodes the image generated by the image generation unit 420 at an appropriate compression ratio using the adjusted quantization parameters. The compression encoding processing unit 454 performs all of these processes on a partial image basis for the frames generated by the image generation unit 420 and sequentially supplies them to the packetization unit 424. At this time, by supplying the frame rate, resolution, and quantization parameter data in association with each other, the packetization unit 424 packets the partial image data along with this data.
[0304] As a result, the compressed and encoded partial image data is transmitted to the image processing device 200 via the communication unit 426, as described above. The unit for adjusting the data size and the amount of adjustment used by the compression encoding processing unit 454 may be the unit of the partial image, which is the unit of compression encoding, or it may be one frame or a predetermined number of frames. The decoding and decompression unit 242 of the image processing device 200 performs dequantization based on the quantization parameters transmitted along with the image data.
[0305] Furthermore, if the resolution or frame rate is reduced from its original value, the image is displayed by adjusting its size to match the screen size of the connected display device or by outputting the same frame several times. General techniques can be used for such display processing procedures that respond to changes in resolution and frame rate. If the compression encoding unit 422 of server 400 reduces the resolution, server 400 may also transmit additional data that cannot be identified from the low-resolution image alone.
[0306] Here, the additional data includes, for example, feature quantities in the original image generated by the image generation unit 420 and various parameters determined by the compression encoding unit 422 during compression encoding. The feature quantities may include at least one of the following: feature points of the original image, edge intensity, depth of each pixel in the original image, texture type, optical flow, and motion estimation information. Alternatively, the additional data may include data indicating the object represented by the original image, identified by the object recognition processing performed by the compression encoding unit 422. In this case, the decoding and decompression unit 242 of the image processing device 200 accurately generates a high-resolution display image based on the transmitted image data and the additional data. This feature has been filed by the inventors as Japanese Patent Application No. 2019-086286.
[0307] Figures 27-31 illustrate examples of determining scores based on the content of a video. Figure 27 shows an example of applying the scoring rule described in 3 above, which is based on the size of the image on the video. The three images shown represent different frames from the same video. In each frame, objects 152a, 152b, and 152c, which use the same object model, are drawn. The object model actually contains three-dimensional information composed of polygons and textures.
[0308] The image generation unit 420 renders objects 152a, 152b, 152c, etc., by placing the object model in the rendering space and projecting it onto the view screen. In this process, the apparent size of objects 152a, 152b, and 152c changes depending on the distance from the view screen, so the image generation unit 420 adjusts the number of polygons used for rendering and the texture resolution as appropriate.
[0309] For example, while object model 150 is defined with 1 million polygons, object 152c, which is far from the view screen, is rendered with 10,000 polygons, and object 152a, which is close to the view screen, is rendered with 100,000 polygons. Also, the texture resolution is object 152a > object 152b > object 152c. The image content acquisition unit 450 acquires such information from the image generation unit 420 as information related to the content represented by the image, namely the LOD of the object as seen from the view screen, the level of the mipmap texture, the level of tessellation, and the size of the image.
[0310] The compression encoding unit 454 then adjusts the frame rate, resolution, and quantization parameters based on information related to the content represented by the image in order to achieve a data size that corresponds to the communication conditions. Specifically, it aggregates the LOD, mipmap texture level, tessellation level, and image size of all objects appearing in the scene and calculates a score to be given to the parameters to be adjusted. In the illustrated example, the image on the left has more finely rendered objects and occupies a larger proportion of the total area of the image (people and avatars are rendered in high resolution overall), so the score representing the weight of quantization quality and resolution is increased.
[0311] Figure 28 shows an example of applying the scoring rule based on object detail, as described in 4 above. The two images shown represent different frames from the same video. Each is a zoomed-in and zoomed-out frame of the same scene. When the size and quantity of objects represented in the images are aggregated, the image on the left has a larger quantity of small objects, and they occupy a larger proportion of the total area of the image (overall, there are many detailed objects, letters, symbols, signs, and pixel art). Therefore, in this example, the compression encoding processing unit 454 assigns a higher score, representing the resolution weight, to the image on the left.
[0312] Figure 29 shows an example of applying the scoring rules based on contrast and dynamic range described in item 5 above. When the contrast and dynamic range of each unit region, which is formed by dividing the image plane into predetermined sizes, are aggregated, the area of high-contrast or high-dynamic-range regions within the overall image is larger in the left image. Therefore, in this example, the compression encoding processing unit 454 assigns a higher score, which represents the weight of quantization quality, to the left image.
[0313] Figure 30 shows an example of applying the scoring rule based on image motion described in 6 above. The two images shown represent different frames of the same video. The images represent the absolute value of the amount of movement (vector) of objects in the image plane, and the magnitude and amount of optical flow and motion estimation for each unit region obtained by dividing the image plane into predetermined sizes. As a result, due to the presence of objects with large movements 152d to 152g, the left image has more objects with large movements and a larger proportion of their area in the overall image, meaning the sum of the absolute values of the amount of movement (vector) in the entire image is larger. Therefore, in this example, the compression encoding processing unit 454 assigns a higher score, which represents the weight of the frame rate, to the left image.
[0314] Figure 31 shows an example of applying the scoring rule based on texture type described in item 7 above. The two images shown represent different frames of the same video. When the types of image textures for each unit region, which is formed by dividing the image plane into predetermined sizes, are aggregated, the difference between objects 152h~152j and object 152k results in the left image having a larger proportion of dense, detailed, and crowded areas within the overall image. Therefore, in this example, the compression encoding processing unit 454 assigns a higher score, representing the resolution weight, to the left image.
[0315] Figure 32 is a flowchart showing the processing procedure by which the server 400 adjusts the data size according to the communication status. This flowchart is initiated when the user selects a game to play or a video to watch from the image processing device 200. In response, the communication status acquisition unit 452 of the server 400 starts acquiring the communication status to be used for streaming to the image processing device 200 (S50).
[0316] As described above, the communication status is determined by the transmission and reception of signals with the image processing device 200, so the communication status acquisition unit 452 acquires the necessary information from the communication unit 426. Specifically, the communication status acquisition unit 452 acquires the arrival delay time and the arrival rate of the image data transmitted to the image processing device 200, and derives an index representing the communication status from this information. The derivation rules are set in advance. Meanwhile, the image generation unit 420 starts generating moving images (S52). However, the image generation unit 420 is not limited to drawing computer graphics; it may also acquire captured images from the camera.
[0317] Meanwhile, the image content acquisition unit 450 of the compression encoding unit 422 starts acquiring information related to the content of the moving image (S54). As described above, the image content acquisition unit 450 acquires predetermined information at any time by acquiring information from the image generation unit 420 and the compression encoding processing unit 454, or by analyzing the image itself. Next, the compression encoding processing unit 454 determines the transmission size of the image data according to the most recent communication status acquired by the communication status acquisition unit 452 (S56).
[0318] The compression encoding processing unit 454 first obtains a score from the content of the video image acquired by the image content acquisition unit 450 at that time, in accordance with at least one of the rules described above (S58). Next, the compression encoding processing unit 454 derives the data size adjustment means and adjustment amount based on the obtained score (S60). That is, it calculates an overall weight for the frame rate, resolution, and quantization parameters by summing the scores, and determines the value of each parameter in a balance according to that weight to match the target data size.
[0319] The means for obtaining a score from parameters indicating the content of the video, and the means for deriving adjustment means and adjustment amounts from the score, may be provided in the form of a table representing the correspondence between them, or the derivation rules may be modeled and implemented in a program. In this case, the adjustment means and adjustment amounts may be changed depending on the level of communication conditions. Even when the communication conditions are stable and the transmittable data size does not change, the compression encoding processing unit 454 may continue to derive a score corresponding to the content of the video, and even if the transmission size remains the same, the combination of values for the frame rate, resolution, and quantization parameters may be changed.
[0320] Except in abnormal conditions, as long as transmission can be continued, no specific threshold is set for the communication status. Instead, the data transmission size is determined in multiple stages according to the communication status, and the combination of frame rate, resolution, and quantization parameters is adjusted accordingly. Furthermore, as mentioned above, the tables and computational models that define the derivation rules may be optimized as needed using machine learning or deep learning. In the illustrated procedure, a score value is obtained from the video content in S58, and the adjustment means and adjustment amount are determined based on that in S60, which is a two-step process. However, the decision rules may be prepared so that the adjustment means and adjustment amount can be determined directly from the video content.
[0321] The compression encoding processing unit 454 compresses and encodes the data in partial image units while adjusting the data size according to the adjustment means and adjustment amount determined in this way, and sequentially supplies the data of the image to the packetization unit 424 (S62). At this time, information on resolution, frame rate, and quantization parameters is also supplied along with it. The packetization unit 424 packets it and transmits it to the image processing device 200 from the communication unit 426. The compression encoding processing unit 454 may perform score acquisition and derivation of adjustment means and adjustment amount in the background of the compression encoding process. In practice, the frequency of data size determination in S56 and the frequency of compression encoding and transmission in partial image units in S62 may be the same or different.
[0322] As a result, the compression encoding processing unit 454 may reset the means and amount used to adjust the data size at predetermined time intervals. If there is no need to stop image transmission due to user operation on the image processing device 200 (N in S64), the compression encoding and transmission process is repeated for subsequent frames, changing the adjustment means and amount as needed (S56, S58, S60, S62). However, as described in 11 above, the compression encoding processing unit 454 makes adjustments within an acceptable range that does not worsen the user experience, after referring to the history of changes in the adjustment means and amount during the previous predetermined period. If it becomes necessary to stop image transmission, the server 400 terminates all processing (Y in S64).
[0323] As described above, the optimization of the data size reduction means adjusts the image data size in a manner and quantity appropriate to the information represented by the transmitted video, depending on the communication status used for streaming from the server 400 to the image processing device 200. By performing adjustments using a combination of multiple adjustment means such as resolution, frame rate, and quantization parameters, the variations in possible state changes can be significantly increased compared to adjusting only one parameter. As a result, the degradation of image quality can be kept to a level acceptable to the user as much as possible.
[0324] Furthermore, by obtaining information related to the content of the moving images from the image generation unit, which generates images in real time, and the compression encoding unit, which performs compression encoding, countermeasures can be taken based on more accurate and detailed data. In this case, the image generation unit and the compression encoding unit only need to perform their normal processing, so the data size can be adjusted in the most optimal way at each point in time without increasing the processing load.
[0325] 7. Compression ratio control for each region 7-1. Control based on the content represented by moving images. In step 6 above, data size was controlled by adjusting the combination of resolution, frame rate, and quantization parameters, using frames as the smallest unit, based on various information related to the content represented by the moving image. On the other hand, since the area of the displayed image that the user is focusing on is limited, allocating more communication bandwidth to that area can improve the user's perceived image quality even under the same conditions.
[0326] Figure 33 shows the configuration of a functional block of server 400, which has the function of changing the compression ratio depending on the region on the frame based on the content represented by the moving image. Server 400 includes an image generation unit 420, a compression encoding unit 422, a packetization unit 424, and a communication unit 426. The image generation unit 420, the packetization unit 424, and the communication unit 426 have the same functions as those described in Figures 5, 16, and 24. Note that, as mentioned above, the image generation unit 420 may draw the moving image to be transmitted in place, that is, it may dynamically draw video that did not exist before.
[0327] Furthermore, when transmitting multiple images corresponding to each frame, the compression coding unit 422 may further include a splitting unit 440, a second coding unit 444, a transmission target adjustment unit 362, etc., as shown in Figures 20 and 23. The compression coding unit 422 includes an image content acquisition unit 450, an attention estimation unit 460, a communication status acquisition unit 462, and a compression coding processing unit 464. As explained in Figure 26, the image content acquisition unit 450 acquires information relating to the content represented by the video image to be processed from the image processing device 200, the image generation unit 420, and the compression coding processing unit 464, or by performing image analysis itself.
[0328] The attention level estimation unit 460 estimates the user's level of attention for each unit region formed by dividing the frame plane of the video, based on information relating to the content represented by the video. Here, attention level is an indicator such as a numerical value that shows the degree to which the user is paying attention. For example, a high level of attention is estimated for unit regions that include areas where the main object is shown or areas where text is displayed. For this reason, the image content acquisition unit 450 may acquire information relating to the content of the video as illustrated in Figure 26, as well as information relating to the type and position of objects represented in the image.
[0329] The image content acquisition unit 450 may acquire this information by performing image recognition processing itself, or it may receive it as a result of image rendering from the image generation unit 420 or the like. In the former case, the image content acquisition unit 450 may use at least one of the acquired optical flow, motion estimation, coding unit allocation status, whether or not it is a scene change timing, the type of image texture represented in the frame, the distribution of feature points, depth information, etc., to perform image recognition or estimate the degree of attention of the recognized object.
[0330] The communication status acquisition unit 462 has the same function as the communication status acquisition unit 452 in Figures 23 and 26. The compression encoding processing unit 464 compresses and encodes the image data by varying the compression ratio in the image plane based on the distribution of attention levels estimated by the attention level estimation unit 460. In this process, the compression encoding processing unit 464 determines the compression ratio in the frame plane, and consequently the distribution of quantization parameters, based on the combination of attention level, the communication bandwidth available for data transmission, the frame rate, and the resolution.
[0331] Qualitatively, the compression encoding processing unit 464 reduces the compression ratio by lowering the value of the quantization parameter for unit quantities where a high level of attention is estimated, thereby allocating a larger bitrate. However, if the quantization parameter is determined solely based on the estimated level of attention, it may result in excessive compression or insufficient compression relative to the available communication bandwidth.
[0332] As described above, the compression coding processing unit 464 takes into account the frame rate and resolution and determines the distribution of compression ratios in the image plane so that the total data size of the frame is commensurate with the available communication bandwidth. However, as previously stated, the compression coding processing unit 464 performs compression coding on a sub-image basis for the frame generated by the image generation unit 420. Here, a sub-image is defined as, for example, an integer multiple of the size of the unit area used to estimate the degree of attention. The compression coding processing unit 464 determines quantization parameters for each unit area and performs compression coding using the determined quantization parameters for each unit area contained within the sub-image. At this time, the compression coding processing unit 464 grasps the total data size for each sub-image.
[0333] The compression encoding processing unit 464 sequentially supplies the compressed and encoded partial image data to the packetization unit 424. At this time, the packetization unit 424 packets the partial image data together with the quantization parameters applied to the unit region of the partial image by supplying them in association with the quantization parameters applied to the unit region of the partial image. As a result, the compressed and encoded partial image data is transmitted to the image processing device 200 via the communication unit 426, as described above. The decoding and decompression unit 242 of the image processing device 200 performs dequantization based on the quantization parameters transmitted along with the image data.
[0334] Figure 34 is a diagram illustrating the process by which the attention estimation unit 460 estimates the distribution of attention levels on the image plane. Image 160 represents a frame from the data of a moving image. Image 160 shows objects 162a, 162b and GUI 162c. The attention estimation unit 460 divides the image plane into predetermined intervals in both vertical and horizontal directions, as shown by dashed lines, to form unit regions (for example, unit region 164). Then, it estimates the attention level for each unit region.
[0335] The figure shows that the set of unit regions 166a and 166b, which represent objects 162a and 162b and are enclosed by a dashed line, are estimated to have a higher level of attention than other unit regions. However, in reality, given that more objects are represented, the level of attention is estimated based on various information, such as those exemplified in "6. Optimization of Data Size Reduction Methods" above. The level of attention may also be a binary value of 0 or 1, i.e., not being noticed or being noticed, or it may be represented by more gradations. For example, the level of attention estimation unit 460 may derive a score representing the level of attention according to at least one of the following rules, and the level of attention for each unit region may be comprehensively determined by summing these scores.
[0336] 1. Scoring rules based on object importance Unit areas containing objects that meet the criteria for high importance are given a higher score than others. These include objects that are the target of user interaction, such as people, avatars, robots, vehicles, and machines, as well as objects that confront such primary objects, and unit areas containing indicators intended to notify the player of information, such as the player's life status, weapon status, and game score in a game. These objects are estimated by the image content acquisition unit 450 either by obtaining information from the image generation unit 420 or by the correlation between objects identified by image recognition and user operations. The user may be a single person, or multiple people who share the same image processing device 200, or who have different image processing devices 200 connected to the same server 400.
[0337] 2. Scoring rules based on the area ratio of the object's image The larger the proportion of the image area occupied by objects such as people, avatars, or robots with similar characteristics, the higher the score of the unit region containing them. These objects are identified by the image content acquisition unit 450, which obtains information from the image generation unit 420, or by image recognition. General face detection algorithms may be used for image recognition. For example, faces can be detected by searching within the image for locations where the features of the eyes, nose, and mouth are relatively positioned in a T-shape.
[0338] 3. Scoring rules based on the presented information A unit area containing indicators intended to notify the player of information such as the player's life status, weapon status, and game score in the game is given a higher score than others. These indicators are identified by the image content acquisition unit 450, which obtains information from the image generation unit 420 or by image recognition.
[0339] 4. Scoring rules based on object detail When objects such as intricate characters, symbols, signs, and pixel art occupy a small proportion of the total area of an image, the unit region containing them is given a higher score than the others. These objects are identified by the image content acquisition unit 450, either by obtaining information from the image generation unit 420 or by image recognition.
[0340] 5. Scoring rules based on the size of the image on the picture. The system references LOD, mipmap texture, tessellation, and object size. The more objects with higher resolution than a predetermined value exist, and the larger the proportion of the overall image they occupy (i.e., the closer the large, high-resolution object is to the view screen), the higher the score of the unit region containing that object will be compared to others.
[0341] 6. Scoring rules based on contrast and dynamic range The system references contrast based on the distribution of pixel values and dynamic range based on the distribution of brightness, assigning higher scores to unit areas with higher contrast or dynamic range. 7. Scoring rules based on the movement of the image The system references the movement of the object's image, optical flow, and the magnitude and amount of motion estimation. Unit regions with a large number of objects whose movement is greater than a predetermined value, and which occupy a large proportion of the overall image area, are given higher scores.
[0342] 8. Score determination rules based on texture type The types of image textures in each unit region are tallied, and when a unit region containing a high-density texture, a detailed texture, or a crowded texture occupies a large proportion of the total area in the image, that unit region is given a higher score than others.
[0343] 9. Score control rules at scene transition timings Information about the timing of the switch is obtained from the image generation unit 420. Alternatively, the score is switched when any of the following change abruptly or are reset in the time series direction: object quantity, feature points, edges, optical flow, motion estimation, pixel contrast, dynamic range of brightness, presence or absence of audio, number of audio channels, or audio mode. In this case, at least two frames are referenced. The compression encoding processing unit 454 can also detect scene changes based on correlation with the previous frame in order to understand the need for frames to be used as intraframes in compression encoding.
[0344] In this case, the accuracy of scene change detection can be improved by using the score determination rules described above in addition to the determination by the conventional compression encoding processing unit 454. When it is determined that there is no scene change, the score of the high-scoring unit area set based on the other score determination rules for the frames up to that point is further increased. When it is determined that there is a scene change, the score of the unit area that was set as high is reset. Since an intraframe is required for the detected scene change, and the data size tends to increase dramatically, the total score of all unit areas is reduced for the target frame and a predetermined number of subsequent frames.
[0345] 10. Scoring rules based on user actions In content such as games generated by the image generation unit 420, if the amount of change in user operation during the preceding predetermined period is greater than or equal to a predetermined value, it is presumed that the user expects high responsiveness from the content and that there is an area of interest. Therefore, the score of a unit area that has been given a score greater than or equal to a predetermined value by other scoring rules is further amplified. If the amount of change is small even if an operation has been performed, the level of the score is lowered by one step from the amplified result. User operation is acquired from the image processing device 200 by input information acquired from input devices (not shown) such as a game controller, keyboard, and mouse, the position, posture, and movement of the head-mounted display 100, gesture (hand sign) instructions based on the results of analyzing images taken by the camera of the head-mounted display 100 or an external camera (not shown), and voice instructions acquired by a microphone (not shown).
[0346] 11. Controlling the switching frequency of the compression ratio distribution The system refers to the history of attention levels determined in the preceding specified period and adjusts the score so that the switch is within an acceptable range from a user experience perspective. The "acceptable range" is determined based on predefined tables and models.
[0347] The attention estimation unit 460 comprehensively determines the attention level of each unit area based on one of the above rules, or the sum of the scores obtained from multiple rules. The compression encoding processing unit 454 may refer to the Z value in the rules 1 to 9 above to determine the object's position, projection order, size, and relative distance from the view screen. The above determination rules are prepared in advance in the form of a table or calculation model and stored internally by the attention estimation unit 460.
[0348] The compression encoding processing unit 464 determines the compression ratio (quantization parameter) for each unit region based on the distribution of attention levels estimated in this way. For example, in the illustrated example, if one row of a unit region is considered a sub-image, the uppermost sub-image 168a and the lowermost sub-image 168b, which do not contain any unit regions determined to have a high level of attention, will have a higher compression ratio than the other intermediate sub-images. Furthermore, if object 162a has a higher level of attention than object 162b, the set of unit regions 166b representing object 162b may also have a higher compression ratio than the set of unit regions 166a.
[0349] In any case, the compression encoding processing unit 464 updates the distribution of the compression ratio at one of the following intervals: per partial image, per frame, per predetermined number of frames, or at a predetermined time interval, so that the data size matches the latest communication conditions. For example, if the available communication bandwidth, resolution, and frame rate are the same, the more unit regions that are estimated to be of high importance, the smaller the difference in compression ratio between the areas of interest and the areas of non-interest. Such adjustments may be made individually for each set of unit regions 166a and 166b, depending on the level of interest. In some cases, some sets of unit regions (for example, set of unit regions 166b) may be excluded from the reduction in compression ratio based on the determined priority of importance.
[0350] Furthermore, given the same available bandwidth, if the resolution and frame rate are high, the data size per frame is adjusted by increasing the overall compression ratio or reducing the difference in compression ratio between the region of interest and the region of non-interest, thereby optimizing the bitrate. While the transmission target is in the form of partial images, the data size control can be in the form of partial images, per frame, per predetermined number of frames, or per predetermined time unit. Rules for determining the distribution of quantization parameters based on the combination of the degree of interest of each unit region, the available bandwidth, the frame rate, and the resolution are prepared in advance in the form of a table or computational model.
[0351] However, as mentioned above, the information indicating the content of the video is diverse, so in order to determine the optimal quantization parameters by comprehensively considering all of this information, it is desirable to optimize the decision rule itself. Therefore, as explained in "6. Optimization of Data Size Reduction Methods" above, the decision rule may be optimized by machine learning or deep learning while collecting adjustment results from various past cases.
[0352] Figure 35 is a flowchart showing the processing procedure by which the server 400 controls the compression ratio for each region of the image plane. This flowchart is initiated when the user selects a game to play or a video to watch from the image processing device 200. In response, the communication status acquisition unit 462 of the server 400 starts acquiring the communication status used for streaming to the image processing device 200 (S70). As described above, the communication status is determined by the transmission and reception of signals with the image processing device 200, so the communication status acquisition unit 462 acquires the necessary information from the communication unit 426.
[0353] The image generation unit 420 also starts generating the corresponding video (S72). However, the image generation unit 420 is not limited to drawing computer graphics; it may also acquire captured images from the camera. Meanwhile, the image content acquisition unit 450 of the compression encoding unit 422 starts acquiring information related to the content of the video (S73). Next, the compression encoding processing unit 464 determines the transmission size of the image data according to the most recent communication status acquired by the communication status acquisition unit 462 (S74). Then, the attention level estimation unit 460 of the compression encoding unit 422 estimates the distribution of attention levels for the frames to be processed based on the information related to the content of the video (S75). That is, it divides the plane of the frame into unit regions and derives the attention level for each.
[0354] To derive the attention score, a table is referenced that allows obtaining scores according to at least one of the rules described above, or a calculation model is used. The scores obtained from various perspectives are then summed up to determine the final attention score for each unit area. The compression coding processing unit 464 determines the distribution of quantization parameters based on the distribution of attention scores (S76). As described above, the quantization parameters are determined considering not only the attention score but also the available communication bandwidth, resolution, and frame rate.
[0355] The compression coding processing unit 464 derives an appropriate distribution of quantization parameters from those parameters by referring to a table or using a computational model. Then, the compression coding processing unit 464 compresses and encodes the unit region contained within the partial image using the determined quantization parameters and sequentially supplies it to the packetization unit 424, which the communication unit 426 transmits to the image processing device 200 (S80).
[0356] The compression encoding processing unit 464 repeatedly compresses, encodes, and transmits all partial image data for the frame to be processed (S82N, S80). If there is no need to stop image transmission due to user operation on the image processing device 200 or other reasons (S84N), the determination of data size, estimation of attention distribution, determination of quantization parameter distribution, compression encoding, and transmission of partial images are repeated for subsequent frames (S84N, S74, S75, S76, S80, S82).
[0357] In practice, since the processing is pipelined on a partial image basis, the attention distribution estimation and quantization parameter distribution determination for the next frame can be performed while the compression encoding and transmission processing of the partial image of the previous frame are in progress. Also, in practice, the frequency of determining the data size in S74 and the frequency of compression encoding and transmission on a partial image basis in S80 may be the same or different. If it becomes necessary to stop transmitting images, the server 400 terminates all processing (Y in S84).
[0358] As described above, with its region-specific compression ratio control based on the content of the video, the server 400 estimates the user's level of attention for each unit region based on the content of the video. It then determines the distribution of quantization parameters so that the compression ratio decreases as the level of attention increases, compresses and encodes the image, and transmits it to the image processing device 200. This improves the quality of the user experience even in environments with limited communication bandwidth and resources.
[0359] Furthermore, by obtaining multifaceted information related to the content of the moving images from the image generation unit, which generates images in real time, and the compression encoding unit, it is possible to estimate the distribution of attention levels more accurately and in detail. By treating attention levels as a distribution and further determining the distribution of compression ratios according to this distribution, the compression ratio can be finely controlled in response to changes in resolution, frame rate, and communication conditions, allowing for flexible responses to various changing circumstances.
[0360] 7-2. Control based on user gaze point In section 7-1 above, attention levels were estimated based on the content represented by the video, but it is also possible to estimate attention levels based on the areas that the user is actually focusing on. In this case, the server 400 obtains location information of the user's gaze points on the screen from a device that detects the gaze points of the user looking at the head-mounted display 100 or the flat panel display 302, and estimates the distribution of attention levels based on this information. Alternatively, the distribution of attention levels may be derived based on both the content represented by the video and the user's gaze points.
[0361] Figure 36 shows the configuration of a functional block of server 400, which has the function of changing the compression ratio depending on the region on the frame based on the user's point of focus. Server 400 includes an image generation unit 420, a compression encoding unit 422, a packetization unit 424, and a communication unit 426. The image generation unit 420, the packetization unit 424, and the communication unit 426 have the same functions as those described in Figures 5, 9, 16, 26, and 33. Furthermore, when transmitting images from multiple viewpoints, the compression encoding unit 422 may further include a splitting unit 440, a second encoding unit 444, a transmission target adjustment unit 362, etc., as shown in Figures 20 and 23.
[0362] The compression coding unit 422 includes a gaze point acquisition unit 470, a focus estimation unit 472, a communication status acquisition unit 474, and a compression coding processing unit 476. The gaze point acquisition unit 470 acquires position information of the user's gaze point on a rendered moving image displayed on a display device connected to the image processing device 200, such as a head-mounted display 100 or a flat panel display 302. For example, a gaze point detector may be installed inside the head-mounted display 100, or a gaze point detector may be attached to the user looking at the flat panel display 302, and the measurement results from the gaze point detector may be acquired.
[0363] Then, the image data acquisition unit 240 of the image processing device 200 transmits the position coordinate information of the point of fixation to the communication unit 426 of the server 400 at a predetermined rate, and the point of fixation acquisition unit 470 acquires it. The point of fixation detector can be a general device that identifies the point of fixation from the direction of the pupil obtained by irradiating the user's eyeball with reference light such as infrared light and detecting the reflected light with a sensor. Preferably, the point of fixation acquisition unit 470 acquires the position information of the point of fixation at a frequency higher than the frame rate of the moving image.
[0364] Alternatively, the gaze point acquisition unit 470 may generate position information at a suitable frequency by temporally interpolating the position information of the gaze point acquired from the image processing device 200. In preparation for transmission failures from the image processing device 200, the gaze point acquisition unit 470 may simultaneously acquire a predetermined number of historical values along with the latest position information of the gaze point from the image processing device 200.
[0365] The attention estimation unit 472 estimates the level of attention for each unit region formed by dividing the frame plane of the moving image based on the position of the point of gaze. Specifically, the attention estimation unit 472 estimates the level of attention based on at least one of the following: the frequency with which the point of gaze is included, the dwell time of the point of gaze, the presence or absence of saccades, and the presence or absence of blinking. Qualitatively, the attention estimation unit 472 assigns a higher level of attention to a unit region the more frequently the point of gaze is included or the longer its dwell time is per unit time.
[0366] Saccades are rapid eye movements that occur when focusing gaze on an object, and it is known that the processing of visual signals in the brain is interrupted during the period when saccades occur. Naturally, images are not recognized during blinking. Therefore, the attention estimation unit 472 changes the difference in attention levels to reduce or eliminate them during these periods. A method for detecting saccades and blinking periods is disclosed, for example, in U.S. Patent Application Publication No. 2017 / 0285736.
[0367] If the gaze point acquisition unit 470 acquires or generates gaze point location information at a frequency (e.g., 240 Hz) higher than the frame rate of the video (e.g., 120 Hz), the attention level estimation unit 472 may update the attention level each time new location information is obtained. In other words, the distribution of attention levels may change during the compression encoding process for one frame. As described above, the attention level estimation unit 472 may further estimate the attention level based on the content shown in the video for each unit region, similar to the attention level estimation unit 460 described in 7-1, and integrate it with the attention level based on the gaze point location to obtain the final distribution of attention levels. The communication status acquisition unit 474 has the same function as the communication status acquisition unit 452 in Figure 26.
[0368] The compression coding processing unit 476 basically has the same function as the compression coding processing unit 464 in Figure 33. That is, the compression coding processing unit 476 compresses and encodes the image data by applying different compression ratios in the image plane based on the distribution of attention levels estimated by the attention level estimation unit 472. In this process, the compression coding processing unit 476 determines the compression ratio in the frame plane, and consequently the distribution of quantization parameters, based on the combination of communication bandwidth, frame rate, and resolution that can be used for data transmission.
[0369] For example, the compression coding processing unit 476 adjusts the distribution of the compression ratio based on a comparison between the data size of a predetermined number of recent frames and the available communication bandwidth acquired by the communication status acquisition unit 474. On the other hand, the movement of the point of attention is generally complex, and it may not be possible to obtain attention with stable accuracy. Therefore, in addition to the distribution of attention estimated by the attention estimation unit 472, the compression coding processing unit 476 may also refer to the compression results of recent frames to improve the accuracy of determining the compression ratio of the frame to be processed.
[0370] For example, the compression coding unit 476 may adjust the distribution of compression ratios of the frames to be processed based on the compression coding results of a predetermined number of recent frames, i.e., the regions with increased compression ratios and their compression ratios. This can mitigate bias in the distribution of compression ratios. Alternatively, the compression coding unit 476 may determine the effectiveness of the level of attention given to each unit region, and if it determines that it is not effective, it may refer to the compression coding results of a predetermined number of recent frames to determine the distribution of compression ratios.
[0371] For example, if areas of high importance are severely dispersed, or if areas of interest are outside the effective range, the compression coding unit 476 determines the compression ratio distribution for the frame to be processed by either adopting the compression ratio distribution determined for the most recent predetermined number of frames, or by extrapolating the changes in the distribution. Furthermore, as described above in 7-1, the compression coding unit 476 takes into account the frame rate and resolution, and determines the compression ratio distribution in the image plane so that the total data size of the frame is commensurate with the available communication bandwidth. Subsequent processing is the same as described above in 7-1.
[0372] Figure 37 is a diagram illustrating the process by which the attention estimation unit 472 estimates the distribution of attention in the image plane. Image 180 represents a frame from the data of a moving image. The circles (e.g., circle 182) shown on the image represent the positions where the point of focus remained, and the size of the circle represents the length of the stay. Here, "stay" refers to the point of focus remaining within a predetermined range considered to be the same position for a predetermined time or longer. The lines (e.g., line 184) represent the movement paths of the point of focus that occurred with a frequency of a predetermined value or more.
[0373] The attention estimation unit 472 acquires positional information of the gaze point from the gaze point acquisition unit 470 at predetermined intervals and generates information as shown in the figure. For each unit region, it estimates the level of attention based on the frequency in which the gaze point is included, the dwell time, the presence or absence of saccades, the presence or absence of blinking, etc. In the figure, for example, the unit regions included in the regions 186a, 186b, and 186c enclosed by the dashed lines are estimated to have a higher level of attention than the other unit regions. Furthermore, unit regions that become part of the viewpoint movement path at a frequency above a threshold may also be assigned a high level of attention. In this embodiment as well, the level of attention may be a binary value of 0 or 1, i.e., not being noticed or being noticed, or it may be represented by more gradations.
[0374] Figure 38 illustrates the method by which the compression coding unit 476 determines the distribution of compression ratios based on gaze points. Basically, when the compression coding unit 476 starts compressing a target frame, it determines the compression ratio (quantization parameters) for each partial image based on the distribution of the latest attention points and the available communication bandwidth, and then compresses and codes the image. However, the resulting data size may vary depending on the image content, and gaze point information may be updated in the unprocessed partial images within the target frame.
[0375] Therefore, the expected data size at the start of compression encoding of the target frame may differ from the actual result. For this reason, if the available communication bandwidth is temporarily exceeded, or is likely to be exceeded, the specification for reducing the compression ratio based on importance is canceled, and the data size is adjusted during the compression encoding process of the target frame or subsequent frames. In the figure, as shown in the upper panel, the plane of image 190 is divided horizontally into seven equal sub-regions (e.g., sub-region 192).
[0376] For example, when compressing and encoding each sub-image, if a sub-image contains a unit region where the compression ratio should be reduced, or if the sub-image is a unit region where the compression ratio should be reduced, the compression encoding processing unit 476 reduces the compression ratio by lowering the quantization parameters for that sub-image. Also, while compressing and encoding the sub-images of the image 190 sequentially from the upper section, if new position information of a point of focus is obtained, the attention estimation unit 472 updates the attention level for each unit region accordingly, and the compression encoding processing unit 476 determines the compression ratio of the sub-image based on the most recent attention level. In other words, when the attention level is updated after the start of processing of the frame to be processed, the compression encoding processing unit 476 compresses and encodes the sub-image with a compression ratio based on the distribution of the most recent attention levels.
[0377] As mentioned above, when determining the compression ratio, the compression ratio may be adjusted by referring to the results of previous compression encoding. The bar graph in the lower part of the figure shows the compressed data size of each partial image in four consecutive frames ("Frame0" to "Frame3") in bitrate (bit / sec) as the result of compression encoding. In addition, the bitrate per frame after compression encoding is shown in line graph 196 for each frame.
[0378] Here, bitrate "A" represents the communication bandwidth available for communication between the server 400 and the image processing device 200. However, in reality, this communication bandwidth fluctuates over time. The compression encoding processing unit 476 adjusts the compression ratio and, consequently, the quantization parameters of each partial image by comparing the bitrate per frame with the available communication bandwidth. In the example shown in the figure, the bitrate per frame of "Frame0" is well below the communication bandwidth, while in the next "Frame1" it approaches the communication bandwidth.
[0379] The compression encoding processing unit 476 predicts the bitrate per frame during the compression encoding process of each sub-image. If it determines that the difference with the communication bandwidth is smaller than a predetermined threshold, as in "Frame 1", it increases the compression ratio of one of the sub-images. In the figure, the arrows indicate that the bitrate decreased by increasing the compression ratio of the seventh sub-image in "Frame 1" from the initial decision.
[0380] Furthermore, in the next "Frame 2," the bitrate per frame exceeds the communication bandwidth. At this point, the compression encoding processing unit 476 may adjust the compression ratio for the next frame, "Frame 3." In the figure, the arrows indicate that the bitrate decreased by increasing the compression ratio of the first partial image of "Frame 3" from the initially determined value.
[0381] In this example, attention is represented by a binary value of "not attracting attention" or "attracting attention," and it corresponds to the operation of excluding the 7th subimage of "Frame1" and the 1st subimage of "Frame3," which were initially considered areas of attention, from the areas of attention. In other words, when the bitrate per frame exceeds the communication bandwidth, or when the difference between the two becomes smaller than a predetermined value, the compression encoding processing unit 476 cancels the reduction of the compression ratio even for subimages that contain a unit area of high attention where the compression ratio should be reduced.
[0382] In this case, the same compression ratio as other unfocused regions may be applied, or the compression ratio may be increased by a predetermined value. However, the adjustment method is not limited to those shown in the diagram, and especially when the level of focus is represented in multiple stages, the amount of change in the compression ratio may be varied according to the level of focus, or the entire distribution of the compression ratio may be changed. Thus, the compression coding processing unit 476 may determine the quantization parameters based only on the frame to be processed, or it may determine the quantization parameters by also referring to the compression coding results of a predetermined number of recent frames.
[0383] At this time, the compression encoding processing unit 476 may comprehensively check the time changes of parameters such as the distribution of quantization parameters in past frames, the available communication bandwidth, and the total data size of the frame, and determine the distribution of quantization parameters in the frame to be processed based on these. The rules for determining the distribution of quantization parameters based on attention level, or the rules for determining the distribution of quantization parameters based on the time changes of parameters such as the distribution of quantization parameters in past frames, the available communication bandwidth, and the total data size of the frame, may be optimized by machine learning or deep learning while collecting adjustment results from various past cases.
[0384] The processing procedure for server 400 to control the compression ratio for each region of the image plane may be the same as shown in Figure 35. However, the attention estimation process may be updated as needed whenever the position information of the point of focus is obtained, and the quantization parameters may be changed accordingly as needed.
[0385] As described above, with compression ratio control for each region based on the user's gaze point, server 400 acquires the user's actual eye movements and estimates the user's level of attention for each unit region based on the results. This makes it possible to accurately determine the object that is actually being focused on, and by preferentially allocating resources to that object, image quality can be improved in the user's perception.
[0386] Furthermore, by using the most recent compression encoding result to determine the compression ratio, the accuracy of the compression ratio setting can be maintained even if there are errors in the estimation of attention level. In addition, by monitoring the data size after compression encoding for the entire frame, even if there are too many areas judged to be of high attention in the initial settings, or if the compression ratio for those areas is set too low, these values can be appropriately adjusted. Also, similar to 7-1, by treating attention level as a distribution and further determining the distribution of the compression ratio according to that distribution, the compression ratio can be finely controlled in response to changes in resolution, frame rate, and communication conditions, allowing for flexible response to various situational changes.
[0387] As a variation, the gaze point acquisition unit 470 and the attention level estimation unit 472 may be provided on the image processing device 200 side, and the server 400 may acquire the attention level information estimated for each unit area from the image processing device 200. In this case, the gaze point acquisition unit 470 acquires the gaze point position information, which is the measurement result, from the gaze point detector described above. The attention level estimation unit 472 estimates the attention level for each unit area based on the gaze point position information, in the same manner as when it is provided on the server 400.
[0388] As mentioned above, the location information of the point of focus may be acquired at a higher frequency than the frame rate, and the distribution of attention levels may be updated accordingly. The attention level information estimated by the attention level estimation unit 472 is transmitted to the communication unit 426 of the server 400 at a predetermined rate, for example, via the image data acquisition unit 240. In this case as well, the attention level estimation unit 472 may simultaneously transmit a predetermined number of historical values along with the latest estimated value of the attention level to prepare for transmission failure. The operation of the communication status acquisition unit 474 and the compression encoding processing unit 476 of the server 400 is the same as described above. The same effects as described above can be achieved with this configuration as well.
[0389] The present invention has been described above based on embodiments. The embodiments are illustrative, and it will be understood by those skilled in the art that various modifications are possible in combinations of their components and processing processes, and that such modifications also fall within the scope of the present invention. [Explanation of Symbols]
[0390] 1 Image display system, 100 Head-mounted display, 200 Image processing device, 202 Input / output interface, 204 Partial image storage unit, 206 Control unit, 208 Video decoder, 210 Partial image storage unit, 212 Control unit, 214 Image processing unit, 216 Partial image storage unit, 218 Control unit, 220 Display controller, 240 Image data acquisition unit, 242 Decoding / decompression unit, 244 Image processing unit, 246 Display control unit, 248 Data acquisition status identification unit, 250 Output target determination unit, 252 Output unit, 260 Position and orientation tracking unit, 262 First correction unit, 264 Synthesis unit, 266 Second correction unit, 270a First forming unit, 270b Second forming unit, 272a First control unit, 272b Second control unit, 280 First decoding unit, 282 Second decoding unit, 302 Flat panel display, 400 Server, 402 Drawing control unit, 404 Image drawing unit, 406 Frame buffer, 408 Video encoder, 410 Partial image storage unit, 412 Control unit, 414 Video stream control unit, 416 Input / output interface, 420 Image generation unit, 422 Compression encoding unit, 424 Packetization unit, 426 Communication unit, 430 Drawing unit, 432 Formation content switching unit, 434 Data formation unit, 440 Splitting unit, 442 First encoding unit, 444 Second encoding unit, 450 Image content acquisition unit, 452 Communication status acquisition unit, 454 Compression encoding processing unit, 460 Attention level estimation unit, 462 Communication status acquisition unit, 464 Compression encoding processing unit, 470 Point of focus acquisition unit, 472 Attention level estimation unit, 474 Communication status acquisition unit, 476 Compression encoding processing unit.
Claims
1. A drawing unit that draws the moving image to be displayed, An image content acquisition unit that acquires information relating to the content represented by the aforementioned moving image, A unit for estimating the degree of attention a user is paying to for each unit region obtained by dividing the frame plane of the video at predetermined intervals, based on the content of the image in each unit region and the information relating to the content, A compression encoding unit compresses and encodes the video data by varying the compression ratio in the frame plane based on the distribution of attention in the frame plane, A communication unit that transmits compressed and encoded video data to a device that enables display, Equipped with, The attention estimation unit determines a score representing the attention level for each unit region based on a plurality of rules, reflecting the content represented by the moving image, determines a compression ratio for each unit region based on the sum of the scores, and, as one of the plurality of rules, amplifies the score of the unit region that has been given a score of a predetermined value or higher by another of the plurality of rules when the amount of change in the input value of user operations for the content defining the moving image during the immediately preceding predetermined period is greater than or equal to a predetermined value.
2. The image data transfer device according to claim 1, characterized in that the drawing unit dynamically draws a moving image to be transmitted that did not exist before.
3. The image data transfer device according to claim 1 or 2, characterized in that the image content acquisition unit acquires information relating to the content represented by the moving image by at least one of the following: notification from the drawing unit and the analysis results of the drawn moving image.
4. The image data transfer device according to claim 3, characterized in that the image content acquisition unit utilizes the information acquired by the compression encoding unit in the compression encoding process as an analysis result of the moving image.
5. The image data transfer device according to any one of claims 1 to 4, characterized in that the image content acquisition unit acquires at least one of the following as information relating to the content: the timing of the user operation, the interval of the timing, the content of the user operation, the status of the content, and the status of the audio generated by the content.
6. The image data transfer device according to claim 5, characterized in that the image content acquisition unit acquires information from the drawing unit to determine, as the status of the content, at least one of the following: a) a scene that requires user operation and affects the processing of the content, b) a movie scene that does not require user operation, c) a scene other than a movie scene that does not require user operation, or d) a scene that requires user operation but is not the main part of the content.
7. The image data transfer device according to any one of claims 1 to 6, characterized in that the attention estimation unit estimates the attention level based on at least one of the following: the amount of optical flow, the amount of motion estimation, the allocation status of encoding units, whether or not it is a scene switching timing, the type of image texture represented in the frame, feature points, depth information, the amount of objects, the amount of each level of mipmap texture used for 3D graphics, the amount of each level of LOD (Level Of Detail) and tessellation, the amount of characters and symbols, the type of scene represented, and the type of object represented.
8. The system further includes a communication status acquisition unit that acquires at least one of the data arrival rate and delay time of the video at the receiving device at predetermined time intervals to determine the communication status used for streaming video. The image data transfer device according to any one of claims 1 to 7, characterized in that the compression encoding unit determines the distribution of quantization parameters in the frame plane based on a combination of the attention level, the communication status, the frame rate, and the resolution.
9. The image data transfer device according to any one of claims 1 to 8, characterized in that the attention estimation unit makes the score of a unit region containing an object that satisfies the conditions for high importance larger than that of others.
10. The image data transfer device according to any one of claims 1 to 9, characterized in that the compression encoding unit determines the compression ratio based on rules optimized by machine learning or deep learning.
11. The image data transfer device according to any one of claims 1 to 10, characterized in that the compression encoding unit determines the distribution of compression ratios in the frame plane based on the total data size of the frame.
12. The image data transfer device according to any one of claims 1 to 11, characterized in that the compression encoding unit compresses and encodes the moving image in units of partial images smaller than one frame, and updates the distribution of the compression ratio in units of such partial images, units of one frame, units of a predetermined number of frames, or predetermined time intervals.
13. The communication unit sequentially transmits the data that the compression encoding unit has compressed and encoded for each compression unit formed by dividing the frame plane. The image data transfer device according to any one of claims 1 to 12, characterized in that the compression encoding unit reduces the compression ratio of a unit region when the compression unit includes a unit region in which the compression ratio should be reduced.
14. The steps include drawing the video image to be displayed, The steps include: obtaining information relating to the content represented by the aforementioned video image; The steps include: for each unit region obtained by dividing the frame plane of the video image at predetermined intervals, estimating the degree of attention the user is paying to each unit region based on the content of the image in that unit region, according to the information relating to the content; The steps include compressing and encoding the video data by varying the compression ratio in the frame plane based on the distribution of attention in the frame plane, The steps include: transmitting compressed and encoded video data to a device that enables display; Includes, The estimation step is characterized by determining a score representing the level of attention for each unit region based on a plurality of rules, reflecting the content represented by the moving image, determining a compression ratio for each unit region based on the sum of the scores, and, as one of the plurality of rules, amplifying the score of the unit region that has been given a score of a predetermined value or higher by another of the plurality of rules when the amount of change in the input value of user operations for the content defining the moving image during the immediately preceding predetermined period is greater than or equal to a predetermined value.
15. A function to draw the moving image to be displayed, A function to acquire information relating to the content represented by the aforementioned video image, A function that estimates the degree of attention a user is paying to each unit region, which is formed by dividing the frame plane of the video image into predetermined intervals, based on the content of the image in each unit region, using information related to the content. A function to compress and encode the video data by varying the compression ratio in the frame plane based on the distribution of attention in the frame plane, A function to transmit compressed and encoded video data to a device that enables display, To make this a reality on a computer, The estimation function is a computer program characterized by determining a score representing the level of attention for each unit region based on a plurality of rules, which reflects the content represented by the moving image, determining a compression ratio for each unit region based on the sum of the scores, and, as one of the plurality of rules, amplifying the score of the unit region that has been given a score of a predetermined value or higher by another of the plurality of rules when the amount of change in the input value of user operations for the content defining the moving image during the immediately preceding predetermined period is greater than or equal to a predetermined value.
Citation Information
Patent Citations
Image encoding device
JP1998079948A
High image quality medium synchronous reproducing method and device for encoding video image
JP1998200900A
Image processor, image processing method, semiconductor device, computer program and record medium
JP2004005452A
Video coding apparatus
JP2006080663A
Systems and methods for reducing the number of hops associated with head-mounted systems
JP2016533659A