Image generation method, display device and server
Through CloudVR technology, VR game rendering and interaction are separated, and cloud servers and display devices work together, the problems of high cost of terminal equipment and strict processing performance requirements in existing VR game technology are solved, and users need to play VR games anytime, anywhere, and improve display effect.
Patent Information
- Application Number
- PCT/CN2023/128393
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-10-31
- Publication Date
- 2025-09-04
AI Technical Summary
The existing virtual reality (VR) gaming technology has led to high cost and strict processing performance requirements for terminal devices, which cannot meet users' needs to play VR games anytime, anywhere.
CloudVR technology is used to separate VR game rendering and interaction. Cloud servers complete game rendering based on the received interactive instructions and transmit images to terminal devices through wireless networks. The display device obtains the human eye position information through the eye tracking module, and the server predicts the human eye position information to determine the dynamic high-definition extended area, and generates low-resolution and high-definition image data for shunt transmission.
It significantly reduces the cost and processing performance requirements of VR game terminal devices, enables users to play VR games through the Internet anytime, anywhere, and reduces the probability of lag in display devices and improves the display effect.
Smart Images

Figure CN2023128393_04092025_PF_FP_ABST
Abstract
Description
Image generation method, display device and server Technical Field
[0001] The present disclosure belongs to the field of virtual reality technology, and particularly relates to an image generation method, a display device, a server, an electronic device, and a computer non-transitory readable storage medium. Background Art
[0002] The virtual reality (VR) gaming industry is rapidly developing. High-load tasks such as rendering require significant computing resources, often requiring high-performance gaming consoles. This results in high costs for VR gaming and hinders the ability to play VR games anytime, anywhere. CloudVR technology leverages end-to-end cloud collaboration to separate VR game rendering from game interaction. Cloud servers render the game based on received interaction commands (such as eye tracking data) and transmit the game screen to the device via a wireless network. CloudVR technology can significantly reduce the cost of VR gaming devices and lower the performance requirements of VR processors, allowing users to access the internet and play VR games anytime, anywhere.
[0003] The head-mounted display device (HMD) starts the eye tracking module and obtains eye tracking data, and then uploads the tracking data to the image source end (source server, i.e., cloud server) through the encoding and network transmission module. After receiving the eye tracking data, the network transmission and encoding and decoding module at the image source end calculates the dynamic high-definition expansion area. The image processing module then renders a full low-definition image and a high-definition image based on the calculated area, which is then transmitted to the display device. The display device decodes the image, splices the gaze area and the non-gaze area, and displays the final image. After the HMD's viewpoint map is transmitted to the cloud server and displayed, the position of the human eye may have changed, resulting in a delay in the eye tracking area and causing dizziness.
[0004] Summary of the Invention
[0005] The present invention aims to solve at least one of the technical problems existing in the prior art and provides an image generation method, a display device, a server, an electronic device and a computer non-transitory readable storage medium.
[0006] In a first aspect, an embodiment of the present disclosure provides an image generation method, applied to a display device, comprising:
[0007] Send the acquired eye position information of the user at the current time t1 to the server;
[0008] Receiving an initial gaze area, second-resolution image data, and initial first-resolution image data of an image to be displayed from a server; the resolution of the second-resolution image data is smaller than the resolution of the initial first-resolution image data; the image generated corresponding to the initial first-resolution image data has a first resolution, and the image generated corresponding to the second-resolution image data has a second resolution; the first resolution is greater than the second resolution;
[0009] The initial gaze area, the second-resolution image data, and the initial first-resolution image data of the image to be displayed are obtained by the server through the following steps:
[0010] The server predicts the human eye position information at time t2 based on the human eye position information at time t1 and the source data, and obtains an initial gaze area and an initial non-gaze area of the image to be displayed based on the human eye position information at time t1 and the predicted human eye position information at time t2; there is a first time difference Δt between time t2 and time t1; and obtains the second-resolution image data and the initial first-resolution image data based on the initial gaze area and the source data.
[0011] In response to the acquired actual human eye position information at time t2 being located in the initial gaze area, using the initial gaze area of the image to be displayed as the actual gaze area of the image to be displayed, and using the initial first-resolution image data of the image to be displayed as the actual first-resolution image data;
[0012] The image to be displayed is generated based on at least the second resolution image data of the image to be displayed and the actual first resolution image data.
[0013] The method further includes: in response to the actual human eye position information obtained at time t2 not being located in the initial gaze area, determining the actual gaze area and actual first-resolution image data of the image to be displayed at time t2 based on the actual gaze area and actual first-resolution image data of the first N frames of the image to be displayed; wherein 0<N≤10, and N is a positive integer.
[0014] The step of determining the actual gaze area and the actual first-resolution image data of the image to be displayed at time t2 based on the actual gaze area and the actual first-resolution image data of the previous N frames of the image to be displayed includes:
[0015] similarity between the determined actual first-resolution image data of the preceding N frames of the image to be displayed and the initial first-resolution image data of the image to be displayed, and a positional relationship between the actual eye position information at time t2 and the actual gaze area of the preceding N frames of the image to be displayed;
[0016] The display image containing the actual eye position information at time t2 in the actual gaze area of the first N frames of the display image to be displayed is used as the candidate display image;
[0017] The actual first resolution image data of a frame having the highest similarity between the actual first resolution image data in the candidate display image and the initial first resolution image data of the image to be displayed is used as the actual first resolution image data of the image to be displayed.
[0018] The step of generating the image to be displayed based at least on the second resolution image data of the image to be displayed and the actual first resolution image data comprises:
[0019] Acquiring the processing performance required by the display device to process the image to be displayed;
[0020] In response to the required processing performance not exceeding a preset processing performance threshold, generating an image to be displayed based on the actual first-resolution image data and the second-resolution image data; in response to the required processing performance exceeding the preset processing performance threshold, cropping the actual gaze area, determining actual first-resolution image data after cropping based on the actual first-resolution image data, and generating an image to be displayed based on the actual first-resolution image data and the second-resolution image data after cropping.
[0021] The processing performance includes a processing speed, and obtaining the processing performance required by the display device to process the image to be displayed includes:
[0022] The processing speed required by the display device is calculated based on the actual gaze area, the actual non-gaze area, the actual first-resolution image data and the second-resolution image data of the image to be displayed.
[0023] The processing performance includes a processing speed, and obtaining the processing performance required by the display device to process the image to be displayed includes:
[0024] Based on obtaining a processing speed of the display device for processing the first M frames of the image to be displayed;
[0025] The processing speed required by the display device to process the image to be displayed is determined according to the processing speed of the first M frames of display image of the image to be displayed; 0<M≤10, and M is a positive integer.
[0026] The determining, based on the processing speed of the first M frames of the image to be displayed, the processing speed required by the display device to process the image to be displayed includes:
[0027] The processing speed required by the display device to process the image to be displayed is determined according to an average value of the processing speeds of the first M frames of display image of the image to be displayed.
[0028] The processing performance includes rendering capability; and obtaining the processing performance required by the display device to process the image to be displayed includes:
[0029] Acquire first data information and second data information of the image to be displayed; wherein the first data information includes at least the number of objects and the number of layers in the actual viewing area; and the second data information includes at least the number of objects and the number of layers in the actual non-viewing area;
[0030] Determine the size of the image of the actual viewing area according to the product of the number of objects in the first data information and the first value; determine the size of the image of the non-viewing area according to the product of the number of objects in the second data information and the second value;
[0031] determining, based on the first data information and the second data information, a processing complexity of an actual gaze area and a processing complexity of an actual non-gaze area of an image to be displayed;
[0032] The rendering capability required by the display device is calculated based on the processing complexity of the actual gaze area of the image to be displayed, the processing complexity of the actual non-gaze area, the size of the image of the actual gaze area, the size of the image of the actual non-gaze area, the first target resolution and the first target refresh frequency of the actual gaze area, and the second target resolution and the second target refresh frequency of the actual non-gaze area.
[0033] The clipping of the actual gaze area includes:
[0034] Determine the position information of the four vertices of the actual gaze area after clipping based on the actual eye position information of the human eye at time t2 and the pre-stored width and height of the gaze area;
[0035] The actual gaze area is clipped according to the position information of the four vertices of the determined clipped actual gaze area.
[0036] The display device includes a first mode and a second mode; and the method further includes:
[0037] When the processing performance required by the display device does not exceed the preset processing performance threshold, predicting the remaining play time of the display device in the first mode and the second mode respectively based on the actual non-gaze area, the actual gaze area, the actual first-resolution image data, the actual second-resolution image data, and the current remaining battery power of the image to be displayed;
[0038] When the predicted remaining play time of the display device in the first mode is greater than or equal to the play time preset by the user, generating an image to be displayed based on the actual first resolution image data and the second resolution image data;
[0039] When the predicted remaining play time of the display device in the first mode is less than the play time preset by the user, and the predicted remaining play time of the display device in the second mode is greater than or equal to the play time preset by the user, the actual gaze area is cropped, actual first-resolution image data after cropping is determined based on the actual first-resolution image data, and an image to be displayed is generated based on the actual first-resolution image data after cropping and the second-resolution image data;
[0040] When the predicted remaining play time of the display device in the second mode is less than the play time preset by the user, the actual gaze area is cut, and the actual first resolution image data after cut is determined based on the actual first resolution image data, and the image to be displayed is generated based on the actual first resolution image data after cut and the second resolution image data, and a prompt message is sent to the user.
[0041] In a second aspect, an embodiment of the present disclosure provides an image generation method, applied to a server, comprising:
[0042] Predicting the eye position information at time t2 based on the user's current eye position information at time t1 and the source data sent by the display device; there is a first time difference Δt between the time t2 and the time t1;
[0043] Based on the human eye position information at time t1 and the predicted human eye position information at time t2, obtaining an initial gaze area and an initial non-gaze area of the image to be displayed;
[0044] Obtaining second-resolution image data and initial first-resolution image data based on the initial gaze area and the source data; an image generated corresponding to the initial first-resolution image data has a first resolution, and an image generated corresponding to the second-resolution image data has a second resolution; the first resolution is greater than the second resolution;
[0045] The predicted human eye position information at time t2, the initial gaze area, the initial non-gaze area, the second resolution image data and the initial first resolution image data are sent to a display device, so that the display device determines the second resolution image data and the actual first resolution image data of the image to be displayed to generate the image to be displayed.
[0046] The method further comprises:
[0047] Get the network transmission rate;
[0048] In response to a network transmission speed being lower than a preset rate, only the second-resolution image data is sent to the display device.
[0049] In a third aspect, an embodiment of the present disclosure provides a display device, comprising:
[0050] An eye tracking module, configured to obtain position information of a human eye in real time;
[0051] A first sending module is configured to send the position information of the human eye to a server;
[0052] A first receiving module receives an initial gaze area, second-resolution image data, and initial first-resolution image data of an image to be displayed sent by the server; wherein the initial gaze area, second-resolution image data, and initial first-resolution image data of the image to be displayed are obtained by the server through the following steps:
[0053] The server predicts the human eye position information at time t2 based on the human eye position information at time t1 and the source data, and obtains an initial gaze area and an initial non-gaze area of the image to be displayed based on the human eye position information at time t1 and the predicted human eye position information at time t2; there is a first time difference Δt between time t2 and time t1; the second-resolution image data and the initial first-resolution image data are obtained based on the initial gaze area and the source data; the image generated corresponding to the initial first-resolution image data has a first resolution, and the image generated corresponding to the second-resolution image data has a second resolution; the first resolution is greater than the second resolution;
[0054] a judgment module configured to judge whether the actual human eye position information obtained at time t2 is located in the initial gaze area;
[0055] a data determining module, in response to the acquired actual human eye position information at time t2 being located in the initial gaze area, using the initial gaze area of the image to be displayed as the actual gaze area of the image to be displayed, and using the initial first-resolution image data of the image to be displayed as the actual first-resolution image data;
[0056] The image generation module is configured to generate the image to be displayed based on at least the second resolution image data of the image to be displayed and the actual first resolution image data.
[0057] In a fourth aspect, an embodiment of the present disclosure provides a server, comprising:
[0058] The prediction module is configured to predict the eye position information of the user at time t1 and the source data at time t2 according to the current eye position information of the user at time t1 sent by the display device; there is a first time difference Δt between the time t2 and the time t1;
[0059] an area determination module configured to obtain an initial fixation area and an initial non-fixation area of the image to be displayed based on the human eye position information at time t1 and the predicted human eye position information at time t2;
[0060] a data rendering module configured to obtain second-resolution image data and initial first-resolution image data based on the initial gaze area and source data; an image generated corresponding to the initial first-resolution image data has a first resolution, and an image generated corresponding to the second-resolution image data has a second resolution; the first resolution is greater than the second resolution;
[0061] The second sending module is configured to send the predicted human eye position information at time t2, the initial gaze area, the initial non-gaze area, the second resolution image data and the initial first resolution image data to the display device, so that the display device determines the second resolution image data and the actual first resolution image data of the image to be displayed to generate the image to be displayed.
[0062] In a fifth aspect, an embodiment of the present disclosure provides an electronic device, comprising:
[0063] one or more processors;
[0064] a memory for storing one or more programs;
[0065] When the one or more programs are executed by the one or more processors, the one or more processors implement any of the above-mentioned image generation methods.
[0066] In a sixth aspect, an embodiment of the present disclosure provides a computer non-volatile readable storage medium, wherein a computer program is stored on the computer non-volatile readable storage medium, and when the computer program is executed by a processor, the steps of the image generation method as described in any one of the above items are executed. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] FIG1 is a flowchart of an image generating method according to an embodiment of the present disclosure (applied to a display device).
[0068] FIG2 is a flowchart of another image generating method according to an embodiment of the present disclosure (applied to a display device).
[0069] FIG3 is a flowchart of another image generating method according to an embodiment of the present disclosure (applied to a display device).
[0070] FIG4 is a flowchart of an image generation method according to an embodiment of the present disclosure (applied to a server).
[0071] FIG5 is a schematic diagram of a display device according to an embodiment of the present disclosure.
[0072] FIG6 is a schematic diagram of a server according to an embodiment of the present disclosure.
[0073] FIG7 is a schematic diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0074] In order to enable those skilled in the art to better understand the technical solution of the present invention, the present invention is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0075] Unless otherwise defined, the technical or scientific terms used in this disclosure should have the usual meanings understood by people with ordinary skills in the field to which this disclosure belongs. The words "first", "second" and similar words used in this disclosure do not indicate any order, quantity or importance, but are only used to distinguish different components. Similarly, words such as "one", "an" or "the" do not indicate a quantity limitation, but rather indicate the existence of at least one. Words such as "include" or "comprise" mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. Words such as "connect" or "connected" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0076] Before introducing the embodiments of the present disclosure, it should be noted that the image generation method of the embodiments of the present disclosure is applied to a display device on the one hand and to a cloud server on the other hand, but the two application scenarios provided to achieve the final image display are ultimately achieved through information interaction between the display device and the cloud server. Next, the application of the image generation method of the embodiments of the present disclosure to a display device and a cloud server will be described respectively. In addition, for the sake of convenience of description, the image data corresponding to the first resolution image is referred to as high-definition image data, and the image data corresponding to the second resolution image is referred to as low-definition image data. The corresponding gaze area corresponds to the area where the high-definition image data is located, and the rest of the display device corresponds to the non-gaze area.
[0077] In a first aspect, embodiments of the present disclosure provide an image generation method. FIG1 is a flowchart of an image generation method according to an embodiment of the present disclosure (applied to a display device). As shown in FIG1 , the method is applied to a display device. The method may specifically include the following steps: S11, transmitting the acquired eye position information of the user at the current time t1 to a cloud server.
[0078] Specifically, the display device is integrated with an eye tracking module that can acquire real-time eye tracking data, i.e., the position of the user's eye at time t1. The position of the eye can be expressed using coordinates, with the position of the eye at time t1 being denoted as T1(x1, y1).
[0079] S12: Receive the initial gaze area, low-definition image data, and initial high-definition image data of the image to be displayed sent by the cloud server.
[0080] The initial gaze area, low-definition image data, and initial high-definition image data of the image to be displayed received from the cloud server in step S12 are obtained by the server using the following steps:
[0081] S01. The server predicts the eye position information at time t2 based on the eye position information at time t1 and the received source data. Based on the eye position information at time t1 and the predicted eye position information at time t2, the server obtains the initial gaze area of the image to be displayed. There is a first time difference Δt between time t2 and time t1, i.e., t2 = t1 + Δt. It should be noted that Δt primarily includes the round-trip network data transmission time and the image processing time. Source data refers to the image data in the video stream used to generate the image to be displayed.
[0082] In some examples, step S01 may include:
[0083] Step 1) Based on the human eye position information of n time nodes before time t1, the acquired human eye position information at time t1 is corrected using source data.
[0084] The specific step 1) includes: a. extracting edges within a circle with a radius of r, with the eye position at time t1 as the center, to obtain an edge image. b. determining the location with the highest edge density on the edge image, and correcting this location to the eye's gaze position at time t1, i.e., the actual eye position at time t1.
[0085] Step 2) Use a nonlinear equation to fit the corrected eye position information at time t1 to determine the gaze position p of the eye at time t, and obtain the eye movement trajectory function p=f(t).
[0086] Step 3) interpolate the eye movement trajectory function to obtain the predicted result p2=f(t2) of the eye position information at time t2, that is, obtain the predicted eye position information T2(x2, y2) at time t2.
[0087] Step 4) Based on the eye position information at time t1 and the predicted eye position information at time t2, the initial gaze area of the image to be displayed is obtained. For example: the eye position information at time t1 is T1 (x1, y1), the predicted eye position information at time t2 is T2 (x2, y2), the preset high-definition range is (±P, ±P), the correction coefficient is k, and T1 (x1, y1) is used as the center, and combined with the preset high-definition range of (±P, ±P), the first area is determined. Similarly, T2 (x2, y2) is used as the center, and combined with the preset high-definition range of (±P, ±P), the second area is determined. The first area and the second area are fitted to obtain the initial gaze area. At the same time, the initial non-attention area of the image to be displayed is also obtained (the remaining areas except the initial attention area are the initial non-attention areas).
[0088] S02. Rendering to obtain an initial low-definition image and an initial high-definition image according to the initial gaze area and the frame image data to be displayed at time t2.
[0089] When determining the initial gaze area of the image to be displayed, the cloud server fully considers the eye's position at time t1 (T1(x1, y1)) and predicts the eye's position at time t2 (T2(x2, y2)). This eliminates the possibility of prediction accuracy issues that could lead to loss of the original coordinates. The cloud server then sends the determined initial gaze area, the initial low-definition image, and the initial high-definition image to the display device.
[0090] S13. Obtain the actual eye position information at time t2, and determine whether the actual eye position information at time t2 is within the initial gaze area; if the actual eye position information at time t2 is within the initial gaze area, execute step S141; if the actual eye position information at time t2 is not within the initial gaze area, execute step S142.
[0091] Specifically, the cloud server sends the initial gaze area, low-definition image data, and initial high-definition image data of the image to be displayed to the display device. The eye tracking module on the display device then obtains the actual eye position information T2' (x2', y2') at the current time t2. The display device can then determine whether the actual eye position information is within the initial gaze area. Specifically, the display device needs to determine whether x2' and y2' simultaneously satisfy the following size relationship:
[0092] If the actual eye position information is within the initial gaze area, it means that the eye position information predicted by the cloud server at time t2 is relatively accurate, and the initial gaze area obtained based on the predicted eye position information is also relatively accurate. If the actual eye position information is not within the initial gaze area, it means that there is a deviation between the eye position information predicted by the cloud server at time t2 and the actual eye position information, so the initial gaze area obtained based on the predicted eye position information is also biased.
[0093] S141: Using the initial gaze area sent by the cloud server as the actual gaze area, and using the initial high-definition image data as the actual high-definition image data.
[0094] S142. Acquire the actual gaze area and actual high-definition image data of the first N frames of the image to be displayed, and use the actual gaze area and actual high-definition image data of the frame of display image whose actual eye position information at time t2 is closest to the initial high-definition image data as the actual gaze area and actual high-definition image data of the image to be displayed, where 0 < N ≤ 10, and N is a positive integer.
[0095] In some examples, in step S142, the display device may first determine the similarity between the actual high-definition image data of the first N frames of display images and the initial high-definition image data of the image to be displayed, and whether the actual human eye position information at time t2 is located within the actual gaze area of the first N frames of display images, and use the frame display image whose actual human eye position information at time t2 is located within the actual gaze area of the first N frames of display images as an alternative display image, and use the actual gaze area and actual high-definition image data of the alternative display image with the highest similarity to the initial high-definition image data of the image to be displayed as the actual gaze area and actual high-definition image data of the image to be displayed.
[0096] S15 . Generate an image to be displayed based on the actual high-definition image data and the low-definition image data and display it.
[0097] Specifically, in step S15, the processor of the display device may synthesize the obtained actual high-definition image data and the low-definition image data into an image to be displayed, and display the image.
[0098] The image generation method provided by the embodiment of the present disclosure reduces the processing task load of the display device by rendering the frame data of the image to be displayed in the cloud, and can greatly reduce the probability of the display device being stuck.
[0099] The present disclosure also provides an image generation method. FIG2 is a flowchart of another image generation method according to the present disclosure (applied to a display device). As shown in FIG2 , the method is also applied to a display device and includes steps S21-S25. Only step S25 differs from step S15 described above. The remaining steps are the same as those in the above example, so only step S25 will be described below. Step S25 may specifically include:
[0100] S251: Obtain the processing performance required by the display device to process the image to be displayed, and determine whether the required processing performance exceeds a preset processing performance threshold of the display device. If the required processing performance of the display device does not exceed the preset processing performance threshold, execute step S252; if the required processing performance of the display device exceeds the preset processing performance threshold, execute step S253.
[0101] S252: Generate an image to be displayed based on the actual high-definition image data and the low-definition image data, and display it.
[0102] S253: Crop the actual gaze area, determine cropped actual high-definition image data based on the actual high-definition image data, generate an image to be displayed based on the cropped actual high-definition image data and the low-definition image data, and display the image.
[0103] In some examples, the processing performance of the display device in step S251 may be a processing speed of the display device.
[0104] In one example, when the processing performance of the display device can be the processing speed of the display device, step S251 is a step of obtaining the processing performance required for the display device to process the image to be displayed. Specifically, the processing speed required for the display device can be calculated based on the actual gaze area, actual non-gaze area, actual high-definition image data and number of low-definition images of the image to be displayed.
[0105] In one example, when the processing performance of the display device can be the processing speed of the display device, step S251 of obtaining the processing performance required for the display device to process the image to be displayed can be specifically based on obtaining the processing speed of the display device for processing the first M frames of the image to be displayed; and determining the processing speed required for the display device to process the image to be displayed based on the processing speed of the first M frames of the image to be displayed; 0<M≤10, and M is a positive integer. For example, M=5. In some examples, determining the processing speed required for the display device to process the image to be displayed based on the processing speed of the first M frames of the image to be displayed includes: determining the processing speed required for the display device to process the image to be displayed based on an average of the processing speeds of the first M frames of the image to be displayed.
[0106] In some examples, the processing performance of the display device may be the rendering capability of the display device processor. When the processing performance of the display device is the rendering capability of the display device processor, the above step S251 may specifically include:
[0107] S2511. Obtain first data information and second data information for the image to be displayed. The first data information includes at least the number of objects and layers in the actual viewing area; the second data information includes at least the number of objects and layers in the actual non-viewing area. It should be noted that objects include, but are not limited to, people and objects.
[0108] In some examples, in step S2511, the display device can cluster the low-resolution image data separately to obtain the number of image clusters in the actual gaze area and the actual non-gaze area, thereby obtaining the number of objects in the actual gaze area and the actual non-gaze area. Simultaneously, the display device can parse the actual high-resolution image data and the low-resolution image data separately to obtain the number of layers in the high-resolution image data and the number of layers in the low-resolution image data. In other words, the first data information and the second data information can be obtained in this manner.
[0109] S2512. Determine the size of the image of the actual gaze area based on the product of the number of objects in the first data information and the first numerical value; determine the size of the image of the actual non-attention area based on the product of the number of objects in the second data information and the second numerical value; wherein the first numerical value is the ratio of the size of the gaze area to the size of the display device; and the second numerical value is the ratio of the size of the non-attention area to the size of the display device.
[0110] S2513: Determine, based on the first data information and the second data information, the processing complexity of the actual gaze area and the processing complexity of the actual non-gaze area of the image to be displayed.
[0111] In some embodiments, step S2513 may be a display device that determines the processing complexity of the image to be displayed based on the first data information and the second data information, as well as a preset complexity database. The preset complexity database may be pre-configured by the display device. The preset complexity database includes the correspondence between the first data information and the second data information and the processing complexity. For example, if the first data information and the second data information are both the number of objects, then the preset complexity database includes the correspondence between the number of objects and the processing complexity. The larger the number of objects, the greater the corresponding processing complexity. In this way, the display device can quickly and accurately determine the processing speed of the image to be displayed based on the preset complexity database.
[0112] S2514. Calculate the rendering capability required by the display device processor based on the processing complexity of the actual gaze area of the image to be displayed, the processing complexity of the actual non-gaze area, the size of the image in the actual gaze area, the size of the image in the actual non-gaze area, the first target resolution and the first target refresh frequency of the actual gaze area, and the second target resolution and the second target refresh frequency of the actual non-gaze area.
[0113] S2515: Determine the required rendering capability and compare it with a preset rendering capability threshold of the display device. If the required rendering capability of the display device does not exceed the preset rendering capability threshold, generate and display the image to be displayed based on the actual high-definition image data and the low-definition image data. If the required processing performance of the display device exceeds the preset processing performance threshold, crop the actual gaze area, determine cropped actual high-definition image data based on the actual high-definition image data, generate and display the image to be displayed based on the cropped actual high-definition image data and the low-definition image data.
[0114] In some examples, the step of cropping the actual gaze area may specifically include: determining, based on the actual eye position information T2' (x2', y2') at time t2 and the pre-stored width and height (w, h) of the gaze area, the four vertex coordinates of the cropped actual gaze area as (x2'-w / 2, y2'-h / 2), (x2'-w / 2, y2'+h / 2), (x2'+w / 2, y2'+h / 2), and (x2'+w / 2, y2'-h / 2), and cropping the actual gaze area based on the four vertex coordinates to serve as the cropped actual gaze area. By reducing the size of the actual gaze area, display smoothness is further ensured, thereby improving the display effect.
[0115] The present disclosure provides another image generation method. FIG3 is a flow chart of another image generation method according to the present disclosure (applied to a display device). As shown in FIG3 , the method is applied to a display device. The method further includes step S26 based on the above method. The display device includes a first mode and a second mode. The first mode is to directly display the actual high-definition image data and the low-definition image data. The second display mode is to cut the actual gaze area and display the actual gaze area after cutting. Specifically, step S26 includes:
[0116] Step S261: When the processing performance required by the display device does not exceed the preset processing performance threshold, the remaining play time of the display device in the first mode and the second mode is predicted respectively based on the actual non-gaze area, the actual gaze area, the actual high-definition image data, the low-definition image data, and the current remaining power of the image to be displayed.
[0117] It should be noted that when the device is started, it will perform power tests in the first mode and the second mode in two phases for a short period of time. That is, the average power consumption rate v1 of the first mode is tested within the Δt time, and then the average power consumption rate v2 of the second mode is tested within the next Δt time. Assuming that the total power is S, the time occupied by the first mode is calculated as The first mode takes time
[0118] Step S262: Determine the relationship between the remaining play time of the display device in the first mode and the second mode and the play time preset by the user.
[0119] When the predicted remaining play time of the display device in the first mode is greater than or equal to the play time preset by the user, step S252 is executed, that is, the image to be displayed is generated based on the actual high-definition image data and the low-definition image data.
[0120] When the predicted remaining play time of the display device in the first mode is less than the play time preset by the user, and the predicted remaining play time of the display device in the second mode is greater than or equal to the play time preset by the user, step S253 is executed, that is, the actual gaze area is cut, and the actual high-definition image data after cutting is determined based on the actual high-definition image data, and the image to be displayed is generated based on the actual high-definition image data after cutting and the low-definition image data.
[0121] When the predicted remaining play time of the display device in the second mode is less than the play time preset by the user, step S253 is executed, that is, the actual gaze area is cut, and the actual high-definition image data after cutting is determined based on the actual high-definition image data, and the image to be displayed is generated based on the actual high-definition image data after cutting and the low-definition image data.
[0122] In some examples, when the predicted remaining play time of the display device in the second mode is less than the user's preset play time, a prompt message may be sent to the user. The user can then choose to continue playing for the remaining time until the battery is depleted, or choose not to play for the time being, thereby ensuring that the user can maintain a sense of enjoyment while playing and avoiding a loss of enjoyment due to a sudden shutdown of the device during play.
[0123] In some examples, when the processing performance required by the display device exceeds the preset processing performance threshold, regardless of the relationship between the user's preset play time and the predicted play time in the first mode and the second mode, the actual gaze area of the image to be displayed needs to be cropped.
[0124] The image processing method provided by the embodiment of the present disclosure not only reduces the processing task of the display device by rendering the frame data of the image to be displayed in the cloud, which can greatly reduce the probability of the display device being stuck, but also can further process the data of the image to be displayed according to the required processing performance of the image to be displayed and the playing time, thereby improving the user experience after use.
[0125] In a second aspect, an embodiment of the present disclosure provides an image generation method. FIG4 is a flow chart of an image generation method according to an embodiment of the present disclosure (applied to a server). As shown in FIG4 , the method is applied to a cloud server and specifically includes:
[0126] S31. Predicting eye position information at time t2 based on the user's current eye position information at time t1 and source data sent by the display device; where there is a first time difference Δt between time t2 and time t1. Determining an initial gaze area and an initial non-gaze area of the image to be displayed based on the eye position information at time t1 and the predicted eye position information at time t2.
[0127] Step S31 is the same as the above-mentioned step S01, so it will not be repeated here.
[0128] S32: Rendering to obtain low-definition image data and initial high-definition image data according to the initial gaze area and the source data.
[0129] Step S32 is the same as the above-mentioned step S02, so it will not be repeated here.
[0130] S33. Send the predicted human eye position information at time t2, the initial gaze area, the initial non-gaze area, the low-definition image data and the initial high-definition image data to the display device, so that the display device determines the low-definition image data and the actual high-definition image data of the image to be displayed to generate the image to be displayed.
[0131] In some examples, before the cloud server sends the predicted eye position information at time t2, the initial gaze area, the initial non-gaze area, the low-definition image data, and the initial high-definition image data to the display device, the cloud server further includes:
[0132] The disclosed embodiment further includes the steps of obtaining a network transmission rate and determining a relationship between the network transmission rate and a preset rate. If the network transmission rate is less than the preset rate, only the low-definition image data is transmitted to the display device. If the network transmission rate is greater than or equal to the preset rate, both the low-definition image data and the high-definition data are transmitted to the display device.
[0133] That is to say, a network detection module is added to the cloud server. When it is detected that the network transmission rate between the display device and the cloud server is lower than the preset rate, only low-definition image data is transmitted, and high-definition image data is dropped frame processing. The display device only displays low-definition image data, and reduces the display resolution to avoid display freezes. Specifically, the display device obtains the human eye position information and transmits it to the cloud server. The cloud server predicts the initial gaze area based on the human eye position information and then renders and dedistorts the full-field image. After the rendering is completed, the eye movement information is predicted based on the display time to complete the initial high-definition image data rendering and low-definition image data rendering. The image is transmitted after judging the network rate. When the network transmission rate is lower than the preset rate, only the full-field image (low-definition image data) is transmitted. After the local end receives it, it is directly output to the system frame buffer for display. When the network transmission rate is higher than the preset rate, the full-field image and high-definition image data are transmitted locally. After receiving the image information, the local end synthesizes it based on the fitted eye movement information, and outputs it to the system frame buffer for display after completion.
[0134] In a third aspect, an embodiment of the present disclosure provides a display device. FIG5 is a schematic diagram of a display device according to an embodiment of the present disclosure. As shown in FIG5 , the display device can execute any of the image generation methods described in the first aspect. The display device includes an eye tracking module, a first sending module, a first receiving module, a determination module, a data determination module, and an image generation module.
[0135] The eye tracking module is configured to obtain the position information of the human eye in real time.
[0136] The first sending module is configured to send the position information of the human eye to the server.
[0137] The first receiving module receives the initial gaze area, low-definition image data, and initial high-definition image data of the image to be displayed sent by the server; wherein the initial gaze area, low-definition image data, and initial high-definition image data of the image to be displayed are obtained by the server through the following steps:
[0138] The server predicts the eye position information at time t2 based on the eye position information at time t1 and the source data, and obtains the initial gaze area and the initial non-gaze area of the image to be displayed based on the eye position information at time t1 and the predicted eye position information at time t2; there is a first time difference Δt between time t2 and time t1; based on the initial gaze area and the source data, low-definition image data and initial high-definition image data are rendered.
[0139] The judgment module is configured to judge whether the acquired actual human eye position information at time t2 is located in the initial gaze area.
[0140] The data determination module responds to the actual human eye position information at time t2 obtained and is located in the initial gaze area, and uses the initial gaze area of the image to be displayed as the actual gaze area of the image to be displayed, and uses the initial high-definition image data of the image to be displayed as the actual high-definition image data.
[0141] The image generation module is configured to generate an image to be displayed based on at least the low-definition image data and the actual high-definition image data of the image to be displayed.
[0142] In a fourth aspect, an embodiment of the present disclosure provides a server. FIG6 is a schematic diagram of the server according to an embodiment of the present disclosure. As shown in FIG6 , the server can execute any of the image generation methods described in the second aspect. The server includes a prediction module, a region determination module, a data rendering module, and a second sending module.
[0143] The prediction module is configured to predict the eye position information at time t2 based on the user's current eye position information at time t1 sent by the display device and source data; there is a first time difference Δt between the time t2 and the time t1.
[0144] The area determination module is configured to obtain an initial gaze area and an initial non-gaze area of the image to be displayed based on the human eye position information at time t1 and the predicted human eye position information at time t2.
[0145] The data rendering module is configured to render low-definition image data and initial high-definition image data according to the initial gaze area and the source data.
[0146] The second sending module is configured to send the predicted human eye position information at time t2, the initial gaze area, the initial non-gaze area, the low-definition image data and the initial high-definition image data to the display device, so that the display device determines the low-definition image data and the actual high-definition image data of the image to be displayed to generate the image to be displayed.
[0147] In a fifth aspect, an embodiment of the present disclosure further provides an electronic device. FIG7 is a schematic diagram of an electronic device provided by an embodiment of the present disclosure. As shown in FIG7 , the electronic device includes one or more processors 701, a memory 702, and one or more I / O interfaces 703. The memory 702 stores one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the display control method as described in any of the above embodiments. The one or more I / O interfaces 703 are connected between the processor and the memory and are configured to implement information exchange between the processor and the memory.
[0148] Among them, the processor 701 is a device with data processing capabilities, including but not limited to a central processing unit (CPU); the memory 702 is a device with data storage capabilities, including but not limited to random access memory (RAM, more specifically such as SDRAM, DDR, etc.), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), and flash memory (FLASH); the I / O interface (read-write interface) 703 is connected between the processor 701 and the memory 702, and can realize information exchange between the processor 701 and the memory 702, including but not limited to a data bus (Bus), etc.
[0149] In some embodiments, the processor 701 , the memory 702 , and the I / O interface 703 are connected to each other via a bus 704 , and further connected to other components of the computing device.
[0150] In some embodiments, the one or more processors 701 include a field programmable gate array (FPGA).
[0151] In a sixth aspect, embodiments of the present disclosure further provide a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the steps of any of the display control methods described in the above embodiments.
[0152] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a machine-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication part, and / or installed from a removable medium. When the computer program is executed by a central processing unit (CPU), the above-mentioned functions defined in the system of the present disclosure are executed.
[0153] It should be noted that the computer non-transitory readable medium shown in the present disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination of the above. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device or device. In the present disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any non-transitory computer-readable storage medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the non-transitory computer-readable storage medium may be transmitted using any suitable medium, including but not limited to wireless, wire, optical cable, RF, etc., or any suitable combination thereof.
[0154] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architectures, functions and operations of the devices, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the aforementioned module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the boxes can also occur in an order different from that marked in the accompanying drawings. For example, two connected boxes can actually represent execution in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0155] It will be understood that the above embodiments are merely exemplary embodiments for illustrating the principles of the present invention, and the present invention is not limited thereto. Those skilled in the art will appreciate that various modifications and improvements can be made without departing from the spirit and substance of the present invention, and such modifications and improvements are also considered to be within the scope of protection of the present invention.
Claims
1. An image generation method, applied to a display device, comprising: Sending the obtained eye position information of the user at the current time t1 to the server; Receiving the initial fixation region of the image to be displayed, the second-resolution image data, and the initial first-resolution image data sent by the server; The image generated corresponding to the initial first-resolution image data has a first resolution, and the image generated corresponding to the second-resolution image data has a second resolution; the first resolution is greater than the second resolution; Wherein, the initial fixation region of the image to be displayed, the second-resolution image data, and the initial first-resolution image data are obtained by the server through the following steps: The server predicts the eye position information at time t2 based on the eye position information at time t1 and the source data, and obtains the initial fixation region and the initial non-fixation region of the image to be displayed based on the eye position information at time t1 and the predicted eye position information at time t2; there is a first time difference Δt between time t2 and time t1; according to the initial fixation region and the source data, the second-resolution image data and the initial first-resolution image data are obtained; In response to the obtained actual eye position information at time t2 being located in the initial fixation region, taking the initial fixation region of the image to be displayed as the actual fixation region of the image to be displayed, and taking the initial first-resolution image data of the image to be displayed as the actual first-resolution image data; Generating the image to be displayed based at least on the second-resolution image data and the actual first-resolution image data of the image to be displayed.
2. The image generation method according to claim 1, wherein, Further comprising: in response to the obtained actual eye position information at time t2 not being located in the initial fixation region, determining the actual fixation region and the actual first-resolution image data of the image to be displayed at time t2 based on the actual fixation regions and the actual first-resolution image data of the previous N frames of the image to be displayed; wherein, 0 < N ≤ 10, and N is a positive integer.
3. The image generation method according to claim 2, wherein, The step of determining the actual fixation region and the actual first-resolution image data of the image to be displayed at time t2 based on the actual fixation regions and the actual first-resolution image data of the previous N frames of the image to be displayed comprises: Determining the similarity between the actual first-resolution image data of the previous N frames of the image to be displayed and the initial first-resolution image data of the image to be displayed, and the positional relationship between the actual eye position information at time t2 and the actual fixation regions of the previous N frames of the image to be displayed; Taking the display image that contains the actual eye position information at time t2 within the actual fixation regions of the previous N frames of the image to be displayed as the alternative display image; Taking the actual first-resolution image data of the frame with the highest similarity between the actual first-resolution image data of the alternative display image and the initial first-resolution image data of the image to be displayed as the actual first-resolution image data of the image to be displayed.
4. The image generation method according to claim 1, wherein, The step of generating the image to be displayed based at least on the second-resolution image data and the actual first-resolution image data of the image to be displayed includes: Obtaining the processing performance required by the display device to process the image to be displayed; In response to the required processing performance not exceeding a preset processing performance threshold, generating an image to be displayed based on the actual first-resolution image data and the second-resolution image data; in response to the required processing performance exceeding the preset processing performance threshold, cropping the actual fixation area, determining the cropped actual first-resolution image data according to the actual first-resolution image data, and generating an image to be displayed according to the cropped actual first-resolution image data and the second-resolution image data.
5. The image generation method according to claim 4, wherein, The processing performance includes processing speed, and the obtaining of the processing performance required by the display device to process the image to be displayed includes: Calculating the processing speed required by the display device based on the actual fixation area, actual non-fixation area, actual first-resolution image data, and second-resolution image data of the image to be displayed.
6. The image generation method according to claim 4, wherein, The processing performance includes processing speed, and the obtaining of the processing performance required by the display device to process the image to be displayed includes: Based on obtaining the processing speed of the previous M display images of the display device for processing the image to be displayed; Determining the processing speed required by the display device to process the image to be displayed according to the processing speed of the previous M display images of the image to be displayed; 0 < M ≤ 10, and M is a positive integer. The determining of the processing speed required by the display device to process the image to be displayed according to the processing speed of the previous M display images of the image to be displayed includes:
7. The image generation method according to claim 6, wherein, Determining the processing speed required by the display device to process the image to be displayed according to the average value of the processing speeds of the previous M display images of the image to be displayed. The processing performance includes rendering ability; the obtaining of the processing performance required by the display device to process the image to be displayed includes:
8. The image generation method according to claim 4, wherein, Obtaining first data information and second data information of the image to be displayed; wherein, the first data information at least includes the number of objects and the number of layers in the actual fixation area; the second data information at least includes the number of objects and the number of layers in the actual non-fixation area; Determining the size of the image in the actual fixation area according to the product of the number of objects in the first data information and a first value; determining the size of the image in the non-fixation area according to the product of the number of objects in the second data information and a second value; Determining the processing complexity of the actual fixation area and the processing complexity of the actual non-fixation area of the image to be displayed according to the first data information and the second data information; Calculate the rendering capability required by the display device based on the processing complexity of the actual fixation area of the image to be displayed, the processing complexity of the actual non-fixation area, the size of the image in the actual fixation area, the size of the image in the actual non-fixation area, the first target resolution and the first target refresh rate of the actual fixation area, and the second target resolution and the second target refresh rate of the actual non-fixation area.
9. The image generation method according to claim 4, wherein, The shearing of the actual fixation area includes: Determine the position information of the four vertices of the sheared actual fixation area according to the actual eye position information of the human eye at time t2 and the width and height of the pre-stored fixation area; Shear the actual fixation area according to the determined position information of the four vertices of the sheared actual fixation area.
10. The image generation method according to claim 4, wherein, The display device includes a first mode and a second mode; the method further includes: When the required processing performance of the display device does not exceed the preset processing performance threshold, predict the remaining play time of the display device in the first mode and the second mode respectively according to the actual non-fixation area, the actual fixation area, the actual first-resolution image data, the second-resolution image data, and the current remaining battery power of the image to be displayed; When the predicted remaining play time of the display device in the first mode is greater than or equal to the preset play time of the user, generate an image to be displayed according to the actual first-resolution image data and the second-resolution image data; When the predicted remaining play time of the display device in the first mode is less than the preset play time of the user, and the predicted remaining play time of the display device in the second mode is greater than or equal to the preset play time of the user, shear the actual fixation area, determine the sheared actual first-resolution image data according to the actual first-resolution image data, and generate an image to be displayed according to the sheared actual first-resolution image data and the second-resolution image data; When the predicted remaining play time of the display device in the second mode is less than the preset play time of the user, shear the actual fixation area, determine the sheared actual first-resolution image data according to the actual first-resolution image data, generate an image to be displayed according to the sheared actual first-resolution image data and the second-resolution image data, and send a prompt message to the user.
11. An image generation method applied to a server, which includes: Predict the eye position information at time t2 according to the eye position information of the user at the current time t1 and the source data sent by the display device; There is a first time difference △t between time t2 and time t1; Based on the eye position information at time t1 and the predicted eye position information at time t2, obtain the initial fixation area and the initial non-fixation area of the image to be displayed; According to the initial fixation area and the source data, obtain the second-resolution image data and the initial first resolution image data; The image generated corresponding to the initial first-resolution image data has a first resolution, and the image generated corresponding to the second-resolution image data has a second resolution; the first resolution is greater than the second resolution; Send the predicted eye position information, the initial fixation region, the initial non-fixation region, the second-resolution image data, and the initial first-resolution image data at the t2 moment to the display device, so that the display device determines the second-resolution image data and the actual first-resolution image data of the image to be displayed, and generates the image to be displayed.
12. The image generation method according to claim 11, wherein, It further includes: Obtain the network transmission rate; In response to the network transmission speed being less than the preset rate, only send the second-resolution image data to the display device.
13. A display device, which includes: An eye tracking module configured to obtain the position information of the human eye in real time; A first sending module configured to send the position information of the human eye to the server; A first receiving module that receives the initial fixation region, the second-resolution image data, and the initial first-resolution image data of the image to be displayed sent by the server; wherein, the initial fixation region, the second-resolution image data, and the initial first-resolution image data of the image to be displayed are obtained by the server through the following steps: The server predicts the eye position information at the t2 moment according to the eye position information at the t1 moment and the source data, and based on the eye position information at the t1 moment and the predicted eye position information at the t2 moment, obtains the initial fixation region and the initial non-fixation region of the image to be displayed; there is a first time difference △t between the t2 moment and the t1 moment; according to the initial fixation region and the source data, render to obtain the second-resolution image data and the initial first-resolution image data; the image generated corresponding to the initial first-resolution image data has a first resolution, and the image generated corresponding to the second-resolution image data has a second resolution; the first resolution is greater than the second resolution; A judgment module configured to judge whether the obtained actual eye position information at the t2 moment is located in The initial fixation region; A data determination module, in response to the obtained actual eye position information at the t2 moment being located in the initial fixation region, uses the initial fixation region of the image to be displayed as the actual fixation region of the image to be displayed, and uses the initial first-resolution image data of the image to be displayed as the actual first-resolution image data; An image generation module configured to generate the image to be displayed based at least on the second-resolution image data and the actual first-resolution image data of the image to be displayed.
14. A server, which includes: A prediction module configured to predict the eye position information at the t2 moment according to the current eye position information of the user at the t1 moment sent by the display device and the source data; There is a first time difference △t between the t2 moment and the t1 moment; An area determination module, configured to obtain an initial fixation area and an initial non-fixation area of the image to be displayed based on the human eye position information at the moment t1 and the predicted human eye position information at the moment t2; A data rendering module, configured to obtain second-resolution image data and initial first-resolution image data according to the initial fixation area and the source data; The image generated corresponding to the initial first-resolution image data has a first resolution, and the image generated corresponding to the second-resolution image data has a second resolution; the first resolution is greater than the second resolution; A second sending module, configured to send the predicted human eye position information at the moment t2, the initial fixation area, the initial non-fixation area, the second-resolution image data, and the initial first-resolution image data to a display device, so that the display device determines the second-resolution image data and the actual first-resolution image data of the image to be displayed, and generates the image to be displayed.
15. An electronic device, wherein, Comprising: One or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the image generation method according to any one of claims 1 to 12.
16. A computer non-transitory readable storage medium, wherein, A computer program is stored on the non-transitory computer-readable storage medium, and when the computer program is run by a processor, it executes the steps of the image generation method according to any one of claims 1 to 12.