Image display method, device, electronic equipment and storage medium

By employing bidirectional weighted prediction and caching techniques in extended reality devices, the problem of uneven image smoothness caused by network disturbances and inconsistent decoding time is solved, resulting in a smoother image display effect.

CN116233444BActive Publication Date: 2025-11-18BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310118527.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-13
Publication Date
2025-11-18
Estimated Expiration
2043-02-13

AI Technical Summary

Technical Problem

In extended reality devices, inaccurate predictions of asynchronous time distortion and asynchronous space distortion caused by network disturbances and uneven decoding time result in unsmooth images.

Method used

By acquiring the current video frame data and the predicted video frame data from the previous cycle, a bidirectional weighted prediction is performed to generate the second predicted video frame data, which is then inserted between the current video frame data. This, combined with caching technology, addresses the discrepancy between the predicted frame and the actual frame, thereby improving the smoothness of the video.

Benefits of technology

It effectively solves the problem of uneven picture quality caused by network disturbances and differences in decoding time, and improves the frame rate within the allowable latency range, thereby enhancing the user's visual experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116233444B_ABST
    Figure CN116233444B_ABST
Patent Text Reader

Abstract

The present disclosure provides an image display method, device, electronic equipment and storage medium. The image display method comprises: acquiring current video frame data and first predicted video frame data inserted in a last period; performing image prediction according to the current video frame data and the first predicted video frame data to obtain second predicted video frame data; inserting the second predicted video frame data between the first predicted video frame data and the current video frame data, and displaying an image corresponding to the second predicted video frame data on an extended reality device. The method of the present disclosure can solve the problem of discontinuous picture display caused by inaccurate one-time prediction frame insertion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of smart terminal technology, and in particular to an image display method, apparatus, electronic device and storage medium. Background Technology

[0002] Extended Reality (XR) refers to the use of computers to combine the real and virtual worlds, creating a virtual environment that allows for human-computer interaction. In extended reality scenarios, network disturbances often result in uneven time intervals between data packets received by extended reality devices, and the duration of decoded single frames is also uneven. This can lead to inaccurate predictions when the decoded video frame data is sent to asynchronous timewarp (ATW) or asynchronous spacewarp (ASW) modules for video frame prediction. Summary of the Invention

[0003] This summary section is provided to briefly introduce the concepts, which will be described in detail in the detailed description section below. This summary section is not intended to identify key or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0004] This disclosure provides an image display method, apparatus, electronic device, and storage medium.

[0005] The following technical solution is adopted in this disclosure.

[0006] In some embodiments, this disclosure provides an image display method, including:

[0007] Obtain the current video frame data and the first predicted video frame data inserted in the previous period;

[0008] Based on the current video frame data and the first predicted video frame data, image prediction is performed to obtain the second predicted video frame data;

[0009] The second predicted video frame data is inserted between the first predicted video frame data and the current video frame data, and the image corresponding to the second predicted video frame data is displayed on the extended reality device.

[0010] In some embodiments, this disclosure provides an image display device, including:

[0011] The acquisition module is used to acquire the current video frame data and the first predicted video frame data inserted in the previous cycle;

[0012] The first processing module is used to perform image prediction based on the current video frame data and the first predicted video frame data to obtain the second predicted video frame data.

[0013] The second processing module is used to insert the second predicted video frame data between the first predicted video frame data and the current video frame data, and to display the image corresponding to the second predicted video frame data on the extended reality device.

[0014] In some embodiments, this disclosure provides an electronic device, including: at least one memory and at least one processor;

[0015] The memory is used to store program code, and the processor is used to call the program code stored in the memory to execute the above method.

[0016] In some embodiments, this disclosure provides a computer-readable storage medium for storing program code that, when run by a processor, causes the processor to perform the methods described above.

[0017] The image display method provided in this embodiment obtains the current video frame data and the first predicted video frame data inserted in the previous cycle, performs image prediction based on the current video frame data and the first predicted video frame data to obtain the second predicted video frame data, inserts the second predicted video frame data between the first predicted video frame data and the current video frame data, and displays the corresponding image of the second predicted video frame data on the extended reality device. This method can solve the problem of inaccurate ATW or ASW prediction pins caused by network disturbances or differences in decoding duration, thereby compensating for the defect of insufficient image smoothness caused by large differences between the predicted frame and the actual frame. Attached Figure Description

[0018] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.

[0019] Figure 1 This is one of the flowcharts of the image display method according to an embodiment of this disclosure.

[0020] Figure 2 This is a schematic diagram of an extended reality device data stream according to an embodiment of this disclosure.

[0021] Figure 3 This is a schematic diagram of frame interpolation and frame supplementation based on the related technologies provided in the embodiments of this disclosure.

[0022] Figure 4This is a schematic diagram of frame interpolation in an embodiment of this disclosure.

[0023] Figure 5 This is one of the schematic diagrams illustrating the positional relationship between video frames in an embodiment of this disclosure.

[0024] Figure 6 This is the second schematic diagram illustrating the positional relationship between video frames in an embodiment of this disclosure.

[0025] Figure 7 This is a schematic diagram illustrating the effect of video frame prediction in an embodiment of this disclosure.

[0026] Figure 8 This is a schematic diagram illustrating the effect of video frame buffering in an embodiment of this disclosure.

[0027] Figure 9 This is the second flowchart of an image display method according to an embodiment of the present disclosure.

[0028] Figure 10 This is a schematic diagram illustrating the effect of the screen display in an embodiment of this disclosure.

[0029] Figure 11 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present disclosure. Detailed Implementation

[0030] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0031] It should be understood that the various steps described in the method embodiments of this disclosure can be performed in sequence and / or in parallel. Furthermore, method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0032] The term "comprising" and its variations as used herein are open-ended inclusion, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Relevant definitions of other terms will be given in the description below. The term "in response to" and related terms refer to a signal or event being affected to some extent by another signal or event, but not necessarily completely or directly. If event x occurs "in response to" event y, then x may be directly or indirectly responsive to y. For example, the occurrence of y may ultimately lead to the occurrence of x, but there may be other intermediate events and / or conditions. In other cases, y may not necessarily lead to the occurrence of x, and x may occur even if y has not yet occurred. Furthermore, the term "in response to" can also mean "at least partially responsive to".

[0033] The term "determine" broadly encompasses a wide variety of actions, including acquisition, calculation, computation, processing, derivation, investigation, search (e.g., searching in a table, database, or other data structure), discovery, and similar actions; it may also include receiving (e.g., receiving information), accessing (e.g., accessing data in memory), parsing, selecting, choosing, building, and similar actions, etc. Definitions for other terms will be provided below.

[0034] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0035] It should be noted that the use of the word "a" in this disclosure is illustrative rather than restrictive, and those skilled in the art should understand that it should be understood as "one or more" unless otherwise expressly indicated in the context.

[0036] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0037] Asynchronous Timewarp (ATW) is a technique for generating intermediate frames. When a game's frame rate cannot be maintained sufficiently, it generates intermediate frames to compensate, thus maintaining a high refresh rate. For example, the Graphics Processing Unit (GPU) renders the images for the left and right eyes of an extended reality device separately, and then inserts an ATW process before the image is displayed. ATW retrieves the previous frame, combines it with changes in head movement to predict the new frame, and displays it, thus maintaining the frame rate. However, ATW only predicts 3-DOF rotational motion and cannot predict translational motion. Therefore, the main function of ATW is to reduce image jitter, improve efficiency, and maintain low latency.

[0038] Asynchronous Spacewarp (ASW) is an advanced version of ATW that predicts new frames based on previous frames, such as predicting the third frame based on the first two. However, ASW doesn't perform well at refresh rates below half the screen's refresh rate. Additionally, depending on the displayed content, imperfect frame prediction can cause visual artifacts. Typical examples are as follows:

[0039] (1) Rapid changes in brightness. Lightning, oscillating lights, fade-in / fade-out, flashes, and other rapid changes in brightness are difficult for ASW to track. When ASW attempts to identify blocks, these parts of the image may appear shaky. Additionally, some semi-transparent animated scenes may also exhibit a similar effect.

[0040] (2) Object De-occlusion Trail. "De-occlusion" means that an object moves away from occluding a certain area. When an object moves, Assassin's Window (ASW) needs a frame to fill the area it no longer occludes. However, ASW doesn't know what to fill this area with, so the background frame will be stretched to fill it, forming a trail. Since the distance an object moves is usually very small at 45 FPS, this trail is generally not noticeable.

[0041] (3) High-speed movement of repetitive patterns. For example, running in front of a gate and looking at it. Because one part of the object looks similar to the others, it is difficult to judge which way it is moving. This kind of misprediction is very rare for ASW, but it still happens occasionally.

[0042] (4) Head-locked elements move too fast to be tracked. Some applications use head-locked elements (meaning they don't move relative to your head), such as cockpits, head-up displays (HUDs), or menus. If an application tries to handle such elements itself, jitter can occur when the background moves too fast relative to the locked element. ASW can compensate for this to some extent, but sometimes the user's head moves too fast, making it difficult to track, resulting in choppy visuals.

[0043] As can be seen, all of these shortcomings are due to the inaccuracy of ASW prediction. Furthermore, both ATW and ASW predict based on already received frames, then use the generated predicted frames for interpolation and display. Figure 2 As shown, the data processing flow for the split-type streaming scheme of extended reality devices includes: (1) The PC (Personal Computer) renders and encodes the game data and transmits it to the extended reality head-mounted device via the network. (2) The head-mounted device receives network data packets. Due to network disturbances, the time interval of the received data packets is uneven, and even the order of the data packets needs to be adjusted. (3) After receiving the network data packets, the head-mounted device sends the received data to the decoder for decoding. However, the time required to decode each frame of data is different, which is related to the complexity of the current frame and the correlation of the reference frame.

[0044] like Figure 3 As shown, the ASW frame interpolation scheme of related technologies predicts frame C' based on two received data frames, A and B. Each predicted frame is based on an actual frame (not a predicted frame); that is, if there is a predicted frame E', then E' is predicted based on the delayed arrival of frame C and the previously displayed frame D. However, due to network disturbances or decoding time, if frame C is not ready when the on-screen signal arrives, frame C' will be pushed to the on-screen signal for image on-screen processing. If frames C and D arrive consecutively before the next on-screen signal arrives, the ASW scheme typically discards frame C and directly pushes frame D to the on-screen signal for processing. The image data after ASW frame interpolation is sent to the ATW module, which performs predictive processing based on head rotation. Finally, the user can see the processed image on the head-mounted display.

[0045] like Figure 5As shown, the ASW frame interpolation scheme of related technologies predicts the ball's movement position in the image as C', but the actual position of the ball in frame C is position C in the image, and the position of the ball in the next frame D is position D in the image. This is because the ASW prediction algorithm controls its threshold to avoid "overprediction," so the actual displacement of the predicted frame is often smaller than the actual displacement. It is understandable that the ASW scheme predicts the interpolated third frame based on the two already arrived frames, but the actual third frame has a significant positional difference from the predicted third frame. This results in the interpolation effect not guaranteeing smooth movement of objects in the image, making the viewing experience less smooth.

[0046] To address the aforementioned technical problems, the solutions provided by the embodiments of this disclosure will be described in detail below with reference to the accompanying drawings.

[0047] like Figure 1 As shown, Figure 1 This is a flowchart of an image display method according to an embodiment of the present disclosure, which includes the following steps.

[0048] Step S01: Obtain the current video frame data and the first predicted video frame data inserted in the previous cycle;

[0049] In some embodiments, the current video frame data is the currently received video frame data (equivalent to the next video frame data), and the first predicted video frame data inserted in the previous cycle is the predicted frame data predicted by ASW in the previous cycle. For example Figure 2 In this diagram, D represents the current video frame data, and C' represents the predicted video frame data from the previous cycle. C' is obtained by predicting video frames based on the previous two frames A and B.

[0050] Step S02: Perform image prediction based on the current video frame data and the first predicted video frame data to obtain the second predicted video frame data;

[0051] In some embodiments, such as Figure 4 As shown, when the D-frame up-screen signal arrives, the image actually displayed in this embodiment of the present disclosure is not... Figure 5 The first predicted video frame (C'D) is the first predicted video frame displayed on the screen. Specifically, it is based on the first predicted video frame shown on the screen. Figure 4 (frame C' in the middle) and the next frame that has been reached ( Figure 4 The D frames in the image are weighted and bidirectionally predicted to generate C'D frames.

[0052] Step S03: Insert the second predicted video frame data between the first predicted video frame data and the current video frame data, and display the image corresponding to the second predicted video frame data on the extended reality device.

[0053] In some embodiments, such as Figure 7 As shown, the second predicted video frame data C' generated by prediction

[0054] D is to address the problem of a large difference between the predicted C' from the previous cycle and the D frame to be displayed in the next frame. Therefore, in this embodiment of the present disclosure, a bidirectional prediction data frame C'D is inserted between the C' frame and the D frame, and the image corresponding to the second prediction video frame data is displayed on the extended reality device.

[0055] The image display method provided in this embodiment obtains the current video frame data and the first predicted video frame data inserted in the previous cycle, performs image prediction based on the current video frame data and the first predicted video frame data to obtain the second predicted video frame data, inserts the second predicted video frame data between the first predicted video frame data and the current video frame data, and displays the corresponding image of the second predicted video frame data on the extended reality device. This method can solve the problem of inaccurate ATW or ASW prediction pins caused by network disturbances or differences in decoding duration, thereby compensating for the defect of insufficient image smoothness caused by large differences between the predicted frame and the actual frame.

[0056] In some embodiments, the first predicted video frame data is obtained by asynchronous spatial warping of the decoded video frame data.

[0057] In some embodiments, such as Figure 7 As shown, the first predicted video frame data (frame C') is obtained by performing ASW prediction processing on the decoded video frame data (frames A and B).

[0058] In some embodiments, the step of performing image prediction based on the current video frame data and the first predicted video frame data to obtain second predicted video frame data includes:

[0059] The current video frame data and the first predicted video frame data are subjected to bidirectional weighted prediction to obtain the second predicted video frame data.

[0060] In some embodiments, the bidirectional weighted prediction of the current video frame data and the first predicted video frame data includes:

[0061] The current video frame data and the first predicted video frame data are weighted and fused according to the preset weights of the current video frame data and the first predicted video frame data to obtain the second predicted video frame data.

[0062] In some embodiments, the current video frame data and the first predicted video frame data are subjected to bidirectional weighted prediction according to the following formula:

[0063] P pred =w1*p ref1+ w2*p ref2 ;

[0064] Among them, P ref1 P represents the first predicted video frame data inserted in the previous period. ref2 This represents the current video frame data, where W1 represents P. ref1 The preset weights, W2 represents P ref2 The preset weight, P pred This represents the second predicted video frame data, where the weight ratio of W1 and W2 can be freely set according to the actual situation, with a default value of 50%. There are no specific restrictions here, but the constraint condition W1+W2=1 must be met.

[0065] It should be noted that the embodiments of this disclosure solve the problem of discontinuous images caused by inaccurate single-frame prediction through bidirectional prediction and secondary frame interpolation. Furthermore, the secondary frame interpolation of the embodiments of this disclosure is not a simple unidirectional prediction based on single-frame interpolation, but a bidirectional prediction frame interpolation based on a single frame (already displayed on the screen) and the next frame that has arrived.

[0066] In some embodiments of this disclosure, the effect of bidirectional prediction after secondary frame interpolation is as follows: Figure 7 As shown, Figure 7 The small ball outlined by the dashed line represents the second predicted video frame data. Figure 6 In the existing scheme, frame C' is the prediction result of frames A and B, E' is the prediction result of frame C (not displayed on the screen, but used for prediction) and frame D, and so on.

[0067] In some embodiments, after obtaining the second predicted video frame data, the method further includes:

[0068] The current frame data is cached, and the cached current frame data is used to display the corresponding image on the extended reality device in the next cycle.

[0069] In some embodiments, such as Figure 4 As shown, after generating the second predicted video frame data C'D, the current frame D is temporarily stored in the buffer. The buffered D frame data can be used to counteract the effects of subsequent network disturbances or excessive delays caused by decoding.

[0070] In some embodiments, such as Figure 8As shown, frames that fail to be displayed in time or have already arrived can be buffered to counteract missing frame data caused by network disturbances or decoding latency. The latency of one frame (11ms at 90Hz) is within the acceptable range, and with appropriate buffering within permissible latency limits, even under network disturbances, it can effectively solve problems such as missing frames and insufficient frame rate.

[0071] In some embodiments, displaying the image corresponding to the second predicted video frame data on the extended reality device includes:

[0072] The frame sequence into which the second predicted video frame data is inserted is sent to the display screen of the extended display device so that the display screen displays the corresponding image according to the frame sequence.

[0073] In some embodiments, the first predicted video frame data is a video frame displayed on the extended display device.

[0074] In some embodiments, this disclosure also provides an image display method, such as... Figure 9 As shown, it includes:

[0075] Step (a): The PC receives 6DoF sensor data from the headset and controllers and renders video images in the game engine;

[0076] In some embodiments, the split-type extended reality streaming solution sends 6DoF sensor data from the headset and controllers to a PC, where the game engine renders game images based on this 6DoF data.

[0077] Step (b): The PC sends the encoded video image data to the headset;

[0078] In some embodiments, the PC encodes and packages the rendered video image data, and then sends the data packets to the headset via wireless or wired transmission. Prior to this processing step, as long as the PC's encoder (GPU performance) is sufficiently powerful, it can ensure that the video frame data is encoded, packaged, and sent at stable frame rate intervals.

[0079] Step (c): The head-mounted device receives the rendered and encoded video data;

[0080] In some embodiments, due to network disturbances, if a wireless transmission method is used, the time interval between data packets received at the head-mounted device will be uneven. Furthermore, if the transmission method is TCP (Transmission Control Protocol), if packet loss and retransmission occur, the order of the received data packets will also be adjusted (TCP is a reliable transmission and will not lose packets; if packets are lost, they will be retransmitted, but the arrival time will be later). Therefore, when the head-mounted device receives data packets, the time interval between each frame is uneven.

[0081] Step (d): The decoder performs video decoding;

[0082] In some embodiments, the head-mounted device pushes each received video frame data to the decoder for decoding. Currently, taking a 1080p resolution video frame as an example, its decoding time is in the range of 3ms to 5ms. The decoding time of each frame is related to the image complexity and the correlation with the reference frame.

[0083] Step (e): After decoding, the data is pushed to ASW for frame interpolation and frame supplementation. Based on the two video frames that have been received in the previous two frames, the next frame to be displayed is predicted.

[0084] In some embodiments, frame interpolation of predicted frames is achieved using ASW (Automatic Frame Layout). Specifically, image motion vector prediction or optical flow prediction is performed based on the two already arrived frames. That is, given two known frames, a third frame is predicted forward, and the resulting predicted frame is inserted into the frame sequence as the third frame data and pushed to the backend for display.

[0085] Step (f): Perform bidirectional prediction interpolation based on the predicted frame data and the next frame data, that is, insert another frame between the predicted frame and the next frame, and buffer the next frame;

[0086] In some embodiments, such as Figure 4 As shown, the D-frame data originally intended for display in the current cycle needs to be buffered and temporarily stored in the current cycle, and then displayed in the next cycle. If the decoding end does not deliver the frame data in time in the next cycle (which usually results in missing frames or insufficient frame rate), the buffered data will be used for display.

[0087] In some embodiments, since this step corresponds to the method embodiment described above, the relevant details can be found in the description of the method embodiment, and will not be repeated here.

[0088] In step (g), the data after frame interpolation is sent to the screen for display according to the on-screen signal.

[0089] In some embodiments, the image sequence after frame interpolation according to the present disclosure can solve the problems of uneven image motion and large differences between the predicted position of the predicted frame C' and the actual C frame in related technologies. Furthermore, the present disclosure can also compensate for insufficient frame rate to a greater extent, that is, by using cached D frames, sacrificing the "latency" of the current frame to achieve this function. However, the "latency" is only for the current frame; throughout the entire processing flow, the cached frame data can effectively combat insufficient frame rate in response to network disturbances.

[0090] As can be seen, the bidirectional predictive frame interpolation implemented in this embodiment can not only compensate for the problem of uneven image smoothness caused by large differences between predicted frames and actual frames, but also ensure that the frame rate will not decrease due to missing frames to a greater extent within the allowable latency range. Notably, the cached frame data is not delayed by a fixed one-frame delay, but is only cached in the case of bidirectional prediction.

[0091] from Figure 10 As can be seen, the relevant technology cannot continuously acquire two predicted frames from two existing frames, which leads to missing frames during the display process, meaning the actual frame rate cannot reach the preset effect. Furthermore, because there is a certain deviation between the position of objects in frame C' and the actual frame C, the playback effect of the entire image sequence often results in a less than smooth visual experience for the user. The image display method provided in this disclosure can satisfy the smoothness requirement while also mitigating the insufficient frame rate to a greater extent. Moreover, in this disclosure, the temporary frame buffer is not a fixed n (≥1) frame buffer, allowing for a more flexible and smoother approach without introducing a fixed image display delay. This disclosure plays a crucial role in improving the overall data streaming effect of the extended reality split-type machine.

[0092] This disclosure also provides an image display device, including:

[0093] The acquisition module is used to acquire the current video frame data and the first predicted video frame data inserted in the previous cycle;

[0094] The first processing module is used to perform image prediction based on the current video frame data and the first predicted video frame data to obtain the second predicted video frame data.

[0095] The second processing module is used to insert the second predicted video frame data between the first predicted video frame data and the current video frame data, and to display the image corresponding to the second predicted video frame data on the extended reality device.

[0096] In some embodiments, the first predicted video frame data is obtained by asynchronous spatial warping of the decoded video frame data.

[0097] In some embodiments, the first processing module is specifically used for:

[0098] The current video frame data and the first predicted video frame data are subjected to bidirectional weighted prediction to obtain the second predicted video frame data.

[0099] In some embodiments, the first processing module is further specifically used for:

[0100] The current video frame data and the first predicted video frame data are weighted and fused according to the preset weights of the current video frame data and the preset weights of the first predicted video frame data to obtain the second predicted video frame data.

[0101] In some embodiments, the first processing module is further specifically used for:

[0102] The current frame data is cached, and the cached current frame data is used to display the corresponding image on the extended reality device in the next cycle.

[0103] In some embodiments, the second processing module is specifically used for:

[0104] The frame sequence into which the second predicted video frame data is inserted is sent to the display screen of the extended display device so that the display screen displays the corresponding image according to the frame sequence.

[0105] In some embodiments, the first predicted video frame data is a video frame displayed on the extended display device.

[0106] For embodiments of the apparatus, since they basically correspond to the method embodiments, relevant details can be found in the descriptions of the method embodiments. The apparatus embodiments described above are merely illustrative, and the modules described as separate modules may or may not be separate. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0107] The methods and apparatus of this disclosure have been described above based on embodiments and application examples. Furthermore, this disclosure also provides an electronic device and a computer-readable storage medium, which are described below.

[0108] The following is for reference. Figure 11The figure illustrates a structural schematic of an electronic device (e.g., a terminal device or server) 800 suitable for implementing embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. The electronic device shown in the figure is merely an example and should not be construed as limiting the functionality and scope of the embodiments of the present disclosure.

[0109] Electronic device 800 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 801, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 802 or a program loaded from storage device 808 into random access memory (RAM) 803. RAM 803 also stores various programs and data required for the operation of electronic device 800. The processing device 801, ROM 802, and RAM 803 are interconnected via bus 804. Input / output (I / O) interface 805 is also connected to bus 804.

[0110] Typically, the following devices can be connected to I / O interface 805: input devices 806 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 807 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 808 including, for example, magnetic tapes, hard disks, etc.; and communication devices 809. Communication device 809 allows electronic device 800 to communicate wirelessly or wiredly with other devices to exchange data. Although an electronic device 800 with various devices is shown in the figure, it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0111] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 809, or installed from a storage device 808, or installed from a ROM 802. When the computer program is executed by a processing device 801, it performs the functions defined in the methods of embodiments of this disclosure.

[0112] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0113] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0114] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0115] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the methods of the present disclosure.

[0116] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0117] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0118] The units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units are not, in some cases, intended to limit the specific unit.

[0119] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0120] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0121] According to one or more embodiments of this disclosure, an image display method is provided, comprising:

[0122] Obtain the current video frame data and the first predicted video frame data inserted in the previous period;

[0123] Based on the current video frame data and the first predicted video frame data, image prediction is performed to obtain the second predicted video frame data;

[0124] The second predicted video frame data is inserted between the first predicted video frame data and the current video frame data, and the image corresponding to the second predicted video frame data is displayed on the extended reality device.

[0125] According to one or more embodiments of this disclosure, a method is provided in which the first predicted video frame data is obtained by asynchronously spatially warping decoded video frame data.

[0126] According to one or more embodiments of this disclosure, a method is provided in which image prediction is performed based on the current video frame data and the first predicted video frame data to obtain second predicted video frame data, comprising:

[0127] The current video frame data and the first predicted video frame data are subjected to bidirectional weighted prediction to obtain the second predicted video frame data.

[0128] According to one or more embodiments of this disclosure, a method is provided in which bidirectional weighted prediction is performed on the current video frame data and the first predicted video frame data, comprising:

[0129] The current video frame data and the first predicted video frame data are weighted and fused according to the preset weights of the current video frame data and the first predicted video frame data to obtain the second predicted video frame data.

[0130] According to one or more embodiments of this disclosure, a method is provided that, after obtaining the second predicted video frame data, further includes:

[0131] The current frame data is cached, and the cached current frame data is used to display the corresponding image on the extended reality device in the next cycle.

[0132] According to one or more embodiments of this disclosure, a method is provided for displaying an image corresponding to the second predicted video frame data on an extended reality device, comprising:

[0133] The frame sequence into which the second predicted video frame data is inserted is sent to the display screen of the extended display device so that the display screen displays the corresponding image according to the frame sequence.

[0134] According to one or more embodiments of this disclosure, a method is provided in which the first predicted video frame data is a video frame to be displayed on the extended display device.

[0135] According to one or more embodiments of the present disclosure, an image display device is provided, comprising:

[0136] The acquisition module is used to acquire the current video frame data and the first predicted video frame data inserted in the previous cycle;

[0137] The first processing module is used to perform image prediction based on the current video frame data and the first predicted video frame data to obtain the second predicted video frame data.

[0138] The second processing module is used to insert the second predicted video frame data between the first predicted video frame data and the current video frame data, and to display the image corresponding to the second predicted video frame data on the extended reality device.

[0139] According to one or more embodiments of the present disclosure, an electronic device is provided, including: at least one memory and at least one processor;

[0140] The at least one memory is used to store program code, and the at least one processor is used to call the program code stored in the at least one memory to execute the method described in any one of the above.

[0141] According to one or more embodiments of the present disclosure, a computer-readable storage medium is provided for storing program code that, when executed by a processor, causes the processor to perform the methods described above.

[0142] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0143] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0144] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.

Claims

1. An image display method, characterized in that, include: Obtain the current video frame data and the first predicted video frame data inserted in the previous cycle, wherein the first predicted video frame data inserted in the previous cycle is the predicted frame data predicted in the previous cycle through asynchronous spatial warp processing; Based on the current video frame data and the first predicted video frame data, image prediction is performed to obtain the second predicted video frame data; The second predicted video frame data is inserted between the first predicted video frame data and the current video frame data, and the image corresponding to the second predicted video frame data is displayed on the extended reality device.

2. The method according to claim 1, characterized in that, The step of performing image prediction based on the current video frame data and the first predicted video frame data to obtain the second predicted video frame data includes: The current video frame data and the first predicted video frame data are subjected to bidirectional weighted prediction to obtain the second predicted video frame data.

3. The method according to claim 2, characterized in that, The step of performing bidirectional weighted prediction on the current video frame data and the first predicted video frame data includes: The current video frame data and the first predicted video frame data are weighted and fused according to the preset weights of the current video frame data and the first predicted video frame data to obtain the second predicted video frame data.

4. The method according to claim 1, characterized in that, After obtaining the second predicted video frame data, the method further includes: The current video frame data is cached, and the cached current video frame data is used to display the corresponding image on the extended reality device in the next cycle.

5. The method according to claim 1, characterized in that, The step of displaying the image corresponding to the second predicted video frame data on the extended reality device includes: The frame sequence into which the second predicted video frame data is inserted is sent to the display screen of the extended reality device so that the display screen displays the corresponding image according to the frame sequence.

6. The method according to claim 1, characterized in that, The first predicted video frame data is a video frame that is displayed on the extended reality device.

7. An image display device, characterized in that, include: The acquisition module is used to acquire the current video frame data and the first predicted video frame data inserted in the previous cycle, wherein the first predicted video frame data inserted in the previous cycle is the predicted frame data predicted in the previous cycle through asynchronous spatial warp processing. The first processing module is used to perform image prediction based on the current video frame data and the first predicted video frame data to obtain the second predicted video frame data. The second processing module is used to insert the second predicted video frame data between the first predicted video frame data and the current video frame data, and to display the image corresponding to the second predicted video frame data on the extended reality device.

8. An electronic device, comprising: At least one memory and at least one processor; The at least one memory is used to store program code, and the at least one processor is used to call the program code stored in the at least one memory to execute the method of any one of claims 1 to 6.

9. A computer-readable storage medium for storing program code, which, when executed by a computer device, causes the computer device to perform the method of any one of claims 1 to 6.

Citation Information

Patent Citations

  • Method and device for reducing virtual reality latency

    CN106658170A