Decoding region of interest first in computer game video and hiding missing portions simultaneously using multiple decoders

By prioritizing the transmission of ROI of computer game videos through a multi-encoder and decoder system and using machine learning models to handle missing parts, the problem of balancing latency and quality under limited network bandwidth is solved, achieving low-latency, high-quality game video transmission.

CN121925833APending Publication Date: 2026-04-24SONY INTERACTIVE ENTERTAINMENT LLC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SONY INTERACTIVE ENTERTAINMENT LLC
Filing Date
2024-10-03
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

With limited network bandwidth, the latency of computer game videos makes it difficult to meet the real-time reaction requirements of gamers, and existing technologies cannot effectively solve the balance between video transmission latency and quality.

Method used

A multi-encoder and multi-decoder system is used to prioritize the transmission of regions of interest (ROI) in computer game videos, and machine learning models are used at the receiving end to reconstruct or fuse missing parts in order to reduce latency and improve video quality.

Benefits of technology

By prioritizing the transmission of ROI portions and utilizing machine learning models to process non-ROI portions, latency is significantly reduced, improving the response speed and quality of game videos and adapting to unstable network environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121925833A_ABST
    Figure CN121925833A_ABST
Patent Text Reader

Abstract

Techniques are described for reducing latency in computer game network streaming through the use of multiple encoders and decoders, where one encoder-decoder pair (304 / 400) is used for a region of interest (ROI) in video and given priority in transmission and rendering higher than background video processed by the other encoder / decoder pair.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to technically inventive and unconventional solutions that are necessarily rooted in computer technology and produce specific technical improvements, and more specifically to first decoding regions of interest in computer game videos and simultaneously using multiple decoders to hide missing parts. Background Technology

[0002] Videos, such as computer simulation videos and computer game videos, can be streamed to end-user terminals over a network. Summary of the Invention

[0003] As understood in this article, network conditions and / or regulatory restrictions on bandwidth for energy conservation may limit network channels available for transmitting video, such as computer game videos. As further understood in this article, latency is a primary concern in such cases, particularly for computer gamers, who prefer near-instantaneous responses to their inputs, such as when shooting weapons in a game. Therefore, video quality may be less of a concern compared to transmitting video with little or no latency.

[0004] Therefore, a device includes at least one processor component configured to encode a first portion of a video using a first encoder; encode a second portion of the video using a second encoder; and transmit the first portion to at least one receiver via a network before transmitting the second portion, such that the first portion is preferentially transmitted via the network.

[0005] In an example implementation, the video may include computer game videos.

[0006] In some implementations, the first part may include a region of interest (ROI). The ROI may be identified at the receiver using gaze tracking, and / or by the source of the video, and / or by a machine learning (ML) model. If desired, the processor component may be configured to send an indication to the receiver of any unencoded portions of the video.

[0007] In another aspect, an apparatus includes at least one processor component configured to decode a first portion of a video using a first decoder and to decode a second portion of the video using a second decoder. The processor component is configured to present the first portion and the second portion together on at least one video display in response to the second portion being available for display. Furthermore, the processor component is configured to present the first portion on at least one video display without presenting the second portion in response to the second portion being unavailable for display during a delay period.

[0008] In some examples, the processor component may be configured to reuse at least a portion of a previous frame to replace the second portion in response to the second portion being unavailable for display during the delay period. In other examples, the processor component may be configured to blur portions of the video surrounding the first portion in response to the second portion being unavailable for display during the delay period. Furthermore, the processor component may be configured to reconstruct the second portion in response to the second portion being unavailable for display.

[0009] On the other hand, one method includes transmitting the ROI portion of a video frame to a receiver before transmitting the portion of the video frame located outside the ROI portion, and presenting the ROI portion on a video display, regardless of whether the portion of the frame located outside the ROI region is decoded for presentation.

[0010] The details of this disclosure regarding both its structure and operation can be best understood with reference to the accompanying drawings, in which the same reference numerals refer to the same parts, and in the accompanying drawings: Attached Figure Description

[0011] Figure 1 This is a block diagram of an example system that includes examples consistent with the principles of the present invention; Figure 2 An example encoder-decoder system is shown; Figure 3 Alternative example encoding systems are shown; Figure 4 An alternative example decoding system is shown; Figure 5 The example transmitter / encoder logic is shown in the example flowchart format; Figure 6 The example receiver / decoder logic is illustrated in an example flowchart format; Figure 7 Example logic for training a machine learning (ML) model to identify regions of interest (ROI) in a video is shown in the example flowchart format. Figure 8 Additional example logic for signaling missing frame portions is shown in an example flowchart format; Figure 9 Additional example logic for signaling missing frame portions is shown in an example flowchart format; Figure 10 Example logic for training an ML model to reconstruct missing parts of a video is shown in an example flowchart format; and Figure 11 An example flowchart is shown to illustrate sample logic for merging the ROI and non-ROI portions of a video. Detailed Implementation

[0012] This disclosure generally relates to a computer ecosystem that includes various aspects of consumer electronics (CE) device networks, such as, but not limited to, computer gaming networks. Systems described herein may include server components and client components that can be networked to enable the exchange of data between the client components and the server components. Client components may include one or more computing devices, including game consoles (such as Sony PlayStation® or game consoles made by Microsoft, Nintendo, or other manufacturers), extended reality (XR) headsets (such as virtual reality (VR) headsets, augmented reality (AR) headsets), portable televisions (e.g., smart TVs, internet-enabled TVs), portable computers (such as laptops and tablets), and other mobile devices (including smartphones and additional examples discussed below). These client devices may operate in a variety of operating environments. For example, some client computers may use operating systems such as Linux, operating systems from Microsoft, or Unix, or operating systems produced by Apple, Inc., Google, Berkeley Software Distribution, or Berkeley Standard Distribution (BSD) OS (including descendants of BSD). These operating environments can be used to execute one or more browsing programs, such as browsers made by Microsoft, Google, or Mozilla, or other browser programs that can access websites hosted on internet servers discussed below. Furthermore, the operating environment according to the principles of the present invention can be used to execute one or more computer game programs.

[0013] A server and / or gateway may be used, which may include one or more processors that execute instructions to configure the server to receive and transmit data over a network such as the Internet. Alternatively, the client and server may connect via a local intranet or virtual private network. The server or controller may be instantiated from a game console such as a Sony PlayStation®, a personal computer, etc.

[0014] Information can be exchanged between clients and servers over a network. For this purpose, and for security reasons, servers and / or clients may include firewalls, load balancers, temporary storage and proxies, as well as other network infrastructure for reliability and security. One or more servers can form a device that enables methods for providing secure communities (such as online social networking sites or gaming networks) to network members.

[0015] A processor can be a single-chip or multi-chip processor, which can execute logic using various lines (such as address lines, data lines, and control lines) as well as registers and shift registers. A processor, including a digital signal processor (DSP), can be an implementation of a circuit system. A processor component can include one or more processors.

[0016] Components included in one embodiment can be used in any suitable combination in other embodiments. For example, any of the various components depicted herein and / or in the accompanying drawings can be combined, interchanged, or excluded from other embodiments.

[0017] "A system having at least one of A, B and C" (similarly, "a system having at least one of A, B or C" and "a system having at least one of A, B and C") includes: a system having only A; a system having only B; a system having only C; a system having both A and B; a system having both A and C; a system having both B and C; and / or a system having both A, B and C.

[0018] Now for reference Figure 1 An example system 10 is illustrated, which may include one or more of the example devices mentioned above and further described below according to the principles of the invention. The first device among the example devices included in system 10 is a consumer electronics (CE) device, such as an audio-visual device (AVD) 12, such as, but not limited to, a cinema display system (which may be projector-based) or an internet-enabled TV with a TV tuner (equivalently, a set-top box controlling a TV). Alternatively, the AVD 12 may also be a computerized internet-enabled (“smart”) phone, tablet computer, laptop computer, head-mounted device (HMD) and / or head-mounted equipment (such as smart glasses or VR headsets), another wearable computerized device, a computerized internet-enabled music player, a computerized internet-enabled headset, a computerized internet-enabled implantable device (such as an implantable skin device), etc. In any case, it should be understood that the AVD 12 is configured to implement the principles of the invention (e.g., to communicate with other CE devices to implement the principles of the invention, to perform the logic described herein, and to perform any other functions and / or operations described herein).

[0019] Therefore, to implement such principles, the AVD 12 can be constructed using some or all of the components shown. For example, the AVD 12 may include one or more touch-enabled displays 14, which may be implemented using a high-definition or ultra-high-definition "4K" or higher resolution flat panel screen. The touch-enabled display 14 may include, for example, a capacitive or resistive touch sensing layer with an electrode grid for touch sensing, consistent with the principles of the present invention.

[0020] AVD 12 may also include: one or more speakers 16 for outputting audio according to the principles of the invention; and at least one additional input device 18 (such as an audio receiver / microphone) for inputting audible commands to control AVD 12. Example AVD 12 may also include one or more network interfaces 20 for communicating over at least one network 22 (such as the Internet, WAN, LAN, etc.) under the control of one or more processors 24. Therefore, interface 20 may be, but is not limited to, a Wi-Fi transceiver, which is an example of a wireless computer network interface, such as, but not limited to, a mesh network transceiver. It should be understood that processor 24 controls AVD 12 to implement the principles of the invention, including controlling other elements of AVD 12 described herein, such as controlling display 14 to present images on the display and receiving input from the display. Furthermore, it should be noted that network interface 20 may be a wired or wireless modem or router, or other suitable interface, such as a wireless telephone transceiver or a Wi-Fi transceiver as mentioned above.

[0021] In addition to the foregoing, AVD 12 may also include one or more input and / or output ports 26, such as a High Definition Multimedia Interface (HDMI) port or a Universal Serial Bus (USB) port for physically connecting to another CE device and / or a headphone port for connecting headphones to AVD 12 to present audio from AVD 12 to a user via headphones. For example, input port 26 may be connected via a cable or satellite source 26a to audio / video content, either wired or wirelessly. Therefore, source 26a may be a separate or integrated set-top box or satellite receiver. Alternatively, source 26a may be a game console or disk player containing content. When implemented as a game console, source 26a may include some or all of the components described below with respect to CE device 48.

[0022] AVD 12 may also include one or more computer memory / computer-readable storage media 28 that are not transient signals, such as disk-based storage or solid-state storage. In some cases, the one or more computer memory / computer-readable storage media may be embodied as a stand-alone device within the AVD's housing, or as a personal video recording device (PVR) or video disk player for playing back AV programs, either inside or outside the AVD's housing, or as removable storage media or a server as described below. Furthermore, in some embodiments, AVD 12 may include a location or positioning receiver, such as, but not limited to, a mobile phone receiver, a GPS receiver, and / or an altimeter 30, configured to receive geographic location information from a satellite or mobile phone base station and provide this information to processor 24 and / or, in conjunction with processor 24, determine the altitude at which AVD 12 is positioned.

[0023] Continuing the description of AVD 12, in some embodiments, AVD 12 may include one or more cameras 32, which may be a thermal imaging camera, a digital camera (such as a webcam), an IR sensor, an event-based sensor, and / or a camera integrated into AVD 12 and controllable by processor 24 to acquire pictures / images and / or videos according to the principles of the present invention. AVD 12 may also include a Bluetooth® transceiver 34 and other near-field communication (NFC) elements 36 for communicating with other devices using Bluetooth and / or NFC technologies, respectively. An example NFC element may be a radio frequency identification (RFID) element.

[0024] Furthermore, the AVD 12 may include one or more auxiliary sensors 38 that provide input to the processor 24. For example, one or more of the auxiliary sensors 38 may include one or more pressure sensors forming a layer of the touch-enabled display 14 itself, and may be, but are not limited to, piezoelectric pressure sensors, capacitive pressure sensors, piezoresistive strain gauges, optical pressure sensors, electromagnetic pressure sensors, etc. Other sensor examples include pressure sensors, motion sensors (such as accelerometers, gyroscopes, odometers, or magnetic sensors), infrared (IR) sensors, optical sensors, speed and / or rhythm sensors, event-based sensors, and gesture sensors (e.g., for sensing gesture commands). Thus, the sensor 38 may be implemented by one or more motion sensors, such as individual accelerometers, gyroscopes, and magnetometers and / or inertial measurement units (IMUs), which typically include a combination of accelerometers, gyroscopes, and magnetometers to determine the position and orientation of the AVD 12 in three dimensions, or by event-based sensors (such as event detection sensors (EDS)). Consistent with this disclosure, the EDS provides an output indicating a change in light intensity sensed by at least one pixel of the light sensing array. For example, if the light sensed by a pixel is decreasing, the EDS output can be -1; if the light sensed by a pixel is increasing, the EDS output can be +1. Outputting a binary signal of 0 can indicate that there is no light intensity change below a certain threshold.

[0025] The AVD 12 may also include an over-the-air (OTA) TV broadcast port 40 that provides input to the processor 24 for receiving OTA TV broadcasts. In addition to the foregoing, it should be noted that the AVD 12 may also include an infrared (IR) transmitter and / or an IR receiver and / or an IR transceiver 42, such as an IR data association (IRDA) device. A battery (not shown) may be provided to power the AVD 12, such as a kinetic energy harvester that converts kinetic energy into electrical energy to charge the battery and / or power the AVD 12. A graphics processing unit (GPU) 44 and a field-programmable gate array (FPGA) 46 may also be included. One or more tactile / vibration generators 47 may be provided to generate tactile signals that can be sensed by a person holding or touching the device. Therefore, the haptic generator 47 can use an electric motor to vibrate all or part of the AVD 12, the electric motor being connected via a rotatable shaft to an eccentric and / or unbalanced counterweight, such that the shaft can be rotated under the control of the motor (which can in turn be controlled by a processor such as processor 24) to generate vibrations of various frequencies and / or amplitudes as well as force simulations in various directions.

[0026] It may also include light sources, such as projectors, such as infrared (IR) projectors.

[0027] In addition to AVD 12, System 10 may also include one or more other CE device types. In one example, the first CE device 48 may be a computer game console that can be used to send computer game audio and video to AVD 12 via commands sent directly to AVD 12 and / or via a server described below, while the second CE device 50 may include components similar to the first CE device 48. In the example shown, the second CE device 50 may be configured as a computer game controller operated by a player or a head-mounted display (HMD) worn by a player. The HMD may include a head-up transparent or opaque display for presenting AR / MR content or VR content (more generally, extended reality (XR) content), respectively. The HMD may be configured as a glasses-type display or as a larger VR-type display sold by computer game equipment manufacturers.

[0028] In the example shown, only two CE devices are illustrated; it should be understood that fewer or more devices may be used. The devices described herein can implement some or all of the components shown for AVD 12. Any component shown in the following figures may be combined with some or all of the components shown in the case of AVD 12.

[0029] Referring now to the aforementioned at least one server 52, which includes at least one server processor 54, at least one tangible computer-readable storage medium 56 (such as disk-based storage or solid-state storage), and at least one network interface 58, which, under the control of the server processor 54, allows communication with other illustrated devices via network 22 and, in practice, facilitates communication between the server and client devices according to the principles of the invention. It should be noted that the network interface 58 may be, for example, a wired or wireless modem or router, a Wi-Fi transceiver, or other suitable interface (such as, for example, a wireless telephone transceiver).

[0030] Therefore, in some implementations, server 52 may be an internet server or an entire server "farm," and in an exemplary implementation for, for example, a network gaming application, the server may include and perform "cloud" functionality, enabling devices of system 10 to access a "cloud" environment via server 52. Alternatively, server 52 may be implemented by one or more game consoles or other computers in the same room or nearby as the other devices shown.

[0031] The components shown in the diagram below may include some or all of the components shown in this document. Any user interface (UI) described herein may be combined and / or extended, and UI elements may be mixed and matched between UIs.

[0032] The principles of this invention can be applied to various machine learning models, including deep learning models. Machine learning models consistent with the principles of this invention can use various algorithms trained in ways including: supervised learning, unsupervised learning, semi-supervised learning, reinforcement learning, feature learning, self-learning, and other forms of learning. Examples of such algorithms that can be implemented by computer circuit systems include one or more neural networks, such as convolutional neural networks (CNNs), recurrent neural networks (RNNs), and a type of RNN called a long short-term memory (LSTM) network. Generative pre-trained Transformers (GPTT) can also be used. Support vector machines (SVMs) and Bayesian networks can also be considered examples of machine learning models. In addition to the network types mentioned above, the model in this paper can also be implemented using a classifier.

[0033] As understood in this paper, performing machine learning can therefore involve accessing training data and then training a model on that training data so that the model can process additional data to make inferences. Thus, an artificial neural network / AI model trained via machine learning can include an input layer, an output layer, and multiple hidden layers located between them, which are configured and weighted to make inferences about the appropriate output.

[0034] Figure 2 A system including a video encoder 200 for encoding / compressing video 202 is shown. A video decoder 204 can receive encoded video and decode / decompress it into output video 206.

[0035] Figure 3 Another example encoder / transmitter system is shown in which video, such as computer-simulated video (such as computer game video), is received from a video source 300 (such as a game server or console). A Region of Interest (ROI) identifier engine 302, which may include one or more processors, uses various techniques described herein to identify ROIs in the video, including injecting information from a computer game engine 303A and / or using gaze tracking information 303B from a receiver (e.g., from a head-mounted display camera or other gaze-tracking camera).

[0036] The Region of Interest (ROI) portion of each video frame is encoded by a first encoder 304, while the rest of the frame outside the ROI is encoded by one or more other encoders 306. Prioritizing the transmission of portions encoded by the one or more other encoders 306, such as by prioritizing the transmission of the ROI of the frame over the background portion, the encoded ROI portion is preferentially transmitted by the transmitter 308 (e.g., via a computer network) to one or more receivers. Note that the non-ROI portions of the frame can be encoded at a lower bit rate and / or frame rate and / or lower resolution than the rate / resolution at which the ROI is encoded.

[0037] Figure 4 A mirror receiver-side decoding system is illustrated, comprising an ROI decoder 400 for decoding the Region of Interest (ROI) portion of each video frame and one or more other decoders 402 for decoding the portions of the frame outside the ROI. Each decoder may include its own independent clock, and the decoders may use direct memory access (DMA) to output their signals to a frame buffer / priority engine 404, which may include storage and processing capabilities. The outputs of the frame buffers are presented on one or more video displays 406. Note that decoders 400, 402 may be configured as decoders with independent cores 400, 402, each having its own clock to receive the stream and output to a single frame buffer.

[0038] Figure 5 The transmitter-side logic is shown. Starting at state 500, the ROI portion of the frame is sent to the receiver before the non-ROI portion is transmitted, and the non-ROI portion is subsequently transmitted at state 502.

[0039] Figure 6 The receiver logic, complementary to the transmitter logic, is shown. Starting at state 600, the ROI portion of the received frame is decoded. State 602 indicates that if the non-ROI portion of the frame is not yet available, in order to minimize latency, the non-ROI portion of the frame is either blurred / hidden, or the non-ROI portion of the previous frame is used with the current ROI and updated using motion vectors from the previous non-ROI portion, or at state 604, the non-ROI portion is reconstructed by the ML model.

[0040] On the other hand, if the non-ROI portion of the frame is available, the logic can move from state 602 to state 606 to determine if there is sufficient time to decode the non-ROI portion to render it together with the ROI portion within the delay constraint. That is, even if the non-ROI portion is available for decoding, it may not be rendered with the ROI of the frame if the predetermined delay requirement cannot be met; in this case, the logic flows from state 606 to state 602. However, if there is sufficient time to decode the non-ROI portion to render it together with the ROI portion within the delay constraint, the non-ROI portion is decoded at state 608 and the video frame is rendered at state 610. Furthermore, when the non-ROI portion of the frame cannot be rendered in time with the ROI of the frame to meet the delay requirement, the logic flows from state 604 to the rendering state 610.

[0041] Figure 7 The training logic for the ML model, which can be used in this paper to identify the ROIs of video frames, is shown. Box 700 indicates that the training dataset is fed into the ML model to train the model at box 702. The training set may include video frames and ground truth indices of the ROIs in those frames.

[0042] Figure 8 The receiver / decoder can signal to the transmitter at state 800 that any ROI and / or non-ROI data is missing during reception. At state 802, the transmitter can respond by retransmitting the missing data and / or indicating how to reconstruct the missing data.

[0043] In comparison, Figure 9 This illustrates that the transmitter / encoder can signal to the receiver at state 900 that a portion of a video frame (such as a non-ROI portion) has not yet been transmitted. This signal can accompany the ROI portion of the frame, allowing the decoder to reconstruct the missing non-ROI portion at state 902 to be presented along with the ROI.

[0044] Figure 9 Can be used according to Figure 10 This is achieved using a trained ML model. At state 1000, the training dataset is fed into the ML model to train it at state 1002. The training set may include partial video frames with ground truth indicators of missing parts of the frames.

[0045] Figure 11A technique is illustrated in which the Region of Interest (ROI) portion of a video frame can be stitched to the corresponding non-ROI portion of the frame, such that the non-ROI portion can contain a small portion of the ROI for better fusion. Starting at state 1100, the ROI of the video frame is received. Moving to state 1102, the non-ROI portion of the frame is received. Then, at state 1104, the ROI is fused with the non-ROI portion by placing some portions of the ROI within the non-ROI portion.

[0046] Note that in implementations that decode the entire frame using only a single decoder core, if the non-ROI portion of the frame arrives too late within the latency constraint, only the ROI is rendered, and the non-ROI portion can be discarded. Similarly, if gaze-based focus is used to determine the ROI, if the non-ROI portion of the frame does not arrive at the receiver within the latency constraint, the non-ROI portion can be discarded upon its arrival, and the ROI is rendered within the latency constraint. If the ROI is gaze-based, the receiver / decoder signals to the transmitter / encoder what the ROI is.

[0047] This principle is advantageous for reducing latency and handling unstable network conditions.

[0048] While specific techniques are shown and described in detail herein, it should be understood that the subject matter contained herein is limited only by the claims.

Claims

1. An apparatus comprising: At least one processor component, said at least one processor component being configured to: The first portion of the video is encoded using the first encoder; The second portion of the video is encoded using a second encoder; and Before sending the second part, the first part is sent to at least one receiver via the network, such that the first part is preferentially transmitted via the network.

2. The device of claim 1, wherein the video includes computer game videos.

3. The device of claim 1, wherein the first portion includes a region of interest (ROI).

4. The device of claim 3, wherein the ROI is identified at the receiver using gaze tracking.

5. The device of claim 3, wherein the ROI is identified by the source of the video.

6. The device of claim 3, wherein the ROI is identified by a machine learning (ML) model.

7. The device of claim 1, wherein the processor component is configured to send an indication to the receiver of a portion of the video that was not encoded.

8. An apparatus comprising: At least one processor component, said at least one processor component being configured to: The first part of the video is decoded using the first decoder; The second part of the video is decoded using a second decoder; In response to the second portion being usable for display, the first portion and the second portion are presented together on at least one video display; and In response to the second portion being unavailable for display during the delay period, the first portion is presented on at least one video display without the second portion.

9. The device of claim 8, wherein the video includes computer game video.

10. The device of claim 8, wherein the first portion includes a region of interest (ROI).

11. The device of claim 10, wherein the ROI is identified using gaze tracking.

12. The device of claim 10, wherein the ROI is identified by the source of the video.

13. The device of claim 10, wherein the ROI is identified by a machine learning (ML) model.

14. The device of claim 8, wherein the processor component is configured to reuse at least a portion of a previous frame in place of the second portion in response to the second portion being unavailable for display during a delay period.

15. The device of claim 8, wherein the processor component is configured to partially blur the video surrounding the first portion in response to the second portion being unavailable for display during the delay period.

16. The device of claim 8, wherein the processor component is configured to reconstruct the second portion in response to the second portion being unusable for display.

17. A method comprising: The ROI portion of the video frame is transmitted to the receiver before the portion of the video frame located outside the ROI portion is transmitted; as well as The ROI portion is presented on the video display, regardless of whether the portion of the frame located outside the ROI region is decoded for presentation.

18. The method of claim 17, further comprising using a plurality of encoders to encode the frame.

19. The method of claim 17, further comprising using a plurality of decoders to decode the frame.

20. The method of claim 17, further comprising using at least one machine learning (ML) model at the receiver to reconstruct the portion of the frame located outside the ROI portion.