Changing minimal defectivity (JND) threshold or bitrate for game video coding
By using machine learning models to optimize encoding parameters such as frame rate, bit rate, and resolution, the problem of balancing latency and quality in game videos under limited network bandwidth was solved, achieving low-latency and high-quality video transmission while saving energy consumption.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-02
- Publication Date
- 2026-03-27
AI Technical Summary
With limited network bandwidth, existing technologies struggle to strike a balance between ensuring the real-time responsiveness and video quality of game videos, resulting in a trade-off between latency and video quality.
The dataset is trained using a machine learning model. By adjusting the encoding parameters of frame rate, bit rate, and resolution, video encoding is optimized within the minimum perceptible difference (JND) threshold to ensure that the quality of game videos does not significantly degrade, while reducing network bandwidth and power consumption.
It achieves low-latency game video transmission under limited bandwidth, maintains video quality, reduces network and storage requirements, and saves energy consumption.
Smart Images

Figure CN121753331A_ABST
Abstract
Description
Technical Field
[0001] This application relates to unconventional solutions that must be rooted in computer technology and produce specific technical improvements, and more specifically to changing the minimum perceptible difference (JND) threshold or bit rate of game video encoding based on energy-saving requirements. Background Technology
[0002] Videos, such as computer simulation videos and computer game videos, can be streamed to end-user terminals over a network. Summary of the Invention
[0003] As understood in this article, network conditions and / or regulatory restrictions on bandwidth for energy conservation may limit network channels available for transmitting video, such as computer game videos. As further understood in this article, latency is a primary concern in such cases, particularly for computer gamers, who prefer near-instantaneous responses to their inputs, such as when shooting weapons in a game. Therefore, video quality may be less of a concern compared to transmitting video with little or no latency.
[0004] In a first aspect, an apparatus includes at least one processor component configured to input at least one classification of at least one video (e.g., a computer game video) and the video into at least one machine learning (ML) model. The processor component is further configured to receive at least one encoding parameter (selected from the group consisting of frame rate, bit rate, and resolution) from the ML model, and to encode the computer game video using at least partially the encoding parameter for storage and / or transmission over a network of the encoded video.
[0005] In some examples, the ML model can be trained on a training dataset that includes at least one subjective metric indicating whether a human audience can notice differences in the encoding parameters. Subjective parameters may include the mean opinion score (MOS).
[0006] On the other hand, an apparatus includes at least one computer medium that is not a transient signal, and further includes instructions executable by at least one processor component to train at least one machine learning (ML) model using a dataset including videos, corresponding classifications of the videos, and at least one subjective metric indicating whether a human viewer can notice differences in encoding parameters. The instructions are executable to encode at least a first video using the output from the ML model.
[0007] In another aspect, one method includes inputting at least one classification of at least a first video into at least one machine learning (ML) model. The method also includes receiving at least one value of at least one encoding parameter from the ML model and using that value to encode the first video.
[0008] The details of this disclosure regarding both its structure and operation can be best understood with reference to the accompanying drawings, in which the same reference numerals refer to the same parts, and in the accompanying drawings: Attached Figure Description
[0009] Figure 1 It is a block diagram of an exemplary system that includes examples consistent with the principles of the present invention; Figure 2 An example encoder-decoder system is shown; Figure 3 The example logic for training a machine learning (ML) model to establish encoded parameters within the minimum perceptible difference (JND) constraint is illustrated in the form of an example flowchart. Figure 4 A block diagram of an example encoding system is shown; Figure 5 The example logic for analyzing game video classification is illustrated in the format of an example flowchart; and Figure 6 An alternative example logic for establishing encoding parameters is illustrated in the format of an example flowchart. Detailed Implementation
[0010] This disclosure generally relates to a computer ecosystem, which includes various aspects of consumer electronics (CE) device networks, such as, but not limited to, computer gaming networks. The systems described herein may include server components and client components that are network-connected, enabling the exchange of data between the client components and the server components. Client components may include one or more computing devices, including game consoles (such as Sony PlayStation® or game consoles made by Microsoft, Nintendo, or other manufacturers), extended reality (XR) headsets (such as virtual reality (VR) headsets, augmented reality (AR) headsets), portable televisions (e.g., smart TVs, internet-enabled TVs), portable computers (such as laptops and tablets), and other mobile devices (including smartphones and additional examples discussed below). These client devices may operate in a variety of operating environments. For example, some client computers may use operating systems such as Linux, operating systems from Microsoft, or Unix, or operating systems produced by Apple, Inc., Google, Berkeley Software Distribution, or Berkeley Standard Distribution (BSD) OS (including descendants of BSD). These operating environments can be used to execute one or more browsing programs, such as browsers made by Microsoft, Google, or Mozilla, or other browser programs that can access websites hosted by the Internet servers discussed below. Furthermore, the operating environment according to the principles of the present invention can be used to execute one or more computer game programs.
[0011] A server and / or gateway may be used, which may include one or more processors that execute instructions to configure the server to receive and transmit data over a network such as the Internet. Alternatively, the client and server may connect via a local intranet or virtual private network. The server or controller may be instantiated from a game console (such as a Sony PlayStation®), a personal computer, or the like.
[0012] Information can be exchanged between clients and servers over a network. For this purpose, and for security reasons, servers and / or clients may include firewalls, load balancers, temporary storage and proxies, as well as other network infrastructure for reliability and security. One or more servers may form a device that enables methods for providing secure communities (such as online social networking sites or gaming networks) to network members.
[0013] A processor can be a single-chip or multi-chip processor, which executes logic using various lines (such as address lines, data lines, and control lines) as well as registers and shift registers. A processor, including a digital signal processor (DSP), can be an implementation of a circuit system. A processor component can include one or more processors.
[0014] Components included in one embodiment may be used in any suitable combination in other embodiments. For example, any of the various components described herein and / or depicted in the accompanying drawings may be combined, interchanged, or excluded from other embodiments.
[0015] "A system having at least one of A, B, and C" (similarly, "a system having at least one of A, B, or C" and "a system having at least one of A, B, and C") includes the following systems: having only A; having only B; having only C; having both A and B; having both A and C; having both B and C; and / or having both A, B, and C.
[0016] Now for reference Figure 1 An example system 10 is illustrated, which may include one or more of the example devices mentioned above and further described below according to the principles of the invention. The first device among the example devices included in system 10 is a consumer electronics (CE) device, such as an audio-visual device (AVD) 12, such as, but not limited to, a cinema display system (which may be projector-based) or an internet-enabled TV with a TV tuner (equivalently, a set-top box controlling a TV). Alternatively, the AVD 12 may also be a computerized internet-enabled (“smart”) phone, tablet computer, laptop computer, head-mounted device (HMD) and / or head-mounted equipment (such as smart glasses or VR headsets), another wearable computerized device, a computerized internet-enabled music player, a computerized internet-enabled headset, a computerized internet-enabled implantable device (such as an implantable skin device), etc. In any case, it should be understood that the AVD 12 is configured to implement the principles of the invention (e.g., to communicate with other CE devices to implement the principles of the invention, to perform the logic described herein, and to perform any other functions and / or operations described herein).
[0017] Therefore, to implement such principles, the AVD 12 can be constructed using some or all of the components shown. For example, the AVD 12 may include one or more touch-enabled displays 14, which may be implemented using high-definition or ultra-high-definition "4K" or higher resolution flat panel screens. The touch-enabled displays 14 may include, for example, a capacitive or resistive touch sensing layer with an electrode grid for touch sensing, consistent with the principles of the present invention.
[0018] AVD 12 may also include: one or more speakers 16 for outputting audio according to the principles of the invention; and at least one additional input device 18 (such as an audio receiver / microphone) for inputting audible commands to control AVD 12. Example AVD 12 may also include one or more network interfaces 20 for communicating over at least one network 22 (such as the Internet, WAN, LAN, etc.) under the control of one or more processors 24. Therefore, interface 20 may be, but is not limited to, a Wi-Fi transceiver, which is an example of a wireless computer network interface, such as, but not limited to, a mesh network transceiver. It should be understood that processor 24 controls AVD 12 to implement the principles of the invention, including controlling other elements of AVD 12 described herein, such as controlling display 14 to present images on the display and to receive input from the display. Furthermore, note that network interface 20 may be a wired or wireless modem or router, or other suitable interface, such as a wireless telephone transceiver or a Wi-Fi transceiver as mentioned above.
[0019] In addition to the foregoing, AVD 12 may also include one or more input and / or output ports 26, such as a High Definition Multimedia Interface (HDMI) port or a Universal Serial Bus (USB) port for physically connecting to another CE device and / or a headphone port for connecting headphones to AVD 12 to present audio from AVD 12 to a user via headphones. For example, input port 26 may be wired or wirelessly connected to a cable or satellite source 26a for audio / video content. Therefore, source 26a may be a separate or integrated set-top box or satellite receiver. Alternatively, source 26a may be a game console or disk player containing content. When implemented as a game console, source 26a may include some or all of the components described below with respect to CE device 48.
[0020] AVD 12 may also include one or more computer memory / computer-readable storage media 28 that are not transient signals, such as disk-based storage or solid-state storage. In some cases, the one or more computer memory / computer-readable storage media may be embodied as a stand-alone device within the AVD's housing, or as a personal video recording device (PVR) or video disk player for playing back AV programs, either inside or outside the AVD's housing, or as removable storage media or a server described below. Furthermore, in some embodiments, AVD 12 may include a location or positioning receiver, such as, but not limited to, a cellular phone receiver, a GPS receiver, and / or an altimeter 30, configured to receive geographic location information from a satellite or cellular phone base station and provide said information to processor 24 and / or, in conjunction with processor 24, determine the altitude set by AVD 12.
[0021] Continuing the description of AVD 12, in some embodiments, AVD 12 may include one or more cameras 32, which may be thermal imaging cameras, digital cameras (such as webcams, IR sensors, event-based sensors), and / or cameras integrated into AVD 12 and capable of being controlled by processor 24 to acquire pictures / images and / or videos according to the principles of the present invention. AVD 12 may also include a Bluetooth® transceiver 34 and other near-field communication (NFC) elements 36 for communicating with other devices using Bluetooth and / or NFC technologies, respectively. An example NFC element may be a radio frequency identification (RFID) element.
[0022] Furthermore, the AVD 12 may include one or more auxiliary sensors 38 that provide input to the processor 24. For example, one or more of the auxiliary sensors 38 may include one or more pressure sensors forming a layer of the touch-enabled display 14 itself, and may be, but are not limited to, piezoelectric pressure sensors, capacitive pressure sensors, piezoresistive strain gauges, optical pressure sensors, electromagnetic pressure sensors, etc. Other sensor examples include pressure sensors, motion sensors (such as accelerometers, gyroscopes, odometers, or magnetic sensors), infrared (IR) sensors, optical sensors, speed and / or rhythm sensors, event-based sensors, and gesture sensors (e.g., for sensing gesture commands). Thus, the sensor 38 may be implemented by one or more motion sensors, such as individual accelerometers, gyroscopes, and magnetometers and / or inertial measurement units (IMUs), which typically include a combination of accelerometers, gyroscopes, and magnetometers to determine the position and orientation of the AVD 12 in three dimensions, or by event-based sensors (such as event detection sensors (EDS)). Consistent with this disclosure, the EDS provides an output indicating changes in light intensity sensed by at least one pixel of the light sensing array. For example, if the light sensed by the pixel is decreasing, the EDS output can be -1; if the light sensed by the pixel is increasing, the EDS output can be +1. Outputting a binary signal of 0 can indicate that there is no light intensity change below a certain threshold.
[0023] The AVD 12 may also include a wireless TV broadcast port 40 that provides input to the processor 24 for receiving OTA TV broadcasts. In addition to the foregoing, it should be noted that the AVD 12 may also include an infrared (IR) transmitter and / or an IR receiver and / or an IR transceiver 42, such as an IR data association (IRDA) device. A battery (not shown) may be provided to power the AVD 12, such as a kinetic energy harvester that converts kinetic energy into electricity to charge the battery and / or power the AVD 12. A graphics processing unit (GPU) 44 and a field-programmable gate array 46 may also be included. One or more tactile / vibration generators 47 may be provided to generate tactile signals that can be sensed by a person holding or touching the device. Therefore, the haptic generator 47 can use an electric motor to vibrate all or part of the AVD 12, the electric motor being connected via a rotatable shaft to an eccentric and / or unbalanced counterweight, such that the shaft can be rotated under the control of the motor (the electric motor can then be controlled by a processor (such as processor 24)) to generate vibrations of various frequencies and / or amplitudes as well as force simulations in various directions.
[0024] It may also include light sources, such as projectors, such as infrared (IR) projectors.
[0025] In addition to AVD 12, system 10 may also include one or more other CE device types. In one example, the first CE device 48 may be a computer game console that can be used to send audio and video of a computer game to AVD 12 via commands sent directly to AVD 12 and / or via a server described below, while the second CE device 50 may include components similar to the first CE device 48. In the example shown, the second CE device 50 may be configured as a computer game controller operated by a player or a head-mounted display (HMD) worn by a player. The HMD may include a head-up transparent or opaque display for presenting AR / MR content or VR content (more generally, extended reality (XR) content), respectively. The HMD may be configured as a glasses-type display or as a larger VR-type display sold by computer game equipment manufacturers.
[0026] In the example shown, only two CE devices are illustrated; it should be understood that fewer or more devices may be used. The devices described herein may implement some or all of the components shown for AVD 12. Any of the components shown in the following figures may be combined with some or all of the components shown in the case of AVD 12.
[0027] Referring now to the aforementioned at least one server 52, said at least one server includes at least one server processor 54, at least one tangible computer-readable storage medium 56 (such as a disk-based storage device or a solid-state storage device), and at least one network interface 58, which, under the control of the server processor 54, allows communication with other illustrated devices via network 22 and, in practice, facilitates communication between the server and client devices according to the principles of the invention. Note that the network interface 58 may be, for example, a wired or wireless modem or router, a Wi-Fi transceiver, or other suitable interface (such as, for example, a wireless telephone transceiver).
[0028] Therefore, in some implementations, server 52 may be an internet server or an entire server "farm," and in an example implementation for, for example, online gaming applications, the server may include and perform "cloud" functionality, enabling devices of system 10 to access the "cloud" environment via server 52. Alternatively, server 52 may be implemented by one or more game consoles or other computers in the same room or nearby as the other devices shown.
[0029] The components shown in the following figures may include some or all of the components shown in this document. Any user interface (UI) described herein may be combined and / or extended, and UI elements may be mixed and matched between UIs.
[0030] The principles of this invention can be applied to various machine learning models, including deep learning models. Machine learning models based on the principles of this invention can be trained using various algorithms, including supervised learning, unsupervised learning, semi-supervised learning, reinforcement learning, feature learning, self-learning, and other forms of learning. Examples of such algorithms that can be implemented by computer circuit systems include one or more neural networks, such as convolutional neural networks (CNNs), recurrent neural networks (RNNs), and a type of RNN called a long short-term memory (LSTM) network. Generative pre-trained transformers (GPTTs) can also be used. Support vector machines (SVMs) and Bayesian networks can also be considered examples of machine learning models. In addition to the network types mentioned above, the models in this paper can also be implemented using classifiers.
[0031] As understood in this paper, performing machine learning can therefore involve accessing training data and then training a model based on the training data so that the model can process additional data to make inferences. Thus, an artificial neural network / artificial intelligence model trained by machine learning can include an input layer, an output layer, and multiple hidden layers located between them, which are configured and weighted to make inferences about the appropriate output.
[0032] Figure 2 A system is shown that includes a video encoder 200 for encoding / compressing video 202. A video decoder 204 can receive the encoded video and decode / decompress it into an output video 206.
[0033] Figure 3 An example technique for training a machine learning (ML) model, consistent with the disclosure herein, is shown. Code block 300 instructs that some data be fed into the ML model so that the ML model can be trained in code block 302. The data includes multiple videos, such as computer game videos, each labeled with a ground-based live category. For example, one video might be labeled "shooting game," another "driving game," and a third "esports." The videos might be categorized by internal features such as "high motion," "low motion," "large color distribution," and "small color distribution."
[0034] In addition, the video may also include instructions for encoding parameters used to encode the video, such as bit rate, frame rate, and resolution.
[0035] Furthermore, each training video can be labeled with a subjective metric that indicates whether, on average, human viewers can notice differences in the video's bitrate and / or frame rate and / or resolution. One such metric that can be used is the Mean Opinion Score (MOS), which can be viewed as a subjective measure of video quality.
[0036] The goal of training an ML model is to train it to set the bitrate and / or frame rate and / or resolution as low as possible until the minimum perceptible difference (JND) threshold, at which a human can notice a decrease in video quality, is reached. This reduces power consumption and bandwidth usage without significantly degrading video viewing quality.
[0037] With that in mind, let's now look at... Figure 4 Video 400 (e.g., computer game video or movie) and, in some embodiments, video classification indicators, are preprocessed by an ML model 402 trained according to the principles herein to set the frame rate / bit rate / resolution, thereby achieving encoding slightly above the JND threshold. The ML model outputs the frame rate and / or bit rate and / or resolution required for video encoding to encoder 404. The video may be stored and / or transmitted to a display for presentation.
[0038] Figure 5 A technique that can be used in conjunction with this technique is illustrated. If the video's classification is known at state 500, such as a game video with its associated classification or a movie with a known genre, then at state 504, the classification is passed along with the video to the ML model for processing, as described herein. On the other hand, if the classification is unknown, the logic might proceed to state 502 to analyze the video to classify it. This analysis can be performed by training an ML model that classifies the video based on the video frame sequence and audio. The detected classification is then sent to... Figure 4 The ML model 402 is used to process video.
[0039] Figure 6 Further logic that can be implemented is shown. Assuming that at state 600, the video classification to be preprocessed to establish JND parameters (as described above) is unknown, the logic proceeds to state 602 to determine the value of a first parameter, which is initially a default or standard value. The first parameter can be any of the bitrate, frame rate, and resolution. At state 604, the first parameter is changed, and... Figure 4 The ML model 402 in the code determines when the value of the first parameter reaches a MOS score slightly above the JND threshold at state 606. If this can be achieved using the measured value of the first parameter, then at state 608, the value of the first parameter is output to the encoder for encoding the video.
[0040] On the other hand, if no value for the first parameter is found that makes the MOS score slightly higher than the JND threshold, the logic will proceed to state 610 to change the second encoding parameter until a value that makes the MOS score slightly higher than the JND threshold is determined in state 612. The second parameter is different from the first parameter and can be any of the bit rate, frame rate, and resolution.
[0041] If a value for the second parameter that makes the MOS score slightly higher than the JND threshold can be found, then the value of the second parameter is output to the encoder in state 608 to encode the video.
[0042] On the other hand, if no value for the second parameter is found that makes the MOS score slightly higher than the JND threshold, the logic will proceed to state 614 to change the third encoding parameter until a value that makes the MOS score slightly higher than the JND threshold is determined in state 616. The third parameter is different from the first and second parameters and can be any of the bit rate, frame rate, and resolution.
[0043] Furthermore, the ML model can change and analyze combinations of values for the first and second parameters to determine if a combination of the two values of the first and second parameters achieves an acceptable MOS score. If so, these values are used to encode the video. Similarly, the ML model can change and analyze combinations of values for the first and third parameters to determine if a combination of the two values of the first and third parameters achieves an acceptable MOS score. If so, these values are used to encode the video. Further, the ML model can change and analyze combinations of values for the second and third parameters to determine if a combination of the two values of the parameter achieves an acceptable MOS score. If so, these values are used to encode the video. And, the ML model can change and analyze combinations of values for the first, second, and third parameters to determine if a combination of the three values of the corresponding three parameters achieves an acceptable MOS score. If so, these values are used to encode the video. Parameter combinations can be tested after testing each parameter individually, or they can be tested in the intervals between testing individual parameters. If no parameter value achieving an acceptable MOS score is found, the logic ends at state 618, using the default parameter values.
[0044] Therefore, if most viewers cannot distinguish whether the rendering frame rate of a video scene is low or high, the JND model described above will output a lower frame rate for the encoder to encode at. Conversely, if this criterion is no longer met, the JND model will reverse the process and instruct the encoder to encode at a higher frame rate.
[0045] Therefore, this principle reduces the digital / carbon footprint by reducing the output produced by the encoder, and allows the saved network bandwidth or memory to be used for other purposes that are limited by fixed bandwidth (improving video quality, increasing resolution, ROI).
[0046] While specific techniques are shown and described in detail herein, it should be understood that the subject matter contained herein is limited only by the claims.
Claims
1. An apparatus comprising: At least one processor component is configured as follows: Input at least one class of at least one video and the video into at least one machine learning (ML) model; Receive at least one encoding parameter from the ML model, the encoding parameter being selected from a group including frame rate, bit rate, and resolution; and The computer game video is encoded using the encoding parameters at least in part.
2. The apparatus of claim 1, wherein the processor component is configured to: The encoded video is sent to at least one recipient via a computer network.
3. The apparatus of claim 1, wherein the processor component is configured to: Store the encoded video.
4. The apparatus of claim 1, wherein the video includes computer game video.
5. The apparatus of claim 1, wherein the encoding parameters include the frame rate.
6. The apparatus of claim 1, wherein the encoding parameters include bit rate.
7. The apparatus of claim 1, wherein the encoding parameters include resolution.
8. The apparatus of claim 1, wherein the ML model is trained on a training dataset including at least one subjective metric indicating whether a human audience can notice differences in the encoding parameters.
9. The apparatus of claim 8, wherein the subjective parameter includes the mean opinion score (MOS).
10. An apparatus comprising: At least one computer medium, said at least one computer medium not being a transient signal, and including instructions executable by at least one processor component to: At least one machine learning (ML) model is trained using a dataset comprising videos, corresponding classifications of the videos, and at least one subjective metric indicating whether a human viewer can notice differences in coding parameters; and At least the first video is encoded using the output from the ML model.
11. The apparatus of claim 10, wherein the instructions are executable to transmit the encoded first video over a computer network to at least one receiver.
12. The apparatus of claim 10, wherein the instructions are executable to store the encoded first video.
13. The apparatus of claim 10, wherein the first video comprises a computer game video.
14. The apparatus of claim 10, wherein the output includes a frame rate.
15. The apparatus of claim 10, wherein the output includes a bit rate.
16. The apparatus of claim 10, wherein the output includes resolution.
17. The apparatus of claim 10, wherein the subjective parameter includes the mean opinion score (MOS).
18. A method comprising: Input at least one category of at least the first video into at least one machine learning (ML) model; Receive at least one value of at least one encoded parameter from the ML model; and The first video is encoded using the value.
19. The method of claim 18, further comprising inputting the first video along with the classification into the ML model.
20. The method of claim 18, wherein the first video comprises a computer game video.