Video resolution switching algorithm for network streaming applications
By compressing, decompressing, and amplifying the video using processor components, and identifying switching points using RD curves and fitting equations, the resolution is dynamically adjusted, solving the problem of unstable user experience caused by changes in network bandwidth and achieving high-quality video services.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SONY INTERACTIVE ENTERTAINMENT LLC
- Filing Date
- 2024-09-10
- Publication Date
- 2026-04-21
AI Technical Summary
In online video/game streaming, bandwidth fluctuations lead to unstable user experience, and existing technologies struggle to adaptively adjust resolution and frame rate without compromising quality.
The processor component compresses, decompresses, and amplifies videos at multiple resolutions and bit rates to generate quality metrics. It uses rate-distortion (RD) curves and fitting equations to identify switching points and dynamically adjusts video resolution to adapt to network conditions.
It achieves the goal of maintaining stable user experience quality and avoiding visual quality degradation when network bandwidth changes, thus providing uninterrupted high-quality video services.
Smart Images

Figure CN121909653A_ABST
Abstract
Description
Technical Field
[0001] This application relates to a technically inventive and unconventional solution that is necessarily rooted in computer technology and produces specific technical improvements, and more specifically to a video resolution switching algorithm for network streaming applications. Background Technology
[0002] In video / game streaming applications over a network, constant bandwidth is not always guaranteed. Summary of the Invention
[0003] As understood in this article, during low-latency streaming applications, resolution, frame rate, or combinations thereof need to be adaptively adjusted to provide uninterrupted service without sacrificing user experience. For example, network bandwidth may suddenly decrease significantly or become overloaded due to unexpended events. In such cases, 4K streaming services can be tuned by switching the resolution to 2K or even lower, allowing end users to still experience good service quality with virtually no or minimal quality loss by upgrading their receiving resolution to 4K.
[0004] This principle solves the above problems and provides a systematic method to find the bit rate switching points where the resolution should be adjusted to adapt to the underlying network conditions.
[0005] Therefore, an apparatus includes at least one processor component configured to compress at least some of a plurality of selected videos for each of at least some of a plurality of selected resolutions and for each of at least some of a plurality of bitrates at each resolution to render the compressed selected videos. The processor component is configured to decompress each compressed selected video to generate a corresponding output video, and to determine at least one quality metric between each selected video and the corresponding output video for the corresponding bitrate. The processor component is also configured to generate a first data structure based at least in part on the quality metric and the bitrate.
[0006] Furthermore, the processor component is configured to upscale at least some of the output videos for each of at least some of the selected resolutions and for each of at least some of the multiple bitrates at each resolution to render an upscaled video. The processor component is configured to, for each upscaled video, calculate at least one quality metric between the corresponding selected video and the corresponding upscaled video at a corresponding bitrate, and generate a second data structure based at least in part on the quality metric between the corresponding selected video and the corresponding upscaled video and the corresponding bitrate.
[0007] The processor component is configured to use at least a first data structure and a second data structure to identify switching points for switching the resolution of the streaming video, and to use the switching points to change the resolution of the streaming video.
[0008] In some examples, the first and second data structures include rate distortion (RD) curves. In some examples, the first and second data structures include corresponding equations fitted to the respective rate distortion (RD) curves.
[0009] In an example implementation, the processor component may be configured to at least partially use a first data structure and a second data structure, identifying a switching point for switching the resolution of the streaming video by marking the intersection of the two data structures as a switching point.
[0010] In a non-limiting embodiment, the quality metric may include peak signal-to-noise ratio (PSNR). In other embodiments, the quality metric may include a structural similarity index (SSIM).
[0011] If needed, the processor component can be configured to generate the first data structure using an average value based on the bit rate. Similarly, the processor component can be configured to generate the first data structure using an average value based on a quality metric.
[0012] In another aspect, an apparatus includes at least one computer medium that is not a transient signal, and further includes instructions executable by at least one processor component to establish a first resolution for at least a first stream of video at a first time, at least in part, based on a network bitrate meeting a threshold. The instructions are executable to establish a second resolution for the at least a first stream of video at a second time, at least in part, based on a network bitrate not meeting the threshold, wherein the second resolution is lower than the first resolution. Furthermore, the instructions are executable to establish a third resolution for the at least a first stream of video at a third time, at least in part, based on a network bitrate lower than the network bitrate of the second time, wherein the third resolution is lower than the second resolution. The instructions are executable to establish a second resolution for the at least a first stream of video at a fourth time, at least in part, based on a network bitrate higher than the network bitrate of the third time.
[0013] In another aspect, a method includes determining a first quality metric at multiple bitrates between a selected video and a corresponding output video derived from the compressed and decompressed selected video. The method also includes determining a second quality metric at multiple bitrates between the corresponding selected video and an enlarged video generated by enlarging the output video. The method includes using at least partially the first and second quality metrics and bitrates to identify switching points to change the resolution of at least a first video during its streaming.
[0014] The details of this disclosure regarding its structure and operation can be best understood with reference to the accompanying drawings, wherein like reference numerals denote like parts, and wherein: Attached Figure Description
[0015] Figure 1 It is a block diagram of an example system that includes examples consistent with this principle;
[0016] Figure 2 A diagram of a system conforming to this principle is shown;
[0017] Figure 3 The figure shows a graph illustrating bit rate and resolution relative to time;
[0018] Figure 4 Example logic in the form of a sample flowchart for the first step is shown;
[0019] Figure 5 Example logic in the form of a sample flowchart for the second step is shown;
[0020] Figure 6 This is a table showing the results of the first step;
[0021] Figure 7 This is the first part of the table showing the further results of the first step;
[0022] Figure 8 yes Figure 7 The second part of the table shown;
[0023] Figure 9 This is a table that further illustrates the results of the first step;
[0024] Figure 10 This is a summary table of the aforementioned tables;
[0025] Figure 11 A graph showing the rate-distortion (RD) curve associated with the first step is shown;
[0026] Figure 12 This is a table showing the results of the second step; and
[0027] Figure 13A graph showing the intersection of the curves from the first step and the second step is shown. Detailed Implementation
[0028] This disclosure generally relates to the computer ecosystem, including various aspects of consumer electronics (CE) device networks, such as, but not limited to, computer gaming networks. Systems described herein may include server and client components that can be networked, enabling the exchange of data between client and server components. Client components may include one or more computing devices, including game consoles such as Sony PlayStation® or game consoles manufactured by Microsoft, Nintendo, or other manufacturers; extended reality (XR) headsets such as virtual reality (VR) headsets; augmented reality (AR) headsets; portable televisions (e.g., smart TVs, internet-enabled televisions); portable computers such as laptops and tablets; and other mobile devices including smartphones and additional examples discussed below. These client devices may operate with a variety of operating environments. For example, by way of example, some client computers may run on a Linux operating system, an operating system from Microsoft, or a Unix operating system, or an operating system manufactured by Apple, Inc., or Google, or a Berkeley Software distribution or Berkeley Standard Distribution (BSD) OS (including descendants of BSD). These operating environments can be used to execute one or more browsing programs, such as browsers made by Microsoft, Google, or Mozilla, or other browser programs that can access websites hosted by the internet servers discussed below. Furthermore, operating environments based on this principle can be used to execute one or more computer game programs.
[0029] A server and / or gateway may be used, which may include one or more processors executing instructions that configure the server to receive and send data over a network such as the Internet. Alternatively, the client and server may connect via a local intranet or virtual private network. The server or controller may be instantiated from a game console such as a Sony PlayStation®, a personal computer, etc.
[0030] Information can be exchanged between clients and servers over a network. For this purpose, and for security, servers and / or clients may include firewalls, load balancers, temporary storage, and proxies, as well as other network infrastructure for reliability and security. One or more servers can form a means of implementing methods to provide network members with a secure community, such as an online social networking site or a gaming network.
[0031] A processor can be a single-chip or multi-chip processor that executes logic via various lines such as address lines, data lines, and control lines, as well as registers and shift registers. A processor that includes a digital signal processor (DSP) can be an embodiment of a circuit. A processor assembly can include one or more processors.
[0032] Components included in one embodiment may be used in other embodiments in any suitable combination. For example, any of the various components described herein and / or depicted in the accompanying drawings may be combined, interchanged, or excluded from other embodiments.
[0033] "A system having at least one of A, B and C" (similarly, "a system having at least one of A, B or C" and "a system having at least one of A, B, and C") includes systems having only A, only B, only C, A and B together, A and C together, B and C together and / or A, B and C together.
[0034] Now for reference Figure 1 An example system 10 is illustrated, which may include one or more of the example devices mentioned above and further described below in accordance with this principle. A first example device among the example devices included in system 10 is a consumer electronics (CE) device, such as an audio-visual device (AVD) 12, such as, but not limited to, a projector-based theater display system, or an internet-enabled TV with a TV tuner (equivalently, a set-top box controlling the TV). The AVD 12 may alternatively also be a computerized internet-enabled (“smart”) phone, tablet computer, laptop computer, head-mounted device (HMD) and / or headset (such as smart glasses or VR headsets), another wearable computerized device, a computerized internet-enabled music player, a computerized internet-enabled headset, a computerized internet-enabled implantable device (such as an implantable skin device), etc. In any case, it should be understood that the AVD 12 is configured to take this principle (e.g., communicate with other CE devices to take this principle, perform the logic described herein, and perform any other functions and / or operations described herein).
[0035] Therefore, to implement this principle, the AVD 12 can be constructed from some or all of the components shown. For example, the AVD 12 may include one or more touch-enabled displays 14, which may be implemented using a high-definition or ultra-high-definition "4K" or higher flat panel screen. The touch-enabled displays (one or more) 14 may include, for example, a capacitive or resistive touch sensing layer having an electrode grid for touch sensing consistent with this principle.
[0036] AVD 12 may also include one or more speakers 16 for outputting audio according to these principles, and at least one additional input device 18, such as an audio receiver / microphone, for inputting audible commands to control AVD 12. The exemplary AVD 12 may also include one or more network interfaces 20 for communicating over at least one network 22 (such as the Internet, WAN, LAN, etc.) under the control of one or more processors 24. Thus, interface 20 may be, but is not limited to, a Wi-Fi transceiver, which is an example of a wireless computer network interface, such as, but not limited to, a mesh network transceiver. It should be understood that processor 24 controls AVD 12 to perform these principles, including other elements of AVD 12 described herein, such as controlling display 14 to render images thereon and receive input from it. Furthermore, note that network interface 20 may be a wired or wireless modem or router, or other suitable interface, such as a wireless telephone transceiver or a Wi-Fi transceiver as described above.
[0037] In addition to the above, the AVD 12 may also include one or more input and / or output ports 26, such as an HDMI port or a USB port, for physical connection to another CE device and / or a headphone port for connecting headphones to the AVD 12 to present audio from the AVD 12 to the user via headphones. For example, input port 26 may be connected via a cable or satellite source 26a to audio / video content, either wired or wirelessly. Therefore, source 26a may be a standalone or integrated set-top box or satellite receiver. Alternatively, source 26a may be a game console or disc player containing content. When implemented as a game console, source 26a may include some or all of the components described below with respect to CE device 48.
[0038] AVD 12 may also include one or more computer memory / computer-readable storage media 28, such as non-transient signal disk storage or solid-state storage, which in some cases are embodied in the chassis of the AVD as a standalone device or personal video recording device (PVR) or video disk player, located inside or outside the AVD chassis for playing AV programs, or as a removable storage medium or a server as described below. Furthermore, in some embodiments, AVD 12 may include a location or place receiver, such as, but not limited to, a mobile phone receiver, a GPS receiver, and / or an altimeter 30, configured to receive geolocation information from a satellite or mobile phone base station and provide that information to the processor 24 and / or determine the height at which the AVD 12 is placed with the processor 24.
[0039] Continuing the description of AVD 12, in some embodiments, AVD 12 may include one or more cameras 32, which may be thermal imaging cameras, digital cameras such as webcams, IR sensors, event-based sensors, and / or cameras integrated into AVD 12 and controllable by processor 24 to collect pictures / images and / or videos in accordance with these principles. AVD 12 may also include a Bluetooth® transceiver 34 and other near-field communication (NFC) elements 36 for communicating with other devices using Bluetooth and / or NFC technologies, respectively. An example NFC element may be a radio frequency identification (RFID) element.
[0040] Furthermore, the AVD 12 may include one or more auxiliary sensors 38 that provide input to the processor 24. For example, one or more of the auxiliary sensors 38 may include one or more pressure sensors forming a layer of the touch-enabled display 14 itself, and may be, but are not limited to, piezoelectric pressure sensors, capacitive pressure sensors, piezoresistive strain gauges, optical pressure sensors, electromagnetic pressure sensors, etc. Other sensor examples include pressure sensors, motion sensors (such as accelerometers, gyroscopes, tachometers) or magnetic sensors, infrared (IR) sensors, optical sensors, speed and / or rhythm sensors, event-based sensors, gesture sensors (e.g., for sensing gesture commands). Thus, the sensor 38 may be implemented by one or more motion sensors, such as individual accelerometers, gyroscopes, and magnetometers, and / or an inertial measurement unit (IMU) that typically includes a combination of accelerometers, gyroscopes, and magnetometers, to determine the position and orientation of the AVD 12 in three dimensions, or by event-based sensors, such as an event detection sensor (EDS). Consistent with this disclosure, the EDS provides an output indicating changes in light intensity sensed by at least one pixel of the light sensing array. For example, if the light sensed by the pixel is decreasing, the output of EDS can be -1; if it is increasing, the output of EDS can be +1. No change in light intensity below a certain threshold can be indicated by an output binary signal of 0.
[0041] The AVD 12 may also include an over-the-air (OTA) television broadcast port 40 for receiving OTA television broadcasts that provide input to the processor 24. In addition to the foregoing, it should be noted that the AVD 12 may also include an infrared (IR) transmitter and / or an IR receiver and / or an IR transceiver 42, such as an IR data association (IRDA) device. A battery (not shown) may be provided to power the AVD 12, such as a kinetic energy harvester that can convert kinetic energy into electrical energy to charge the battery and / or power the AVD 12. A graphics processing unit (GPU) 44 and a field-programmable gate array (FPGA) 46 may also be included. One or more tactile / vibration generators 47 may be provided to generate tactile signals that can be felt by a person holding or touching the device. Therefore, the haptic generator 47 can use an electric motor to vibrate all or part of the AVD 12, which is connected to an eccentric and / or unbalanced weight via a rotatable shaft, such that the shaft can be rotated under the control of the motor (which can in turn be controlled by a processor (e.g., processor 24)) to generate vibrations of various frequencies and / or amplitudes as well as force simulations of various directions.
[0042] It may also include a light source such as a projector (such as an infrared (IR) projector).
[0043] In addition to AVD 12, System 10 may include one or more other CE device types. In one example, the first CE device 48 may be a computer game console that can be used to send computer game audio and video to AVD 12 via commands sent directly to AVD 12 and / or via a server described below, while the second CE device 50 may include components similar to the first CE device 48. In the example shown, the second CE device 50 may be configured as a computer game controller operated by a player or a head-mounted display (HMD) worn by a player. The HMD may include a heads-up transparent or opaque display for presenting AR / MR content or VR content (more generally, extended reality (XR) content), respectively. The HMD may be configured as a glasses-type display sold by a computer game device manufacturer or a larger VR-type display.
[0044] In the example shown, only two CE devices are illustrated; it should be understood that fewer or more devices may be used. The devices described herein may implement some or all of the components shown for AVD 12. Any components shown in the following figures may be combined with some or all of the components shown in the case of AVD 12.
[0045] Referring now to at least one server 52 described above, it includes at least one server processor 54, at least one tangible computer-readable storage medium 56 (such as a disk-based or solid-state storage device), and at least one network interface 58, which, under the control of the server processor 54, allows communication with other illustrated devices via network 22 and, in practice, facilitates communication between server and client devices according to this principle. Note that the network interface 58 may be, for example, a wired or wireless modem or router, a Wi-Fi transceiver, or other suitable interface, such as a wireless telephone transceiver.
[0046] Therefore, in some embodiments, server 52 may be an internet server or an entire server "farm" and may include and perform "cloud" functionality, enabling devices of system 10 to access a "cloud" environment via server 52 in example embodiments for, for example, online gaming applications. Alternatively, server 52 may be implemented by one or more game consoles or other computers in the same room as other devices shown or nearby.
[0047] The components shown in the following figures may include some or all of the components shown herein. Any user interface (UI) described herein may be combined and / or extended, and UI elements may be mixed and matched between UIs.
[0048] This principle can be applied to various machine learning models, including deep learning models. Machine learning models consistent with this principle can be trained using a variety of algorithms, including supervised learning, unsupervised learning, semi-supervised learning, reinforcement learning, feature learning, self-learning, and other forms of learning. Examples of such algorithms that can be implemented by computer circuits include one or more neural networks, such as convolutional neural networks (CNNs), recurrent neural networks (RNNs), and RNN types known as long short-term memory (LSTM) networks. Generative pre-trained transformers (GPTTs) can also be used. Support vector machines (SVMs) and Bayesian networks can also be considered examples of machine learning models. In addition to the network types described above, the models in this paper can be implemented using classifiers.
[0049] As understood in this paper, performing machine learning can therefore involve accessing and then training a model on training data so that the model can process further data to make inferences. Thus, an artificial neural network / AI model trained via machine learning can include an input layer, an output layer, and multiple hidden layers in between, which are configured and weighted to make inferences about the appropriate output.
[0050] Figure 2A system including a video encoder 200 for encoding / compressing video 202 is shown. A video decoder 204 can receive the encoded videos and decode / decompress them into an output video 206.
[0051] Figure 3 A scenario showing how an example of resolution switching operates is presented. At time T0, if the network bitrate R > R1, the video resolution is S1. At time T1, if the network bitrate is lower than R1 (i.e., R < R1), then resolution S2 (where S2 < S) is selected for encoding. At time T2, the network bitrate further decreases to below R2, and then resolution S3 (< S2) is selected for encoding. At time T3, the network bitrate increases and > R2, and then a larger resolution S2 is selected for encoding. R1 and R2 in this graph are designated as the bitrates at which video resolution switching occurs, and the proposed algorithm for deriving these bitrates is described below.
[0052] The ultimate goal of the proposed resolution switcher is to provide good quality of service to end - users without experiencing an unpleasant degradation in visual quality. In one implementation, the resolution switcher can be designed in two parts. The first part is a conventional rate - distortion (RD) derivation, as described below and as Figure 4 shown.
[0053] In state 400, in an example switch design, "p" representative videos "V", V k, k=1..p , with different resolutions "S", S1, S2,..., S k , (S1 > S2 >... > S m ) are selected.
[0054] Now cross - reference the following pseudocode with the Figure 4 logic states. In state 402, representative bitrates "R", R1, R 2, ,,, R m (R1 > R2 > …… > R n ) are selected for the expected or anticipated network conditions. Then enter a nested "DO" loop:
[0055] For each selected resolution S i, i=1..m , state 404
[0056] For each selected bitrate R j, j=1..n , state 406
[0057] For each selected video V k, k=1..p , state 408
[0058] At the selected resolution S i and bitrate R jCompress the selected video V k Status 410
[0059] Decompress the bitstream to reconstruct the output video O k Status 412
[0060] Calculate the selected video V k and output video O k Bit rate R between j Visual quality metric at location, state 414
[0061] Calculate bit rate R j The average quality metric for all videos under state 416
[0062] Therefore, quality metrics are calculated for all videos at all resolutions and bit rates.
[0063] State 418 instructs the derivation of the first fitting equation E1 to best describe the RD curve. Note that the quality metric can be an objective measure such as Peak Signal-to-Noise Ratio (PSNR) or Structural Similarity Index (SSIM), or a subjective measure such as Mean Opinion Score (MOS). The RD curve of the source video is available.
[0064] Now that the first step of the example implementation has been described, let’s turn our attention to… Figure 5 To understand the second step, which involves finding the RD curve of the "magnified" visual representation of the output video "O". For the output video O... k-1 Its scaled-up version is represented as O”. k-1 And O” k-1 resolution and O k The resolutions are the same. For example, if the resolution S2 of O2 is 1920 × 1080, then its scaled-up version O”2 has a resolution of 3840 × 2160, which is also the same resolution S1 as O1 and V1. The quality metric is then calculated between O”2 and V1. Note that any common scaling algorithm will be applicable in this case, but ideally, an algorithm that "perfectly" matches the actual use case should be used at the receiving end.
[0065] Starting from state 500, select "p" representative videos (V). k, k=1..p In the design section 2 of the switcher, different resolutions S are selected. 2, .., S k (S2>..>S) m In state 502, select a representative bit rate R1,R for the network conditions. 2, ,, , R m (R1>R2>...>R) n ).
[0066] Status 504 indicates a nested "DO" loop, where
[0067] For each selected resolution S i, i=2..m
[0068] For each selected bit rate R j, j=1..n
[0069] Targeting from Figure 4 Each selected reconstructed output video O k, k=1..p
[0070] Status 506 will be at resolution S i Test O k The output video is enlarged to a resolution of S. i-1 "O" k (that is, from S) i To S i-1 (e.g., from 1080p to 2160p)
[0071] State 508 Calculation Resolution S i-1 The selected video V k With resolution S i-1 The magnified output video O” k Bit rate R between j quality measurement
[0072] State 510 calculates bit rate R j The average quality metric of all reconstructed output videos.
[0073] For all reconstructed output videos at all resolutions and bit rates, a quality metric is calculated. At state 512, a second fitting equation E2 is derived to best describe the quality of the video output. Figure 5 The resulting RD curve. After deriving the first equation E1 and the second equation E2, the final step is to find their intersection point at state 514, which can be used as the optimal switching point for resolution change at state 516. It should be understood that the intersection point of the two equations describing the RD curve, or the graphical intersection point between the curves themselves, can be used to identify the switching point.
[0074] Figures 6 to 9 Presented in tabular form Figure 4 The logic applies to the results at three corresponding resolutions ( Figure 7 and Figure 8 (This is part of the same table). The first column shows the resolution "S" for a specific graph, the second column lists the three selected videos "V", and the third column shows the target bit rate "R". The fourth column shows the actual bit rate, and the fifth column shows the quality metric, which is PSNR in the example shown.
[0075] like Figures 6 to 9 As shown, three videos were selected: FF15-Sunrise, GTSport-Circuit, and Resogun-Decima. The three different resolutions are 3840 × 2160 with ten different bitrates. Figure 6 ), with eleven different bit rates, 1920 × 1080 ( Figure 7 and Figure 8 ) and 1280 × 720 with six different bit rates ( Figure 9 ). Figure 10 This is summarized by showing the resolution "S" in the first column, the associated average actual bit rate "R" in the second column, and the average quality metric in the third column. Figures 6 to 9 .
[0076] Figure 11 It shows Figures 6 to 9 As shown and in Figure 10 The RD curves for the three resolutions summarized in the paper are obtained by plotting the average quality metric on the y-axis relative to the corresponding average actual bit rate on the x-axis. Specifically, Figure 11 The first curve 1100 in the figure shows the RD curve at a resolution of 720p, where the corresponding fitting equation curve 1102 (shown as a dashed line) has a second-order example polynomial approximation. Similarly, Figure 11 The second curve 1104 in the figure shows the RD curve at a resolution of 1080p, where the corresponding fitting equation curve 1106 (represented by the dashed line) has a second-order polynomial approximation. Similarly, Figure 11 The third curve 1108 in the figure shows the RD curve at a resolution of 2160p, where the corresponding fitting equation curve 1110 (shown as a dashed line) has a second-order polynomial approximation. The corresponding actual fitting equations for curves 1102, 1106, and 1110 are shown at 1112, 1114, and 1116.
[0077] Note that any fitting equation can be applied, as long as its error is acceptable.
[0078] As mentioned above, once according to Figure 4Having determined the RD curve with the fitted equation, the logic of this graph implements the second step: to find the RD curve for upscaled data from the generated video "O", for example, finding 2160p data from 1080p data. As discussed, the output 1080p video "O" is upscaled to 2160p (O") using an upscale algorithm, and then quality metrics such as PSNR and / or SSIM are calculated between the upscaled 2160p video "O" and the native 2160p source video "V" to obtain a quality measurement. Any scaling algorithm can be used, such as bilinear or bicubic. In one embodiment, the so-called Lanczos algorithm can be used.
[0079] Figure 12 It shows Figure 5 The result. In Figure 12 In the diagram, the first column shows the resolution, the second column shows the specific video, the third column shows the target bit rate, the fourth column shows the actual bit rate achieved, the fifth column shows the corresponding quality metric, and the sixth column shows the corresponding quality metric for the enlarged video.
[0080] Figure 13 The original 2160p at 1108, the original 1080p at 1104, and the magnified 2160p RD curve at 1300 are shown, along with an approximate equation for its fitted equation at 1302.
[0081] From Figure 4 The fitting equation for the native 2160p curve is E(native 2160p) = 2.9166. ln (bit rate) + 22.105, and from Figure 5 The fitting equation for the magnified 2160p curve is E(magnified 2160p) = 2.3495. ln (bit rate) + 23.816.
[0082] The intersection of these two equations then indicates the critical switching point in bitrate between native 2160p and upscaled 4K (i.e., native 1080p). In our example, the intersection is approximately 20.5 Mbps, meaning that for bitrates below 20.5 Mbps, native 1080p output video exhibits better quality (e.g., PNSR) after being upscaled to 2160p on the receiving side. On the other hand, for bitrates above 20.5 Mbps, native 2160p output is preferred. The target bitrate is then converted to 22.96 Mbps.
[0083] While specific techniques are shown and described in detail herein, it should be understood that the subject matter covered by this application is limited only by the claims.
Claims
1. An apparatus comprising: At least one processor component is configured as follows: For each of at least some of the multiple selected resolutions, and for each of at least some of the multiple bitrates of each resolution, compress at least some of the multiple selected videos to render the compressed selected videos; Decompress each of the selected compressed videos to generate the corresponding output video; Determine at least one quality metric between each selected video and the corresponding output video at the corresponding bit rate; The first data structure is generated based at least in part on the quality metric and bit rate; For each of at least some of the selected resolutions, and for each of at least some of the multiple bitrates at each resolution, at least some of the output videos are upscaled to render an upscaled video; For each amplified video, calculate at least one quality metric for the corresponding bitrate between the selected video and the amplified video. The second data structure is generated based at least in part on the quality metric between the selected video and the corresponding amplified video and the corresponding bit rate; Using at least part of the first and second data structures, the switching points for switching the resolution of the streaming video are identified; as well as Use the switching point to change the resolution of the streaming video.
2. The apparatus according to claim 1, wherein, The first data structure and the second data structure include rate-distortion (RD) curves.
3. The apparatus according to claim 1, wherein, The first data structure and the second data structure include corresponding equations fitted to the corresponding rate-distortion (RD) curve.
4. The apparatus according to claim 1, wherein, The processor component is configured to at least partially use the first data structure and the second data structure, and to identify the switching point by recognizing the intersection between the two data structures as a switching point for switching the resolution of the streaming video.
5. The apparatus according to claim 1, wherein, The quality metric includes peak signal-to-noise ratio (PSNR).
6. The apparatus according to claim 1, wherein, The quality metric includes the structural similarity index (SSIM).
7. The apparatus according to claim 1, wherein, The processor component is configured to generate the first data structure using an average value based on the bit rate.
8. The apparatus according to claim 1, wherein, The processor component is configured to generate the first data structure using an average value based on the quality metric.
9. An apparatus comprising: At least one computer medium, said at least one computer medium being not a non-transitory signal and including instructions, said instructions being executable by at least one processor component to: At the first moment, at least in part based on the network bit rate meeting the threshold, a first resolution is established for at least the first video being streamed; At a second time following the first time, at least in part based on the network bit rate not meeting the threshold, a second resolution is established for at least the first video being streamed, the second resolution being lower than the first resolution; At a third time following the second time, at least in part based on the network bit rate being less than the network bit rate at the second time, a third resolution is established for at least the first video being streamed, the third resolution being lower than the second resolution; as well as At a fourth time following the third time, the second resolution is established for at least the first video being streamed, based at least in part on the network bit rate being greater than the network bit rate at the third time.
10. The apparatus according to claim 9, wherein, The instructions can be executed to: Switch the resolution at the switching point determined by the intersection of the first and second data structures.
11. The apparatus according to claim 10, wherein, The instructions are executable to determine the intersection point at least in part by: For each of at least some of the multiple selected resolutions, and for each of at least some of the multiple bitrates of each resolution, compress at least some of the multiple selected videos to render the compressed selected videos; Decompress each of the selected compressed videos to generate the corresponding output video; Determine at least one quality metric between each selected video and the corresponding output video at the corresponding bit rate; The first data structure is generated based at least in part on the quality metric and bit rate; For each of at least some of the selected resolutions, and for each of at least some of the multiple bitrates at each resolution, at least some of the output videos are upscaled to render an upscaled video; For each amplified video, calculate at least one quality metric for the corresponding bitrate between the selected video and the amplified video. The second data structure is generated based at least in part on the quality metric between the selected video and the corresponding amplified video and the corresponding bit rate; The switching point is identified using at least part of the first and second data structures.
12. The apparatus according to claim 11, wherein, The first data structure and the second data structure include rate-distortion (RD) curves.
13. The apparatus according to claim 11, wherein, The first data structure and the second data structure include corresponding equations fitted to the corresponding rate-distortion (RD) curve.
14. The apparatus according to claim 11, wherein, The quality metric includes peak signal-to-noise ratio (PSNR).
15. The apparatus according to claim 11, wherein, The quality metric includes the structural similarity index (SSIM).
16. A method comprising: Determine a first quality metric at multiple bit rates between the selected video and the corresponding output video derived from the compressed and decompressed selected video; Determine a second quality metric at multiple bit rates between the selected video and the amplified video generated by amplifying the output video; as well as Using at least part of the first quality metric and the second quality metric, along with the bit rate, a switching point is identified to change the resolution of at least the first video during streaming.
17. The method of claim 16, comprising: While streaming the first video, the resolution of the first video is switched using the switching point.
18. The method of claim 16, further comprising identifying the switching point by recognizing an intersection between a first rate-distortion RD data structure derived from the first quality metric and a second RD data structure derived from the second quality metric.
19. The method according to claim 18, wherein, The first RD data structure and the second RD data structure include rate-distortion RD curves.
20. The method according to claim 18, wherein, The first RD data structure and the second RD data structure include corresponding equations fitted to the corresponding rate-distortion RD curves.