Method, system, and medium for selecting a format for streaming a media content item

A server-based system optimizes media content streaming by selecting formats based on network and device information, adapting to changes, and predicting user preferences, enhancing viewing efficiency and quality.

JP7761699B2Active Publication Date: 2025-10-28GOOGLE LLC
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2024062621
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-04-09
Publication Date
2025-10-28
Estimated Expiration
2039-12-11

AI Technical Summary

Technical Problem

Existing user devices struggle with resource-intensive determination of optimal streaming formats for media content items, often requesting inappropriate formats based on network and device resources, leading to suboptimal viewing experiences.

Method used

A server-based system that selects streaming formats for media content items using network and device information, dynamically adapting to changes in network quality and device conditions, utilizing trained models to predict user preferences and optimize viewing duration or quality scores.

Benefits of technology

Enhances streaming efficiency by dynamically adjusting formats to match user preferences and network conditions, optimizing viewing duration and quality, reducing resource wastage and improving user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007761699000001
    Figure 0007761699000001
  • Figure 0007761699000002
    Figure 0007761699000002
  • Figure 0007761699000003
    Figure 0007761699000003
Patent Text Reader

Abstract

To provide a new method for selecting a format for content streaming.SOLUTION: A method includes receiving a request to begin streaming at a server, receiving network information and device information, selecting a first format for a video content item including a first resolution of the multiple resolutions on the basis of the network information and device information, transmitting a first portion of the video content item having the first format, receiving updated network information and device information, selecting a second format for the video content item including a second resolution of the multiple resolutions on the basis of the updated network information and device information, and transmitting a second portion of the video content item having the second format.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The disclosed subject matter relates to methods, systems, and media for selecting a format for streaming a media content item. [Background technology]

[0002] Users frequently stream video content (e.g., videos, movies, television programs, music videos, etc.) from media content streaming services. User devices may use adaptive bitrate streaming, which may allow the user device to request different qualities of a video content item as it is streamed from a server, allowing the user device to continue presenting the video content item even if the quality of the network connection used to stream the video content item changes. For example, the user device may begin presenting a segment of the video content item having a first, relatively high resolution, and then, in response to determining that the network connection has degraded, the user device may request a segment of the video content item having a second, lower resolution from the server. However, it is resource-intensive for the user device to determine the optimal format to request from the server. Furthermore, the user device may request a particular format regardless of the resources available to the server. Summary of the Invention [Problem to be solved by the invention]

[0003] Therefore, it would be desirable to provide new methods, systems, and media for selecting a format for streaming a media content item. [Means for solving the problem]

[0004] SUMMARY OF THE INVENTION Methods, systems, and media are provided for selecting a format for streaming a media content item.

[0005] According to some embodiments of the disclosed subject matter, a method for selecting a format for streaming a media content item is provided, the method including the steps of: receiving, at a server, a request from a user device to start streaming a video content item on the user device; receiving, from the user device, network information indicating a quality of the user device's network connection to a communications network used to stream the video content item and device information associated with the user device; selecting, by the server, a first format for the video content item, the first format including a first resolution from a plurality of resolutions based on the network information and the device information; transmitting, from the server, a first portion of the video content item having the first format to the user device; receiving, at the server, updated network information and updated device information from the user device; selecting, by the server, a second format for the video content item, the second format including a second resolution from a plurality of resolutions based on the updated network information and the updated device information; and transmitting, from the server, a second portion of the video content item having the second format to the user device.

[0006] In some embodiments, selecting a first format for the video content item includes predicting, by the server, a format likely to be selected by a user of the user device.

[0007] In some embodiments, selecting a first format for the video content item includes identifying, by the server, a format that maximizes an expected duration for a user of the user device to stream the video content item.

[0008] In some embodiments, the updated device information includes an indication that the size of the viewport of a video player window running on the user device on which the video content item is being presented has changed.

[0009] In some embodiments, the updated device information indicates that the viewport size has decreased, and the second resolution is lower than the first resolution.

[0010] In some embodiments, the first format for the video content item is selected based on the genre of the video content item.

[0011] In some embodiments, the first format for the video content item is selected based on a format pre-selected by a user of a user device for streaming other video content items from a server.

[0012] According to some embodiments of the disclosed subject matter, there is provided a system for selecting a format for streaming a media content item, the system comprising: a memory and a hardware processor, the hardware processor executing computer-executable instructions stored in the memory, the system including: receiving, at a server, from a user device, a request to begin streaming a video content item on the user device; receiving from the user device network information indicative of a quality of a network connection of the user device to a communications network used to stream the video content item and device information associated with the user device; and selecting, by the server, a first format for the video content item. the first format includes a first resolution of a plurality of resolutions based on the network information and the device information; transmitting from the server a first portion of the video content item having the first format to the user device; receiving at the server updated network information and updated device information from the user device; selecting, by the server, a second format for the video content item, the second format includes a second resolution of the plurality of resolutions based on the updated network information and the updated device information; and transmitting from the server the second portion of the video content item having the second format to the user device.

[0013] According to some embodiments of the disclosed subject matter, there is provided a non-transitory computer-readable medium containing computer-executable instructions that, when executed by a processor, cause the processor to perform a method for selecting a format for streaming a media content item, the method including the steps of receiving, at a server, a request from a user device to begin streaming the video content item on the user device; receiving, from the user device, network information indicative of a quality of a network connection of the user device to a communications network used to stream the video content item and device information associated with the user device; and selecting, by the server, a first format for the video content item. The method includes selecting a format for the video content item, the first format including a first resolution from a plurality of resolutions based on the network information and the device information; transmitting a first portion of the video content item having the first format from the server to the user device; receiving updated network information and updated device information from the user device at the server; selecting a second format for the video content item by the server, the second format including a second resolution from the plurality of resolutions based on the updated network information and the updated device information; and transmitting the second portion of the video content item having the second format from the server to the user device.

[0014] According to some embodiments of the disclosed subject matter, a system for selecting a format for streaming a media content item is provided, the system comprising: means for receiving, at a server, from a user device, a request to start streaming a video content item on the user device; means for receiving, from the user device, network information indicative of a quality of the user device's network connection to a communications network used to stream the video content item and device information associated with the user device; means for selecting, by the server, a first format for the video content item, the first format including a first resolution from a plurality of resolutions based on the network information and the device information; means for transmitting, from the server, a first portion of the video content item having the first format to the user device; means for receiving, at the server, updated network information and updated device information from the user device; means for selecting, by the server, a second format for the video content item, the second format including a second resolution from a plurality of resolutions based on the updated network information and the updated device information; and means for transmitting, from the server, a second portion of the video content item having the second format to the user device.

[0015] Various objects, features, and advantages of the disclosed subject matter may be more fully understood by reference to the following detailed description of the disclosed subject matter when considered in conjunction with the following drawings, in which like reference numerals identify like elements and in which: [Brief explanation of the drawings]

[0016] [Figure 1] 1 shows an illustrative example of a process for selecting a format for a video content item according to some embodiments of the disclosed subject matter. [Figure 2]FIG. 10 shows an illustrative example of a process for training a model to select a format for a video content item based on a pre-selected format, according to some embodiments of the disclosed subject matter. [Figure 3] FIG. 10 shows an illustrative example of a process for training a model to select a format for a video content item based on quality scores associated with streaming video content items having different formats, according to some embodiments of the disclosed subject matter. [Figure 4] FIG. 1 is a schematic diagram of an illustrative system suitable for implementing the mechanisms described herein for selecting a format for streaming a media content item, according to some embodiments of the disclosed subject matter. [Figure 5] 5 illustrates a detailed example of hardware that may be used in the server and / or user device of FIG. 4 in accordance with some embodiments of the disclosed subject matter. DETAILED DESCRIPTION OF THE INVENTION

[0017] According to various embodiments, mechanisms (which may include methods, systems, and media) are provided for selecting a format for streaming a media content item.

[0018] In some embodiments, the mechanisms described herein may select, by a server, a format (e.g., a particular resolution and / or any other suitable format) of a video content item to be streamed to a user device based on information such as the current quality or type of network connection used by the user device to stream the video content item, the type of device associated with the user device, a video content format pre-selected by a user of the user device, and / or any other suitable information. In some embodiments, the server may then begin streaming the video content item to the user device in the selected format. In some embodiments, the server may receive updated information from the user device indicating any appropriate changes, such as a change in the quality of the network connection. In some embodiments, the server may then select a different format based on the updated information and switch to streaming the video content item in the selected different format.

[0019] For example, in some embodiments, the server may begin streaming the video content item to the user device at a first resolution (e.g., 360 pixels by 640 pixels and / or any other suitable first resolution) selected based on network information, device information, and / or any other suitable information. Continuing with this example, in some embodiments, the server may receive information from the user device indicating a change in the status of the streaming of the video content item. For example, the server may receive information indicating a change in the quality of the network connection, a change in the size of a viewport of a video player window running on the user device, and / or any other suitable change. Further continuing with this example, in some embodiments, the server may then select a second resolution based on the change in the status of the streaming of the video content item. For example, in the instance where the server receives information indicating that the quality of the network connection has decreased, the server may switch to streaming the video content item at a resolution lower than the first resolution. Conversely, in the instance where the server receives information indicating that the quality of the network connection has improved, the server may switch to streaming the video content item at a resolution higher than the first resolution. As another example, in the case where a server receives information indicating that the size of the viewport of a video player window has increased, the server may switch to streaming the video content item at a resolution higher than the first resolution. Conversely, in the case where a server receives information indicating that the size of the viewport has decreased, the server may switch to streaming the video content item at a resolution lower than the first resolution.

[0020] In some embodiments, the server may select a format using any suitable technique or combination of techniques. For example, in some embodiments, the server may select a format that is predicted to be the format that would be manually selected by a user when streaming the video content item under similar network conditions and / or using a similar user device. As a more specific example, in some embodiments, the server may use a trained model (e.g., a trained decision tree, a trained neural network, and / or any other suitable model) trained using training samples, each corresponding to a streamed video content item, each training sample including any suitable input features (e.g., the network conditions under which the video content item was streamed, information related to the user device that streamed the video content item, information about the video content item, and / or any other suitable information) and the corresponding user-selected format, as shown in FIG. 2 and described below in connection therewith. As another example, in some embodiments, the server may select a format that is predicted to maximize any suitable quality score that predicts the quality of streaming of the video content item to a user device under certain conditions. As a more specific example, in some embodiments, the server may use a trained model (e.g., a trained decision tree, a trained neural network, and / or any other suitable model) trained using training samples each corresponding to a streamed video content item, each training sample including any suitable input features (e.g., network conditions under which the video content item was streamed, information related to the user device that streamed the video content item, information about the video content item, and / or any other suitable information) and a corresponding quality score, as shown in FIG. 3 and described below in connection therewith.It should be noted that in some embodiments, the quality score may be based on any suitable metric or combination of metrics, for example, the duration that a user watches video content items, including a video content item associated with a training sample, during a video content viewing session, the duration that a user watches a video content item associated with a training sample, the latency experienced by a user between streaming two video content items, and / or any other suitable metric or combination of metrics.

[0021] In some embodiments, by using a trained model to select the format of a video content item to be streamed, the mechanisms described herein may enable the server to select a format that optimizes any appropriate goal, e.g., a format likely to be manually selected by a user of a user device streaming the video content item, a format that maximizes the duration that the user of the user device views the video content item, and / or any other appropriate goal. Additionally, in some embodiments, the mechanisms may enable the server to change the format of a video content item streamed to a user device based on any appropriate change while the video device is streaming, e.g., changes in network conditions, changes in the size of the viewport of a video player window used to present the video content item, and / or any other appropriate change. In some embodiments, by changing the format of a video content item while the video content item is streaming, the mechanism can dynamically adapt the format of the video content item so that any appropriate goal continues to be optimized while streaming conditions change.

[0022] It should be noted that although the mechanisms described herein generally relate to a server selecting a particular quality of a video content item to be streamed to a user device by selecting a particular format or resolution of the video content item, in some embodiments the quality of the video content may be indicated using any other suitable metric, for example Video Multimethod Assessment Fusion (VMAF), and / or any other suitable metric.

[0023] 1 , an illustrative example 100 of a process for selecting a format for a video content item is shown in accordance with some embodiments of the disclosed subject matter. In some embodiments, the blocks of process 100 may be performed by any suitable device. For example, in some embodiments, the blocks of process 100 may be performed by a server associated with a video content streaming service.

[0024] Process 100 begins at 102, when a server may receive a request from a user device to begin streaming a video content item on the user device. As described above, in some embodiments, the server may be associated with any suitable entity or service, e.g., a video content streaming service, a social networking service, and / or any other suitable entity or service. In some embodiments, the server may receive the request to begin streaming the video content item in any suitable manner. For example, in some embodiments, the server may receive the request from the user device in response to determining that a user of the user device has selected an indication of the video content item (e.g., an indication presented on a page for browsing video content items and / or selected in any other suitable manner). As another example, in some embodiments, the server may receive an indication that the video content item is a subsequent video content item in a playlist of video content items that should be presented consecutively on the user device.

[0025] At 104, process 100 may receive network information and device information from the user device. In some embodiments, the network information may include any suitable metric indicative of the quality of the user device's connection to a communications network used by the user device to stream video content items from a server and / or any suitable information about the type of network used. For example, in some embodiments, the network information may include the bandwidth of the connection, the throughput of the connection, the goodput of the connection, the latency of the connection, the type of connection used (e.g., an Ethernet connection, a 3G connection, a 4G connection, a Wi-Fi connection, and / or any other suitable type of network), the type of communications protocol used (e.g., HTTP and / or any other suitable type of protocol), any suitable network identifier (e.g., an Autonomous System Number (ASN) and / or any other suitable identifier), and / or any other suitable network information. In some embodiments, the device information may include any suitable information about the user device or information about the manner in which the user device streams video content items. For example, in some embodiments, the device information may include the type or model of the user device, the operating system used by the user device, the current geographic location of the user device (e.g., as indicated by an IP address associated with the user device, as indicated by current GPS coordinates associated with the user device, and / or as indicated in any other suitable manner), the interface used to stream the video content item by the user device (e.g., a web browser, a particular media content streaming application running on the user device, and / or any other suitable interface information), the screen size or resolution of a display associated with the user device, the size of a viewport associated with a video player window in which the video content item should be presented on the user device, and / or any other suitable information.

[0026] At 106, process 100 may predict a preferred format for the video content item using the trained model and the network and / or device information received at block 104. In some embodiments, the format may include any suitable information, for example, a resolution of the frames of the video content item (e.g., 144 pixels by 256 pixels, 240 pixels by 426 pixels, 360 pixels by 640 pixels, 480 pixels by 854 pixels, 720 pixels by 1280 pixels, 1080 pixels by 1920 pixels, and / or any other suitable resolution).

[0027] In some embodiments, process 100 may predict a preferred format for a video content item in any suitable manner. For example, in some embodiments, process 100 may predict a preferred format for a video content item for streaming by a user device using a model trained with training data that indicates user-selected formats when streaming the video content item on different user devices and / or under different network conditions, as shown in and described below in connection with FIG. 2. In some embodiments, such a model may take user device information and network information as input, and may predict as output a format that is likely to be selected by a user given the input user device information and network information, as shown in and described below in connection with FIG.

[0028] As another example, in some embodiments, process 100 can predict a preferred format for a video content item for streaming by a user device using a model trained using training data that indicates quality metrics associated with streaming video content items using particular formats for different user devices and / or under different network conditions, as shown in FIG. 3 and described below in connection therewith. In some embodiments, such a model can take device information, network information, and a video content item format as input, and can predict as output a predicted quality score associated with streaming a video content item having the video content item format to a user device associated with the device information and network information. In some such embodiments, the model can then be used to predict a format likely to maximize the quality score. Note that in some embodiments, the quality score can include any suitable metric indicative of the quality of streaming of the video content item. For example, as described in more detail below in connection with FIG. 3, in some embodiments, the quality score can include a viewing time metric, which can indicate the average duration a video content item is viewed before presentation of the video content item is stopped. As another example, as described in more detail below in connection with FIG. 3, in some embodiments, the quality score may include an occupancy metric, which may be defined as (elapsed viewing time of the next video content item) / (elapsed viewing time of the next video content item + waiting time between presentation of the current video content item and the next video content item).

[0029] 2 and 3, the trained model may use any suitable features, other than those related to network information or device information, to predict a preferred format for a video content item. For example, in some embodiments, the trained model may use input features related to previous user actions (e.g., a pre-selected format by a user of a user device and / or any other suitable user action), information related to the user's user account used to stream video content (e.g., information related to the user's subscription to a video content streaming service, information related to the billing cycle for payment of the subscription to the video content streaming service, and / or any other suitable user account information), information related to the video content item (e.g., the genre or topic of the video content item, the length of the video content item, the popularity of the video content item, and / or any other suitable video content item information), and / or any other suitable information.

[0030] At 108, process 100 may select a format for the video content item based on a predicted preferred format for the video content item. In some embodiments, process 100 may select a format for the video content item in any suitable manner. For example, in some embodiments, process 100 may select a format for the video content item to be the same as the predicted preferred format. As a more specific example, in some embodiments, in cases where process 100 predicts the preferred format to be a particular resolution, process 100 may select the format as the particular resolution. As another example, in some embodiments, process 100 may select a format for the video content item based on the predicted preferred format and subject to any suitable rules or criteria. As a more specific example, in some embodiments, process 100 may select a format for the video content item subject to a rule dictating the maximum resolution that can be used based on the current viewport size of the video player window used by the user device. As a specific example, in an instance where the predicted preferred format is a particular resolution (e.g., 720 pixels by 1280 pixels) and the user device information received in block 104 indicates that the current size of the viewport is relatively small, process 100 may select a format as a lower resolution than the resolution associated with the predicted preferred format (e.g., 360 pixels by 640 pixels, 240 pixels by 426 pixels, and / or any other suitable lower resolution). As another more specific example, in some embodiments, process 100 may select a format for a video content item that is subject to rules dictating a maximum resolution for a particular type of video content (e.g., music video, lecture, documentary, etc.).As a specific example, in an instance where the predicted preferred format is a particular resolution (e.g., 720 pixels by 1280 pixels) and the video content item is determined to be a music video, process 100 may select a format as a lower resolution than the resolution associated with the predicted preferred format (e.g., 360 pixels by 640 pixels, 240 pixels by 426 pixels, and / or any other suitable lower resolution). As yet another more specific example, in some embodiments, process 100 may select a format for the video content item based on the location of video content items of different resolutions. As a specific example, in an instance where the predicted preferred format is a particular resolution (e.g., 720 pixels by 1280 pixels) and the video content item having the predicted preferred format is stored in a remote cache (e.g., more costly to access than a local cache), process 100 may select a format that corresponds to the version of the video content item stored in the local cache. It should be noted that in some embodiments, process 100 may select a format for a video content item based on any suitable rule that dictates, for example, a maximum or minimum resolution associated with any condition (e.g., network condition, device condition, user subscription condition, condition related to different types of video content, and / or any other suitable type of condition).

[0031] At 110, process 100 may transmit a first portion of the video content item having the selected format to the user device. Note that in some embodiments, the first portion of the video content item may have any suitable size (e.g., a particular number of kilobytes or megabytes, and / or any other suitable size) and / or may be associated with any suitable length of the video content item (e.g., 5 seconds, 10 seconds, and / or any other suitable length). In some embodiments, process 100 may calculate the size or length of the first portion of the video content item in any suitable manner, for example, based on a buffer size used by the user device to store the video content item while streaming the video content item. In some embodiments, process 100 may transmit the first portion of the video content item in any suitable manner. For example, in some embodiments, process 100 may use any suitable streaming protocol to transmit the first portion of the video content item. As another example, in some embodiments, process 100 may transmit the first portion of the video content item in association with an indication of conditions under which the user device should request additional portions of the video content item from a server. As a more specific example, in some embodiments, process 100 may transmit the first portion of the video content item in association with a minimum buffer amount at which the user device should request additional portions of the video content item. As yet another example, in some embodiments, process 100 may transmit the first portion of the video content item in association with instructions to transmit information from the user device to a server in response to detecting a change in the quality of the network connection (e.g., a change in the bandwidth of the network connection, a change in the throughput or goodput of the network connection, and / or any other suitable network connection change) or a change in the state of the device (e.g., a change in the size of the viewport of a video player window presenting the video content item, and / or any other suitable device state change).

[0032] In some embodiments, process 100 may return to block 104 and may receive updated network information (e.g., that the quality of the network connection has improved, that the quality of the network connection has decreased, and / or any other suitable updated network information) and / or updated device information (e.g., that the size of the viewport of a video player window used to present the video content item by the user device has changed, and / or any other suitable device information) from the user device. In some such embodiments, process 100 may loop through blocks 104-110, predict an updated preferred format based on the updated network information and device information, select an updated format based on the predicted updated preferred format, and transmit a second portion of the video content item having the selected updated format to the user device. In some embodiments, process 100 may loop back to block 104 in response to any suitable information and / or with any suitable frequency. For example, in some embodiments, process 100 may loop through blocks 104-110 with any suitable configured frequency such that additional portions of the video content item are transmitted to the user device at the predetermined frequency. As another example, in some embodiments, process 100 may loop back to block 104 in response to receiving a request from a user device for an additional portion of the video content item (e.g., a request sent from the user device in response to the user device determining that the amount of the video content item remaining in the user device's buffer is below a predetermined threshold, and / or a request sent from the user device in response to any other suitable information). As yet another example, in some embodiments, process 100 may loop back to block 104 in response to receiving information from the user device indicating a change in the state of streaming the video content item.It should be noted that in some embodiments, by looping through blocks 104-110, process 100 may cause the format of the video content item to be changed multiple times during presentation of the video content item by the user device.

[0033] 2 , an illustrative example 200 of a process for training a model to predict a video content item format based on a pre-selected format is shown in accordance with some embodiments of the disclosed subject matter. In some embodiments, the blocks of process 200 may be performed by any suitable device, such as a server that stores and / or streams video content items to user devices (e.g., a server associated with a video content sharing service, a server associated with a social networking service, and / or any other suitable server).

[0034] Process 200 may begin at 202 by generating a training set from previously streamed video content items, where each training sample in the training set indicates a user-selected format for the corresponding video content item. In some embodiments, each training sample may correspond to a video content item streamed to a particular user device. In some embodiments, each training sample may include any suitable information. For example, in some embodiments, the training sample may include information indicative of network conditions associated with the network connection over which the video content item was streamed to the user device (e.g., the throughput of the connection, the goodput of the connection, the bandwidth of the connection, the type of connection, the round-trip time (RTT) of the connection, a network identifier (e.g., an ASN, and / or any other suitable identifier), the bitrate used to stream the video content item, the Internet Service Provider (ISP) associated with the user device, and / or any other suitable network information). As another example, in some embodiments, the training sample may include device information associated with the user device used to stream the media content item (e.g., the model or type of the user device, the display size or resolution associated with the display of the user device, the geographic location of the user device, the operating system used by the user device, the viewport size of the video player window used to present the video content item, and / or any other suitable device information). As yet another example, in some embodiments, the training sample may include information related to the video content item to be streamed (e.g., the genre or category of the video content item, the length of the video content item, the resolution at which the video content item was uploaded to the server, the highest available resolution for the video content item, and / or any other suitable information related to the video content item).As yet another example, in some embodiments, the training sample may include information related to a user of a user device (e.g., information about the data plan used by the user, information about the billing cycle for video streaming subscriptions purchased by the user, the average or total duration of video content viewing by the user over any suitable duration, the total number of times a video content item is played by the user over any suitable duration, and / or any other suitable user information).

[0035] In some embodiments, each training sample may include a corresponding format for a video content item that was manually selected by a user while streaming the video content item. For example, in some embodiments, the corresponding format may include a resolution selected by a user for streaming the video content item. Note that in some embodiments, in cases where a user of a user device changes the resolution of a video content item while streaming the video content item, the format may correspond to a weighted resolution indicating, for example, an average resolution of different resolutions at which the video content item was presented, weighted by the duration for which each resolution was used. Furthermore, in cases where a user of a user device changes the resolution of a video content item while streaming the video content item, note that the training sample may include a history of user-selected resolutions for the video content item.

[0036] At 204, process 200 can use the training samples in the training set to train a model that produces as output a selected format for the video content item. That is, in some embodiments, the model may be trained to predict, for each training sample, a format that a user would manually select for streaming the video content item corresponding to the training sample. Note that in some embodiments, the trained model may generate the selected format in any suitable manner. For example, in some embodiments, the model may be trained to output a classification corresponding to a particular resolution associated with the selected format. As another example, in some embodiments, the model may be trained to output a continuous value corresponding to a target resolution, and the model may then subsequently generate the selected format by quantizing the target resolution value to an available resolution. It should be noted that in some embodiments, in cases where process 200 trains a model to generate classifications corresponding to a particular output resolution, the model may generate classifications from any suitable group of possible resolutions (e.g., 144 pixels by 256 pixels, 240 pixels by 426 pixels, 360 pixels by 640 pixels, 480 pixels by 854 pixels, 720 pixels by 1280 pixels, 1080 pixels by 1920 pixels, and / or any other suitable resolution).

[0037] In some embodiments, process 200 can train a model having any suitable architecture. For example, in some embodiments, process 200 can train a decision tree. As a more specific example, in some embodiments, process 200 can train a classification tree to generate a classification of resolutions to be used. As another example, in some embodiments, process 200 can train a regression tree to generate continuous values ​​indicating a target resolution, which can then be quantized to generate available resolutions for the video content item. In some embodiments, process 200 generates a decision tree model that partitions training set data based on any suitable feature using any suitable technique or combination of techniques. For example, in some embodiments, process 200 can use any suitable algorithm that identifies input features along which tree branches should be formed based on information gain, Gini impurity, and / or any other suitable metric. As another example, in some embodiments, process 200 can train a neural network. As a more specific example, in some embodiments, process 200 can train a multi-class perceptron to generate a classification of resolutions to be used. As another more specific example, in some embodiments, process 200 may train a neural network to output continuous values ​​of the target resolution, which may then be quantized to generate available resolutions for the video content item. Note that in some such embodiments, any suitable parameters may be used by process 200 to train the model, such as any suitable learning rate. Further, note that in some embodiments, the training set generated in block 202 may be split into a training set and a validation set, and these sets may be used to refine the model.

[0038] It should be noted that in some embodiments, the model may use any suitable input features and any suitable number of input features (e.g., 2, 3, 5, 10, 20, and / or any other suitable number). For example, in some embodiments, the model may use input features corresponding to a particular network quality metric (e.g., bandwidth of the network connection, throughput or goodput of the network connection, latency of the network connection, and / or any other suitable network quality metric), features corresponding to network type (e.g., Wi-Fi, 3G, 4G, Ethernet, and / or any other suitable type), features corresponding to device information (e.g., display resolution, device type or model, device operating system, device geographic location, viewport size, and / or any other suitable device information), features corresponding to user information (e.g., user subscription information, pre-selected format, and / or any other suitable user information), and / or features corresponding to video content information (e.g., video content item length, video content item popularity, video content item topic or genre). Further, it should be noted that in some embodiments, the input features may be selected in any suitable manner and using any suitable technique.

[0039] At 206, process 200 may receive network information and device information associated with a user device that has requested to stream the video content item. Note that in some embodiments, the user device may or may not be represented in the training set generated at block 202 as described above. As described above in connection with block 104 of FIG. 1 , in some embodiments, the network information may include any suitable information related to the network connection used by the user device to stream the video content item, e.g., the bandwidth of the connection, the latency of the connection, the throughput and / or goodput of the connection, the type of connection (e.g., 3G, 4G, Wi-Fi, Ethernet, and / or any other suitable type of connection), and / or any other suitable network information. As described above in connection with block 104 of FIG. 1 , in some embodiments, the device information may include any suitable information related to the user device, e.g., the model or type of the user device, the operating system used by the user device, the size or resolution of a display associated with the user device, the size of the viewport of a video player window in which the video content item should be presented, and / or any other suitable device information. Additionally, in some embodiments, process 200 may receive any suitable information related to the video content item (e.g., the topic or genre associated with the video content item, the length of the video content item, the popularity of the video content item, and / or any other suitable video content item information), information related to the user of the user device (e.g., billing information associated with the user, information related to the user's subscription to a video content streaming service that provides the video content item, information indicative of the user's previous playback history, and / or any other suitable user information), and / or information indicating a format or resolution of the video content item pre-selected by the user of the user device for streaming other video content items.

[0040] At 208, process 200 may use the trained model to select a format for streaming the video content item. In some embodiments, process 200 may use the trained model in any suitable manner. For example, in some embodiments, process 200 may generate inputs that correspond to input features used by the trained model and that are based on the network and device information received by process 200 in block 206, and may use the generated inputs to generate an output that indicates the resolution of the video content item. As a more specific example, in an instance where the trained model takes as inputs the throughput of the network connection, the geographic location of the user device, the size of the viewport of the video player window presented on the user device, and the operating system of the user device, process 200 may generate an input vector that includes the input information required by the trained model. Continuing with this example, process 200 may then use the trained model to generate as output a target resolution that represents the format that a user is likely to select when streaming the video content item under the conditions indicated by the input vector.

[0041] It should be noted that in instances where the trained model generates continuous values ​​corresponding to a target resolution (e.g., continuous values ​​representing the height and / or width of the target resolution), process 200 may quantize the continuous values ​​to correspond to available resolutions of the video content item. For example, in some embodiments, process 200 may select the available resolution of the video content item that is closest to the continuous value generated by the trained model. As a more specific example, in the case where the trained model generates a target resolution of 718 pixels by 1278 pixels, process 200 may determine that the closest available resolution is 720 pixels by 1280 pixels.

[0042] 3 , an illustrative example 300 of a process for training a model to predict a quality score associated with streaming a video content item having a particular format is shown in accordance with some embodiments of the disclosed subject matter. In some embodiments, the blocks of process 300 may be performed by any suitable device, such as a server that stores and / or streams video content items to user devices (e.g., a server associated with a video content sharing service, a server associated with a social networking service, and / or any other suitable server).

[0043] Process 300 may begin at 302 by generating a training set from previously streamed video content items, where each training sample is associated with a quality score that represents the quality of streaming of the video content item having a particular format. In some embodiments, each training sample may correspond to a video content item streamed to a particular user device. In some embodiments, each training sample may include any suitable information. For example, in some embodiments, a training sample may include information indicative of network conditions associated with the network connection when the video content item was streamed to the user device (e.g., the throughput of the connection, the goodput of the connection, the bandwidth of the connection, the type of connection, the round trip time (RTT) of the connection, a network identifier (e.g., an ASN, and / or any other suitable identifier), the bitrate used to stream the video content item, the Internet Service Provider (ISP) associated with the connection, and / or any other suitable network information). As another example, in some embodiments, the training sample may include device information associated with the user device used to stream the media content item (e.g., the model or type of the user device, the display resolution associated with the user device's display, the geographic location of the user device, the operating system used by the user device, the size of the viewport of the video player window on the user device through which the video content item is presented, and / or any other suitable device information). As yet another example, in some embodiments, the training sample may include information related to the video content item to be streamed (e.g., the genre or category of the video content item, the length of the video content item, the resolution at which the video content item was uploaded to the server, the highest available resolution for the video content item, and / or any other suitable information related to the video content item).As yet another example, in some embodiments, the training sample may include information related to a user of a user device (e.g., information about the data plan used by the user, information about the billing cycle for video streaming subscriptions purchased by the user, the average or total duration of video content viewing by the user over any suitable duration, the total number of times a video content item is played by the user over any suitable duration, and / or any other suitable user information).

[0044] In some embodiments, each training sample may include the format in which the video content item was streamed to the user device. For example, in some embodiments, the format may include the resolution of the video content item as it was streamed to the user device. Note that in cases where the resolution of the video content item was changed during presentation of the video content item, the resolution may be a weighted resolution corresponding to the average of the different resolutions of the video content item, weighted by the duration for which each resolution was used.

[0045] In some embodiments, each training sample may be associated with a corresponding quality score that indicates the quality of streaming the video content item at the resolution (or weighted resolution) indicated in the training sample. In some embodiments, the quality score may be calculated in any suitable manner. For example, in some embodiments, the quality score may be calculated based on a viewing time that indicates the duration for which the video content item was viewed. As a more specific example, in some embodiments, the quality score may be based on the total duration for which the video content item was viewed during a video content streaming session in which the training sample video content item was streamed. As a specific example, in the case where the training sample video content item was streamed during a video content streaming session that lasted 30 minutes, the quality score may be based on the 30-minute session viewing duration. As another more specific example, in some embodiments, the quality score may be based on the duration for which the training sample video content item was viewed or the percentage of the video content item that was viewed prior to ceasing presentation of the video content item. As a specific example, in a case where 50% of the video content items in the training sample are viewed by a user of a user device prior to ceasing presentation of the video content items, the quality score may be based on the 50% of the video content items viewed.

[0046] As another example, in some embodiments, a quality score may be calculated based on an occupancy score, which indicates the impact of the length of latency between the presentation of two video content items. In some embodiments, an example occupancy score that may be used to calculate a quality score is: Occupancy Score = (elapsed viewing time of next video content item) / (elapsed viewing time of next video content item + latency or interval between the current video content item and the next video content item). Note that in some embodiments, the video content item corresponding to the training sample may be either the current video content item referenced in the occupancy score metric or the next video content item referenced in the occupancy score metric.

[0047] It should be noted that in some embodiments, the quality score may be a function of a view time metric (e.g., the duration a video content item was viewed, the percentage of the video content item viewed, the duration a video content item was viewed during a video content streaming session, and / or any other suitable view time metric) or a function of an occupancy metric. In some embodiments, the function may be any suitable function. For example, in some embodiments, the function may be any suitable saturation function (e.g., a sigmoid function, a logistic function, and / or any other suitable saturation function).

[0048] At 304, process 300 can use the training set to train a model that predicts a quality score based on network information, device information, user information, and / or video content information. Similar to what was described above in connection with block 204 of FIG. 2 , in some embodiments, process 300 can train a model having any suitable architecture. For example, in some embodiments, process 300 can train a decision tree. As a more specific example, in some embodiments, process 300 can train a classification tree to generate classifications for quality scores within a particular range (e.g., within 0 to 0.2, within 0.21 to 0.4, and / or any other suitable range). As another example, in some embodiments, process 300 can train a regression tree that generates a continuous value indicating a predicted quality score. In some embodiments, process 300 can generate a decision tree model that partitions the training set data based on any suitable feature using any suitable technique or combination of techniques. For example, in some embodiments, process 300 may use any suitable algorithm that identifies input features along which tree branches should be formed based on information gain, gini impurity, and / or any other suitable metric.

[0049] As another example, in some embodiments, process 300 may train a neural network. As a more specific example, in some embodiments, process 300 may train a multi-class perceptron that generates classifications of quality scores as being within a particular range, as described above. As another more specific example, in some embodiments, process 300 may train a neural network to output continuous values ​​of predicted quality scores. It should be noted that in some such embodiments, any suitable parameters may be used by process 300 to train the model, such as any suitable learning rate. It should also be noted that in some embodiments, the training set generated in block 302 may be split into a training set and a validation set, and these sets may be used to refine the model.

[0050] At 306, process 300 may receive network information and device information associated with a user device that requested streaming of the video content item. Note that in some embodiments, the user device may or may not be represented in the training set generated at block 302 as described above. As described above in connection with block 104 of FIG. 1 , in some embodiments, the network information may include any suitable information related to the network connection used by the user device to stream the video content item, e.g., the bandwidth of the connection, the latency of the connection, the throughput and / or goodput of the connection, the type of connection (e.g., 3G, 4G, Wi-Fi, Ethernet, and / or any other suitable type of connection), and / or any other suitable network information. As described above in connection with block 104 of FIG. 1 , in some embodiments, the device information may include any suitable information related to the user device, e.g., the model or type of the user device, the operating system used by the user device, the size or resolution of a display associated with the user device, the size of the viewport of a video player window on the user device that should be used to present the video content item, and / or any other suitable device information.Additionally, in some embodiments, process 300 may receive any suitable information related to the video content item (e.g., the topic or genre associated with the video content item, the length of the video content item, the popularity of the video content item, and / or any other suitable video content item information), information related to the user of the user device (e.g., billing information associated with the user, information related to the user's subscription to a video content streaming service that provides the video content item, information indicative of the user's previous playback history, and / or any other suitable user information), and / or information indicating a format or resolution of the video content item pre-selected by the user of the user device for streaming other video content items.

[0051] At 308, process 300 may evaluate the trained model using network information, device information, user information, and / or video content item information to calculate a group of predicted quality scores corresponding to different formats in the group of possible formats. Note that in some embodiments, the group of possible formats may include any suitable format, such as different available resolutions of the video content item (e.g., 144 pixels by 256 pixels, 240 pixels by 426 pixels, 360 pixels by 640 pixels, 480 pixels by 854 pixels, 720 pixels by 1280 pixels, 1080 pixels by 1920 pixels, and / or any other suitable resolution). Note that in some embodiments, the group of possible formats may include any suitable number of possible formats (e.g., 1, 2, 3, 5, 10, and / or any other suitable number). An example of a group of predicted quality scores corresponding to a group of possible formats may be: [144 pixels x 256 pixels, 0.2; 240 pixels x 426 pixels, 0.4; 360 pixels x 640 pixels, 0.7; 480 pixels x 854 pixels, 0.5; 720 pixels x 1280 pixels, 0.2; 1080 pixels x 1920 pixels, 0.1], indicating the highest predicted quality score for a format corresponding to a resolution of 360 pixels x 640 pixels.

[0052] It should be noted that in some embodiments, process 300 may evaluate the trained model using any suitable combination of input features. For example, in some embodiments, the trained model may use any combination of input features including any suitable network quality metric (e.g., the bandwidth of the network connection, the latency of the network connection, the throughput or goodput of the network connection, and / or any other suitable network quality metric), device information (e.g., the type or model of the user device, the operating system used by the user device, the display size or resolution of a display associated with the user device, the viewport size of a video player window used to present the video content item, the geographic location of the user device, and / or any other suitable device information), user information (e.g., a pre-selected format, information indicative of subscriptions purchased by a user of the user device, and / or any other suitable user information), and / or video content item information (e.g., the genre or topic of the video content item, the length of the video content item, the popularity of the video content item, and / or any other suitable video content item information). In some embodiments, in instances where the trained model is a decision tree, process 300 may evaluate the trained model using input features selected during training of the decision tree. As a more specific example, the group of input features may include the network connection used by the user device, the operating system used by the user device, and the current goodput of the current geographic location used by the user device.

[0053] At 310, process 300 may select a format from the group of corresponding possible formats based on the predicted quality score for each possible format. For example, in some embodiments, process 300 may select the format from the group of possible formats that corresponds to the highest predicted quality score. As a more specific example, continuing with the above example of possible formats and corresponding predicted quality scores of [144 pixels by 256 pixels, 0.2; 240 pixels by 426 pixels, 0.4; 360 pixels by 640 pixels, 0.7; 480 pixels by 854 pixels, 0.5; 720 pixels by 1280 pixels, 0.2; 1080 pixels by 1920 pixels, 0.1], process 300 may select the format that corresponds to a resolution of 360 pixels by 640 pixels.

[0054] 4, an illustrative example of hardware for selecting a format for streaming a media content item 400 that may be used in accordance with some embodiments of the disclosed subject matter is shown. As shown, the hardware 400 may include a server 402, a communications network 404, and / or one or more user devices 406, such as user devices 408 and 410.

[0055] Server 402 may be any suitable server for storing information, data, programs, media content, and / or any other suitable content. In some embodiments, server 402 may perform any suitable function. For example, in some embodiments, server 402 may select a format for a video content item to be streamed to a user device using a trained model that predicts an optimal video content item format, as shown and described in FIG. 1 . In some embodiments, server 402 may train a model that predicts an optimal video content item format, as shown and described below in connection with FIG. 2 and FIG. 3 . For example, as shown and described below in connection with FIG. 2 , server 402 may train a model that predicts an optimal video content item format based on a format pre-selected by a user. As another example, as shown and described below in connection with FIG. 3 , server 402 may train a model that predicts an optimal video content item format based on quality scores associated with streaming a video content item in different formats.

[0056] The communications network 404, in some embodiments, may be any suitable combination of one or more wired and / or wireless networks. For example, the communications network 404 may include any one or more of the Internet, an intranet, a wide area network (WAN), a local area network (LAN), a wireless network, a digital subscriber line (DSL) network, a frame relay network, an asynchronous transfer mode (ATM) network, a virtual private network (VPN), and / or any other suitable communications network. The user device 406 may be connected by one or more communications links (e.g., communications link 412) to the communications network 404, which may be linked to the server 402 via one or more communications links (e.g., communications link 414). The communications link may be any communications link suitable for communicating data between the user device 406 and the server 402, such as a network link, a dial-up link, a wireless link, a hardwired link, any other suitable communications link, or any suitable combination of such links.

[0057] User device 406 may include any one or more user devices suitable for streaming media content. In some embodiments, user device 406 may include any suitable type of user device, such as a mobile phone, a tablet computer, a wearable computer, a laptop computer, a desktop computer, a smart television, a media player, a game console, a vehicle information and / or entertainment system, and / or any other suitable type of user device.

[0058] Although server 402 is shown as one device, the functions performed by server 402 may be performed using any suitable number of devices in some embodiments. For example, in some embodiments, multiple devices may be used to implement the functions performed by server 402.

[0059] To avoid overcomplicating the diagram, two user devices 408 and 410 are shown in FIG. 4, although in some embodiments, any suitable number and / or type of user devices may be used.

[0060] In some embodiments, the server 402 and the user device 406 may be implemented using any suitable hardware. For example, in some embodiments, the devices 402 and 406 may be implemented using any suitable general-purpose or special-purpose computer. For example, a mobile phone may be implemented using a special-purpose computer. Any such general-purpose or special-purpose computer may include any suitable hardware. For example, as shown in the example hardware 500 of FIG. 5, such hardware may include a hardware processor 502, memory and / or storage 504, an input device controller 506, input devices 508, a display / audio driver 510, display and audio output circuitry 512, a communication interface 514, an antenna 516, and a bus 518.

[0061] The hardware processor 502, in some embodiments, may include any suitable hardware processor, such as a microprocessor, a microcontroller, a digital signal processor, special purpose logic, and / or any other suitable circuitry for controlling the functions of a general-purpose or special-purpose computer. In some embodiments, the hardware processor 502 may be controlled by a server program stored in memory and / or storage of a server, such as the server 402. In some embodiments, the hardware processor 502 may be controlled by a computer program stored in memory and / or storage 504 of the user device 406.

[0062] The memory and / or storage 504, in some embodiments, may be any suitable memory and / or storage for storing programs, data, and / or any other suitable information. For example, the memory and / or storage 504 may include random access memory, read-only memory, flash memory, hard disk storage, optical media, and / or any other suitable memory.

[0063] Input device controller 506, in some embodiments, may be any suitable circuitry for controlling and receiving input from one or more input devices 508. For example, input device controller 506 may be circuitry for receiving input from a touchscreen, from a keyboard, from one or more buttons, from a voice recognition circuit, from a microphone, from a camera, from an optical sensor, from an accelerometer, from a temperature sensor, from a near-field sensor, from a pressure sensor, from an encoder, and / or from any other type of input device.

[0064] Display / audio driver 510, in some embodiments, may be any suitable circuitry for controlling and driving output to one or more display / audio output devices 512. For example, display / audio driver 510 may be circuitry for driving a touch screen, a flat panel display, a cathode ray tube display, a projector, one or more speakers, and / or any other suitable display and / or presentation device.

[0065] Communications interface 514 may be any suitable circuitry for interfacing with one or more communications networks (e.g., computer network 404). For example, interface 514 may include network interface card circuitry, wireless communications circuitry, and / or any other suitable type of communications network circuitry.

[0066] The antenna 516, in some embodiments, may be any suitable antenna or antennas for wirelessly communicating with a communication network (e.g., the communication network 404). In some embodiments, the antenna 516 may be omitted.

[0067] The bus 518 may be any suitable mechanism for communicating between two or more components 502, 504, 506, 510, and 514 in some embodiments.

[0068] According to some embodiments, any other suitable components may be included in hardware 500.

[0069] In some embodiments, at least some of the above-described blocks of the processes of Figures 1-3 may be executed or performed in any order or permutation, including but not limited to the order and permutation shown in and described with respect to the figures. Also, some of the above-described blocks of Figures 1-3 may be executed or performed substantially simultaneously, where appropriate, or in parallel to reduce latency and processing time. Additionally or alternatively, some of the above-described blocks of the processes of Figures 1-3 may be omitted.

[0070] In some embodiments, any suitable computer-readable medium may be used to store instructions for implementing the functions and / or processes herein. For example, in some embodiments, the computer-readable medium may be transitory or non-transitory. For example, a non-transitory computer-readable medium may include media such as magnetic media in non-transitory form (such as hard disks, floppy disks, and / or any other suitable magnetic medium), optical media in non-transitory form (such as compact discs, digital video discs, Blu-ray discs, and / or any other suitable optical medium), semiconductor media in non-transitory form (such as flash memory, electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), and / or any other suitable semiconductor medium), any suitable medium that is not transient or does not lose apparent permanence during transmission, and / or any suitable tangible medium. As another example, a transitory computer-readable medium may include signals on a network, in wires, in conductors, in optical fibers, in circuits, in any suitable medium that is transient and loses any apparent permanence during transmission, and / or in any suitable non-tangible medium.

[0071] Accordingly, methods, systems, and media are provided for selecting a format for streaming a media content item.

[0072] While the present invention has been described and shown in the foregoing illustrative embodiments, it will be understood that the present disclosure is made by way of example only, and that numerous changes in the details of the implementation of the invention may be made without departing from the spirit and scope of the invention, which is limited only by the following claims. The features of the disclosed embodiments may be combined and rearranged in various ways. [Explanation of symbols]

[0073] 400 Hardware 402 Server 404 Communication Networks, Computer Networks 406 User Device 408 User Device 410 User Device 412 Communication Links 414 Communication Links 500 Hardware 502 Hardware Processors, Components 504 Memory and / or Storage, Components 506 Input device controller, component 508 Input Devices 510 Display / Audio Driver, Components 512 Display and audio output circuit configuration, display / audio output device 514 Communication Interfaces, Interfaces, and Components 516 Antenna 518 Bus

Claims

1. 1. A method for selecting a format for streaming a video content item, comprising: receiving, at a server, a request from a user device to begin streaming a video content item on said user device; receiving, from the user device, network information indicative of a quality of a network connection used by the user device to stream the video content item and device information associated with the user device; predicting, by the server, a preferred format for the video content item using a trained model, the network information, and the device information; selecting, by the server, a first format for the video content item, comprising identifying, by the server, a format that maximizes a predicted duration for which a user of the user device will view the video content item, the first format comprising a first resolution of a plurality of resolutions based on the predicted preferred format for the video content item; transmitting, from the server, a first portion of the video content item having the first format to the user device; receiving updated network information and updated device information from the user device at the server; determining, by the server, a second format for the video content item using the trained model, the updated network information, and the updated device information; transmitting, from the server, a second portion of the video content item having the second format to the user device.

2. The method of claim 1 , wherein the predicted preferred format for the video content item is a format likely to be selected by a user of the user device.

3. The method described in claim 1, wherein the updated device information includes an indication that the size of the viewport of a video player window running on the user device on which the video content item is presented has changed.

4. The method of claim 3 , wherein the updated device information indicates that the size of the viewport has decreased, and a second resolution of the plurality of resolutions is lower than the first resolution.

5. The method of claim 1 , wherein the first format for the video content item is selected based on a genre of the video content item.

6. The method of claim 1 , wherein the first format for the video content item is selected based on a format pre-selected by a user of the user device for streaming other video content items from the server.

7. 1. A system for selecting a format for streaming a video content item, comprising: Memory and a hardware processor, the hardware processor executing the computer-executable instructions stored in the memory to: receiving, at a server, a request from a user device to begin streaming a video content item on the user device; receiving, from the user device, network information indicative of a quality of a network connection used by the user device to stream the video content item and device information associated with the user device; predicting, by the server, a preferred format for the video content item using a trained model, the network information, and the device information; selecting, by the server, a first format for the video content item, including identifying, by the server, a format that maximizes a predicted duration for which a user of the user device will view the video content item, the first format including a first resolution of a plurality of resolutions based on the predicted preferred format for the video content item; transmitting, from the server, a first portion of the video content item having the first format to the user device; receiving updated network information and updated device information from the user device at the server; determining, by the server, a second format for the video content item using the trained model, the updated network information, and the updated device information; transmitting, from the server, a second portion of the video content item having the second format to the user device. system.

8. The system of claim 7 , wherein the predicted preferred format for the video content item is a format likely to be selected by a user of the user device.

9. The system described in claim 7, wherein the updated device information includes an indication that the size of the viewport of a video player window running on the user device on which the video content item is presented has changed.

10. 10. The system of claim 9, wherein the updated device information indicates that the size of the viewport has decreased, and a second resolution of the plurality of resolutions is lower than the first resolution.

11. The system of claim 7 , wherein the first format for the video content item is selected based on a genre of the video content item.

12. 8. The system of claim 7, wherein the first format for the video content item is selected based on a format pre-selected by a user of the user device for streaming other video content items from the server.

13. 1. A computer-readable storage medium containing computer-executable instructions that, when executed by a processor, cause the processor to perform a method for selecting a format for streaming a video content item, the method comprising: receiving, at a server, a request from a user device to begin streaming a video content item on said user device; receiving, from the user device, network information indicative of a quality of a network connection used by the user device to stream the video content item and device information associated with the user device; predicting, by the server, a preferred format for the video content item using a trained model, the network information, and the device information; selecting, by the server, a first format for the video content item, comprising identifying, by the server, a format that maximizes a predicted duration for which a user of the user device will view the video content item, the first format comprising a first resolution of a plurality of resolutions based on the predicted preferred format for the video content item; transmitting, from the server, a first portion of the video content item having the first format to the user device; receiving updated network information and updated device information from the user device at the server; determining, by the server, a second format for the video content item using the trained model, the updated network information, and the updated device information; transmitting, from the server, a second portion of the video content item having the second format to the user device; A computer-readable storage medium.

14. 14. The computer-readable storage medium of claim 13, wherein the predicted preferred format for the video content item is a format likely to be selected by a user of the user device.

15. A computer-readable storage medium as described in claim 13, wherein the updated device information includes an indication that the size of the viewport of a video player window running on the user device on which the video content item is presented has changed.

16. 16. The computer-readable storage medium of claim 15, wherein the updated device information indicates that the size of the viewport has decreased, and a second resolution of the plurality of resolutions is lower than the first resolution.

17. The computer-readable storage medium of claim 13 , wherein the first format for the video content item is selected based on a genre of the video content item.

18. 14. The computer-readable storage medium of claim 13, wherein the first format for the video content item is selected based on a format pre-selected by a user of the user device for streaming other video content items from the server.

Citation Information

Patent Citations

  • System and method of data transmission reception, data receiver and data reception method

    JP1999127150A

  • Video distribution system, client terminal and control method thereof

    JP2007151078A

  • Transmission apparatus, transmission method, and reception apparatus

    JP2009290691A

  • Distributed architecture for encoding and delivering video content

    JP2015523805A

  • url parameter insertion and addition in adaptive streaming

    JP2016511954A