Using Bayesian inference to predict review decisions in matching graphs
Automatic prediction of the properties of media projects through Bayesian inference and machine learning models, solving the problems of high resource consumption and manual review of newly uploaded media projects, achieving more efficient media project processing and improved user experience.
Patent Information
- Application Number
- CN201980087180.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-12-31
- Filing Date
- 2019-02-19
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2039-02-28
AI Technical Summary
In the prior art, newly uploaded media items require manual review to determine whether they contain inappropriate content or technical issues, resulting in large processor overhead, high bandwidth requirements and long time consuming, and manual marking may result in errors requiring additional corrections.
By using Bayesian inference and machine learning models, the properties of media projects are automatically predicted, segment prediction values are generated, and media project prediction values are calculated, based on these values, whether to allow playback or further processing is performed, reducing resource consumption for newly uploaded media projects.
Improves efficiency in handling newly uploaded media projects, reduces processor overhead and bandwidth requirements, provides faster media project processing and a better user experience, while reducing the need for manual reviews.
Smart Images

Figure CN113272800B_ABST
Abstract
Description
Technical Field
[0001] Aspects and implementations of the present disclosure relate to predicting review decisions, and in particular, predicting review decisions for media projects. Background Art
[0002] Media projects such as video projects, audio projects, etc. can be uploaded to a media project platform. The media projects can be tagged based on the content type of the media project, the suitability of the media project, the quality of the media project, and the like. Summary of the Invention
[0003] The following is a brief summary of the present disclosure to provide a basic understanding of some aspects of the present disclosure. This summary is not an extensive overview of the present disclosure. It is neither intended to identify key or important elements of the present disclosure nor to depict any scope of a particular implementation of the present disclosure or any scope of the claims. Its sole purpose is to present some concepts of the present disclosure in a simplified form as a prelude to the more detailed description presented later.
[0004] Aspects of the present disclosure automatically determine attributes of a media project. Bayesian inference can be used to predict the attributes of a media project in a matching graph.
[0005] In one aspect of the present disclosure, a method may include identifying a current media project to be processed and processing a plurality of tagged media projects to identify tagged media projects that include at least one corresponding segment that is one of a plurality of segments similar to the current media project. The method may further include, for each of the plurality of segments of the current media project, generating a segment prediction value indicative of a particular attribute associated with the corresponding segment of the current media project based on an attribute associated with the corresponding tagged media project, each tagged media project including a corresponding segment similar to the corresponding segment of the current media project. The method may further include calculating a media project prediction value for the current media project based on the generated segment prediction values for each of the plurality of segments of the current media project and causing the current media project to be processed based on the calculated media project prediction value.
[0006] Each of the tagged media items can be assigned a corresponding tag; and each of the multiple segments of the current media item can at least partially match one or more segments of the tagged media items. Generating a segment prediction value for each of the multiple segments of the current media item can be based on multiple parameters, the multiple parameters including at least one of the length of the corresponding segment, the length of the current media item, or the length of the corresponding tagged media item, each tagged media item including a corresponding segment similar to the corresponding segment of the current media item. The method can further include: determining that a first segment of the current media item matches a first corresponding segment of a first tagged media item; determining that a second segment of the current media item matches a second corresponding segment of a second tagged media item; determining that the first segment is a sub-segment of the second segment, where the second segment includes the first segment and a third segment, where generating a segment prediction value for each of the multiple segments includes: generating a first segment prediction value indicating a first tag for the first segment based on the first corresponding tag of the first tagged media item and the second corresponding tag of the second tagged media item; and generating a second segment prediction value indicating a second tag for the third segment based on the second corresponding tag of the second tagged media item, where calculating a media item prediction value is based on the generated first segment prediction value and the generated second segment prediction value. Causing the current media item to be processed can include one of the following: causing the current media item to be blocked from playback via a media item platform in response to the calculated media item prediction value satisfying a first threshold condition; causing the current media item to be allowed to be played via a media item platform in response to the calculated media item prediction value satisfying a second threshold condition; and causing the current media item to be reviewed to generate a tag indicating whether playback of the current media item via the media item platform is allowed in response to the calculated media item prediction value satisfying a third threshold condition. The segment prediction value for each of the multiple segments is generated based on multiple parameters and one or more weights associated with one or more of the multiple parameters. The method can further include adjusting one or more weights based on the generated tag of the current media item. Adjusting one or more weights includes: training a machine learning model to provide adjusted one or more weights based on a tuning input and a target tuning output for the tuning input; for each of the multiple segments of the current media item, the tuning input includes the length of the corresponding segment, the length of the current media item, and the length of the corresponding tagged media item, each corresponding tagged media item including a corresponding segment similar to the corresponding segment of the current media item; and the tuning target output for the tuning input includes the generated tag for the current media item.
[0007] It will be understood that aspects may be implemented in any convenient form. For example, aspects may be implemented by a suitable computer program which may be carried on a suitable carrier medium which may be a tangible carrier medium (such as a disk) or an intangible carrier medium (such as a communication signal). Aspects may also be implemented using a suitable apparatus which may take the form of a programmable computer running a computer program arranged to implement the invention. Aspects may be combined such that features described in the context of one aspect may be implemented in another aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] In the drawings, the present disclosure is illustrated by way of example and not limitation.
[0009] Figure 1 is a block diagram showing an exemplary system architecture according to an implementation of the present disclosure.
[0010] Figure 2 is an exemplary tuning set generator for creating tuning data for a machine learning model according to an implementation of the present disclosure.
[0011] Figure 3A -D is a flowchart showing an exemplary method of predicting a review decision for a media item according to an implementation of the present disclosure.
[0012] Figure 4A -B is a table showing a review decision for predicting a media item according to an implementation of the present disclosure.
[0013] Figure 5 is a block diagram showing an implementation of a computer system according to an implementation of the present disclosure.
[0014] DETAILED IMPLEMENTATION
[0015] Aspects and implementations of the present disclosure relate to using Bayesian inference to automatically determine attributes of a media item, such as predicting a review decision in a matching graph. A server device may receive a media item uploaded by a user device. The server device may make the media item playable by these or other user devices via a media item platform. In response to a user labeling a media item (e.g., during playback) as having a certain type of content, the server device may submit the media item for review (e.g., manual review). During the review, a user (e.g., an administrator of the media item platform) may play the media item and may mark the media item. Based on the labels from the review, the server may associate the media item with the labels. The labels may indicate that the media item has an inappropriate content type, infringes rights, has a technical problem, has a rating, is suitable for advertising, etc.
[0016] Server devices can perform different actions on media items based on the tags of the media items. For example, playback of media items tagged as containing inappropriate content types (e.g., having a "negative review" tag) can be prevented. Alternatively, playback of media items tagged as not containing inappropriate content types (e.g., having a "positive review" tag) can be allowed. In another example, based on the tags of the media item, advertisements can be included during the playback of the tag-based media item. In yet another example, a message can be transmitted to a user device that has uploaded a media item tagged as having a technical problem.
[0017] Newly uploaded media items can at least partially match tagged media items. Typically, each newly uploaded media item is typically labeled, submitted for manual review, and tagged based on the manual review (e.g., regardless of whether the media item matches a tagged media item). Typically, reviewing a first media item that at least partially matches a tagged media item may require the same amount of time and resources (e.g., processor overhead, bandwidth, power consumption, available human reviewers, etc.) as reviewing a second media item that at least partially does not match a tagged media item. Reviewing media items before allowing them to be available for playback can take a long time and may require peak required resources (e.g., high processor overhead, power consumption, and bandwidth). Allowing media items to be available for playback before reviewing them may allow the playback of media items that include problems and inappropriate content. Incorrectly tagging media items may require time and resources to correct.
[0018] Aspects of the present disclosure address the above and other challenges by automatically determining attributes of media items (e.g., using Bayesian inference to predict review decisions in a match graph). A processing device may identify a current media item to be processed (e.g., a newly uploaded media item), and may identify tagged media items (e.g., previously reviewed media items) to find tagged media items that include at least one corresponding segment similar to one of the segments of the current media item (e.g., a match graph may be created that includes tagged media items that match at least a portion of the current media item). For each of the segments of the current media item (e.g., segments similar to at least a portion of the tagged media item), the processing device may generate a segment prediction value indicative of a particular attribute associated with the corresponding segment of the current media item (e.g., based on the attributes associated with the tagged media items having corresponding segments similar to the corresponding segment of the current media item, based on the match graph). The processing device may calculate a media item prediction value based on the generated segment prediction values for each of the segments of the current media item, and may cause the current media item to be processed based on the calculated media item prediction value. For example, based on the media item prediction value, playback of the current media item via a media item platform may be allowed or blocked, or the current media item may be reviewed to generate a label for the current media item that indicates whether playback of the current media item is allowed.
[0019] As disclosed herein, automatically generating attributes of media items (e.g., predicting review decisions) is advantageous because it improves the user experience and provides technical advantages. Many newly uploaded media items can have segments similar to tagged (e.g., previously reviewed) media items. By performing an initial process to select media items for further processing, attributes can be assigned to media items more efficiently, and fewer media items may be required for further processing. Thus, processing newly uploaded media items based on the predicted values of the media items (e.g., based on tagged media items that at least partially match the newly uploaded media item) can reduce processor overhead, required bandwidth, and energy consumption compared to performing the same process to tag any media item, regardless of whether the media item at least partially matches a previous tagged media item. Allowing or blocking the playback of newly uploaded media items based on the predicted values of the media items can allow for faster processing of the media items and thus provide a better user experience compared to providing all newly uploaded media items for playback and only blocking playback after the user has flagged the media item and subsequent manual review has tagged the media item. Generating predicted values of media items for uploaded media items can be beneficial to the users who upload the media items and to the users of the media item platform. For example, in response to generating a predicted value of a media item indicating a technical issue, an indication based on the predicted value of the media item can be transmitted to the user who uploaded the media item to alert the user that the media item has a technical issue (e.g., suggesting modifications to the media item). In another example, in response to generating a predicted value of a media item, the media item can be processed such that the media item is tagged (e.g., age appropriateness, category, etc.) to improve search results and recommendations for users of the media item platform.
[0020] Figure 1 FIG. 100 shows an exemplary system architecture 100 in accordance with an implementation of the present disclosure. The system architecture 100 includes a media item server 110, a user device 120, a prediction server 130, a content owner device 140, a network 150, and a data store 160. The prediction server 130 can be part of a prediction system 105.
[0021] The media item server 110 may include one or more computing devices (such as rack servers, router computers, server computers, personal computers, mainframe computers, laptop computers, tablet computers, desktop computers, etc.), data storage (e.g., hard disks, memories, databases, etc.), networks, software components, and / or hardware components. The media item server 110 may be used to provide users with access to media items 112 (e.g., tagged media items (“tagged media items” 114), media items currently undergoing a tagging process (“current media items” 116), etc.). The media item server 110 may provide media items 112 to users (e.g., users may select media items 112 and download media items 112 from the media item server 110 in response to a request or purchase of the media items 112). The media item server 110 may be part of a media item platform (e.g., a content hosting platform that provides a content hosting service), which may allow users to consume, develop, upload, download, rate, label, share, search, like, dislike, and / or comment on media items 112. The media item platform may also include a website (e.g., a web page) or application backend software, which may be used to provide users with access to media items 112.
[0022] The media item server 110 may host content, such as media items 112. The media items 112 may be digital content selected by users, digital content available to users, digital content developed by users, digital content uploaded by users, digital content developed by content owners, digital content uploaded by content owners, digital content provided by the media item server 110, etc. Examples of media items 112 include but are not limited to video items (e.g., digital videos, digital movies, etc.), audio items (e.g., digital music, digital audiobooks, etc.), advertisements, slide shows that switch slides over time, text that scrolls over time, graphs that change over time, etc.
[0023] Media item 112 can be consumed via a web browser on user device 120 or via a mobile application (“app”) that can be installed on user device 120 via an app store. The web browser or mobile app can allow a user to perform one or more searches (e.g., for explanatory information, for other media items 112, etc.). As used herein, “application,” “mobile application,” “smart TV application,” “desktop application,” “software application,” “digital content,” “content,” “content item,” “media,” “media item,” “video item,” “audio item,” “contact invitation,” “game,” and “advertisement” can include an electronic file that can be executed or loaded using software, firmware, or hardware configured to present media item 112 to an entity. In one implementation, a media item platform can use data store 160 to store media item 112. Media item 112 can be presented to a user of user device 120 or downloaded by a user of user device 120 from media item server 110 (e.g., a media item platform such as a content hosting platform). Media item 112 can be played via an embedded media player (and other components) provided by the media item platform or stored locally. The media item platform can be, for example, an app distribution platform, a content hosting platform, or a social networking platform, and can be used to provide a user with access to media item 112 or to provide media item 112 to a user. For example, the media item platform can allow a user to consume, label, upload, search, approve (“like”), dislike, and / or comment on media item 112. Media item server 110 can be part of the media item platform, can be a stand-alone system, or can be part of a different platform.
[0024] Network 150 can be a public network that provides user device 120 and content owner device 140 with access to media item server 110, prediction server 130, and other publicly available computing devices. Network 150 can include one or more wide area networks (WANs), local area networks (LANs), wired networks (e.g., Ethernet networks), wireless networks (e.g., 802.11 networks or Wi-Fi networks), cellular networks (e.g., Long Term Evolution (LTE) networks), routers, hubs, switches, server computers, and / or combinations thereof.
[0025] The data storage 160 can be a memory (e.g., random access memory), a drive (e.g., hard disk drive, flash drive), a database system, or another type of component or device capable of storing data. The data storage 160 can include multiple storage components (e.g., multiple drives or multiple databases), which can span multiple computing devices (e.g., multiple server computers). In some implementations, the data storage 160 can store information 162 associated with media items, tags 164, or predicted values 166 (e.g., segment predicted values 168, media item predicted values 169). Each of the tagged media items 114 can have a corresponding tag 164. Each of the tagged media items 114 may have been labeled (e.g., by a user during playback, via a machine learning model trained with image input and tag output). Each of the media items 112 can have a corresponding information 162 (e.g., total length, length of each segment, media item identifier, etc.).
[0026] The user device 120 and the content owner device 140 can include computing devices such as personal computers (PCs), laptop computers, mobile phones, smartphones, tablet computers, netbook computers, network-connected televisions (“smart TVs”), network-connected media players (e.g., Blu-ray players), set-top boxes, over-the-top (OTT) streaming devices, operation boxes, etc.
[0027] Each user device 120 can include an operating system that allows the user to perform playback of media items 112 and label the media items 112. The media items 112 can be presented via a media viewer or a web browser. The web browser can access, retrieve, present, and / or navigate content served by a web server (e.g., web pages such as hypertext markup language (HTML) pages, digital media items, text conversations, notifications, etc.). An embedded media player (e.g., a player or an HTML5 player) can be embedded in a web page (e.g., to provide information about products sold by an online merchant), or can be part of a media viewer (mobile application) installed on the user device 120. In another example, the media items 112 can be presented via a stand-alone application (e.g., a mobile application or an app) that allows the user to view digital media items (e.g., digital video, digital audio, digital images, etc.).
[0028] The user device 120 can include one or more of a playback component 124, a labeling component 126, and a data storage 122. In some implementations, one or more of the playback component 124 or the labeling component 126 can be provided by a web browser or an application (e.g., a mobile application, a desktop application) executing on the user device 120.
[0029] The data store 122 can be a memory (e.g., random access memory), a drive (e.g., hard disk drive, flash drive), a database system, or another type of component or device capable of storing data. The data store 122 can include multiple storage components (e.g., multiple drives or multiple databases), which can span multiple computing devices (e.g., multiple server computers). The data store 122 can include a media item cache 123 and an annotation cache 125.
[0030] The playback component 124 can provide playback of the media item 112 via the user device 120. The playback of the media item 112 can be in response to the playback component 124 receiving a user input to play the media item 112 (via a graphical user interface (GUI) displayed on the user device) and transmitting the request to the media item server 110. In some implementations, the media item server 110 can stream the media item 112 to the user device 120. In some implementations, the media item server 110 can transmit the media item 112 to the user device 120. The playback component 124 can store the media item 112 in the media item cache 123 for playback at a later time point (e.g., subsequent playback regardless of connectivity to the network 150).
[0031] The annotation component 126 can receive user input (e.g., via the GUI, during playback of the media item 112) to annotate the media item 112. The media item 112 can be annotated as one or more types of content. The user input can indicate the type of the content (e.g., by selecting the type of content from a list). Based on the user input, the annotation component 126 can annotate the media item 112 as having an inappropriate content type (e.g., including one or more sexual content, violence or offensive content, hate or abusive content, harmful or dangerous behavior, child abuse, terrorism promotion, spam or misleading, etc.), infringing rights, having a technical problem (e.g., subtitle problem, etc.), having a rating (e.g., age-appropriateness rating, etc.), being suitable for advertising, etc. The annotation component 126 can transmit an indication that the media item 112 has been annotated to one or more of the prediction system 105, the prediction server 130, the data store 160, etc.
[0032] In response to being annotated, the media item 112 can be marked (e.g., by manual review) to generate a marked media item 114. The information 162 and the label 164 associated with the marked media item 114 can be stored in the data store 160 together with the marked media item 114. Alternatively, the marked media item 114 can be stored in a separate data store and associated with the information 162 and the label 164 via a media item identifier.
[0033] The content owner device 140 may include a transmission component 144, a receiving component 146, a modification component 148, and a data store 142.
[0034] The data store 142 may be a memory (e.g., random access memory), a drive (e.g., hard disk drive, flash drive), a database system, or another type of component or device capable of storing data. The data store 142 may include multiple storage components (e.g., multiple drives or multiple databases), which may span multiple computing devices (e.g., multiple server computers). The data store 142 may include a media item cache 143.
[0035] The transmission component 144 may receive media items 112 created, modified, uploaded, or associated by a content owner corresponding to the content owner device 140. The transmission component 144 may store the media items 112 in the media item cache 143. The transmission component 144 may transmit (e.g., upload) the media items 112 to the media item server 110 (e.g., in response to a content owner input to upload the media items 112).
[0036] The receiving component 146 may receive an indication based on a media item prediction value 169 (e.g., generated by a prediction manager 132) from the prediction server 130. The receiving component 146 may store the indication in the data store 142.
[0037] The modification component 148 may modify the media item 112 based on the indication based on the media item prediction value 169. For example, in response to an indication based on the media item prediction value 169 that indicates that the content of the media item 112 is inappropriate or has a technical problem (e.g., an error in the subtitles, etc.), the modification component 148 may cause the content of the media item 112 to be modified (e.g., remove inappropriate content, fix technical problems). In some implementations, in order to cause the content to be modified, the modification component 148 may provide an indication or recommendation to the content owner via the GUI on how to modify the content. In some implementations, in order to cause the content to be modified, the modification component 148 may automatically modify the content (e.g., fix technical problems, remove inappropriate content, etc.).
[0038] The prediction server 130 can be coupled to the user device 120 and the content owner device 140 via the network 150, facilitating the prediction of review decisions for media item 112. In one implementation, the prediction server 130 can be part of a media item platform (e.g., the media item server 110 and the prediction server 130 can be part of the same media item platform). In another implementation, the prediction server 130 can be an independent platform, including one or more computing devices, such as rack servers, router computers, server computers, personal computers, mainframe computers, laptop computers, tablet computers, desktop computers, etc., and can provide review decision prediction services that can be used by the media item platform and / or various other platforms (e.g., social networking platforms, online news platforms, messaging platforms, video conferencing platforms, online meeting platforms, etc.).
[0039] The prediction server 130 can include a prediction manager 132. According to some aspects of the present disclosure, the prediction manager 132 can identify a current media item 116 to be processed (e.g., a newly uploaded media item) (e.g., automatically evaluated without user input based on an indication provided by, for example, the tagging component 126), and can process the tagged media item 114 to find at least one corresponding segment of the tagged media item 114 that includes one of the segments similar to the current media item (e.g., generate a match graph of the tagged media item 114). For each segment of the current media item 116 (e.g., at least partially matching the tagged media item 114), the prediction manager 132 can generate a segment prediction value 168 indicating a specific attribute associated with the corresponding segment of the current media item 116 (e.g., based on the attributes associated with the corresponding tagged media item 114 that teaches the corresponding segment including the corresponding segment similar to the current media item 116, based on the match graph). The prediction manager 132 can calculate a media item prediction value 169 for the current media item 116 based on the generated segment prediction values 168 for each segment of the current media item 116, and can process the current media item 116 based on the calculated media item prediction value 169. In some implementations, the prediction manager (e.g., via a trained machine learning model 190, without a trained machine learning model 190) can use Bayesian inference to predict the review decision in the match graph (e.g., the media item prediction value 169) (e.g., based on the tagged media item 114 that at least partially matches the current media item 116, based on Figure 4A Table 400A, etc.).
[0040] In response to a media item prediction value 169 (e.g., indicating that the current media item 116 has inappropriate content or has a technical issue), the prediction manager 132 can cause the current media item 116 to be blocked from being available for playback via the media item platform. The prediction manager 132 can transmit an indication to the media item platform (or any other platform) or directly to the content owner device 140 based on the media prediction value 169, indicating the determined attributes of the current media item 116. The content owner device 140 can receive the indication (e.g., from the media item platform or from the prediction manager 132) such that the current media item 116 is modified based on the indication and the modified current media item is re-uploaded (e.g., transmitted to the media item server 110, the prediction server 130, or the media item platform).
[0041] In some implementations, the prediction manager 132 can use a trained machine learning model 190 to determine the media item prediction value 169. The prediction system 105 can include one or more of the prediction server 130, the server machine 170, or the server machine 180. The server machines 170 - 180 can be one or more computing devices (such as rack servers, router computers, server computers, personal computers, mainframe computers, laptop computers, tablet computers, desktop computers, etc.), data storage (e.g., hard disks, memories, databases, etc.), networks, software components, or hardware components.
[0042] The server machine 170 includes a tuning set generator 171 that can generate tuning data (e.g., a set of tuning inputs and a set of target outputs) to train the machine learning model. The server machine 180 includes a tuning engine 181 that can use the tuning data from the tuning set generator 171 to train the machine learning model 190. The machine learning model 190 can use the (see Figure 2 and Figure 3C)Tune the input 210 and the target output 220 (e.g., the target tuned output) for training. Then, the trained machine learning model 190 can be used to determine the media item prediction value 169. The machine learning model 190 can refer to a model artifact created by the tuning engine 181 using tuning data that includes the tuning input and the corresponding target output (the correct answer for the corresponding tuning input). Patterns (correct answers) that map the tuning input to the target output can be found in the tuning data, and the machine learning model 190 that captures these patterns is provided. The machine learning model 190 can include parameters 191 (e.g., k, f, g, configuration parameters, hyperparameters, etc.), which can be tuned (e.g., the associated weights can be adjusted) based on subsequent tokens of the current media item 116 for which the media item prediction value 169 is determined. In some implementations, the parameters 191 include one or more of k, f, or g, as described below. In some implementations, the parameters include at least one of the length of the corresponding segment of the current media item 116, the length of the current media item 116, or the length of the corresponding tokenized media item 114 (e.g., each includes a corresponding segment similar to the corresponding segment of the current media item 116).
[0043] The prediction manager 132 can determine information associated with the current media item 116 (e.g., the length of the current media item 116, the length of the segment of the current media item 116 that is similar to the corresponding segment of the tokenized media item 114) and information associated with the tokenized media item 114 (e.g., the length of the tokenized media item 114), and provide this information to the trained machine learning model 190. The trained machine learning model 190 can produce an output, and the prediction manager 132 can determine the media item prediction value 169 based on the output of the trained machine learning model. For example, the prediction manager 132 can extract the media item prediction value 169 from the output of the trained machine learning model 190, and can extract confidence data from the output that indicates the confidence level that the current media item 116 contains the content type indicated by the media item prediction value 169 (e.g., the confidence level that the media item prediction value 169 accurately predicts the manual review decision).
[0044] In an implementation, the confidence data can include or indicate the confidence level that the current media item 116 has a particular attribute (e.g., contains a content type). In one example, the confidence level is a real number between 0 and 1, where 0 indicates disbelief that the current media item 116 has a particular type of content, and 1 indicates absolute belief that the current media item 116 has a particular type of content.
[0045] As described above, for purposes of illustration and not limitation, aspects of the present disclosure describe using information 162 associated with media item 112 to train machine learning model 190 and using the trained machine learning model 190. In other implementations, a heuristic model or a rule-based model is used to determine media item prediction value 169 for current media item 116. It may be noted that any information described with reference to Figure 2 the tuning input 210 can be monitored or otherwise used in a heuristic or rule-based model.
[0046] It should be noted that in some other implementations, the functionality of server machine 170, server machine 180, prediction server 130, or media item server 110 can be provided by a smaller number of machines. For example, in some implementations, server machines 170 and 180 can be integrated into a single machine, and in some other implementations, server machines 170, server machine 180, and prediction server 130 can be integrated into a single machine. Additionally, in some implementations, one or more of server machines 170, server machine 180, and prediction server 130 can be integrated into media item server 110.
[0047] Generally speaking, if appropriate, functionality described in one implementation as being performed by media item server 110, server machine 170, server machine 180, or prediction server 130 can also be performed on user device 120 in other implementations. For example, media item server 110 can stream media item 112 to user device 120 and can receive user input indicating an indication of media item 112.
[0048] Generally speaking, if appropriate, functionality described in one implementation as being performed on user device 120 can also be performed by media item server 110 or prediction server 130 in other implementations. For example, media item server 110 can stream media item 112 to user device 120 and can receive an indication of media item 112.
[0049] Additionally, the functionality belonging to a particular component can be performed by different or multiple components operating together. Media item server 110, server machine 170, server machine 180, or prediction server 130 can also be accessed as a service provided to other systems or devices through an appropriate application programming interface (API) and is thus not limited to use in websites and applications.
[0050] In implementations of the present disclosure, a "user" may be represented as a single individual. However, other implementations of the present disclosure include "users" that are entities controlled by a group of users and / or an automated source. For example, a group of individual users united as a community in a social network may be considered a "user". In another example, an automated consumer may be the automated ingestion pipeline of an application distribution platform.
[0051] Although implementations of the present disclosure are described in terms of a media item server 110, a prediction server 130, and a media item platform, the implementations may generally also apply to any type of platform that provides a connection between content and users.
[0052] In addition to the above description, controls may be provided to a user that allow the user to select whether and when the systems, programs, or features described herein can collect user information (e.g., information about the user's social network, social behavior or activities, occupation, user preferences, or user's current location), and whether to send content or communications to the user from a server (e.g., media item server 110 or prediction server 130). Additionally, certain data may be processed in one or more ways before storage or use to remove personally identifiable information. For example, a user's identity may be processed so that personally identifiable information cannot be determined for the user, or the user's geographical location may be generalized where location information is obtained (such as at the city, postal code, or state level) so that the user's specific location cannot be determined. Thus, the user can control what information about the user is collected, how the information is used, and what information is provided to the user.
[0053] Figure 2 is an exemplary tuning set generator that creates tuning data for a machine learning model using information according to an implementation of the present disclosure. System 200 shows a tuning set generator 171, tuning inputs 210, and target outputs 220 (e.g., target tuning outputs). As referred to Figure 1 as described, system 200 may include components similar to those of system 100. The components described for system 100 with reference to Figure 1 may be used to help describe Figure 2 system 200.
[0054] In an implementation, the tuning set generator 171 generates tuning data that includes one or more tuning inputs 210 and one or more target outputs 220. The tuning data may also include mapping data that maps the tuning inputs 210 to the target outputs 220. The tuning inputs 210 may also be referred to as "features", "attributes", or "information". In some implementations, the tuning set generator 171 may provide the tuning data in a tuning set and provide the tuning set to a tuning engine 181, where the tuning set is used to train a machine learning model 190. Further reference may be made to Figure 3CDescribe some implementations of generating a tuning set.
[0055] In one implementation, the tuning input 210 may include information 162 associated with the current media item 116 and one or more tagged media items 114, where the tagged media item 114 includes at least one corresponding segment similar to one of the segments of the current media item 116. For the current media item 116, the information 162 may include the length 212A of segment 214A of the current media item 116, the length 212B of segment 214B of the current media item 116, etc. (the length 212 of segment 214 below) and the length 216 of the current media item 116. For each tagged media item 114, the information 162 may include the length 218 of the tagged media item 114.
[0056] In an implementation, the target output 220 may include the generated label 164 of the current media item 116. In some implementations, the generated label 164 may have been generated through manual review of the current media item 116. In some implementations, the generated label 164 may have been generated through automatic review of the current media item 116.
[0057] In some implementations, after generating the tuning set and using the tuning set to train the machine learning model 190, the machine learning model 190 may be further trained (e.g., for additional data for the tuning set) or adjusted (e.g., adjusting the weights of the parameters 191) using the generated label 164 of the current media item 116.
[0058] Figure 3A-D depicts a flowchart of illustrative examples of methods 300, 320, 340, and 360 for review decisions of predicted media items according to implementations of the present disclosure. Methods 300, 320, 340, and 360 are exemplary methods from the perspective of a prediction system 105 (e.g., one or more of server machines 170, server machines 180, or prediction server 130) (e.g., and / or a media item platform or media item server 110). Methods 300, 320, 340, and 360 may be performed by a processing device that may include hardware (e.g., circuitry, dedicated logic), software (such as running on a general-purpose computer system or a dedicated machine), or a combination of both. Each of methods 300, 320, 340, and 360 and their individual functions, routines, subroutines, or operations may be performed by one or more processors of a computer device executing the method. In a particular implementation, each of methods 300, 320, 340, and 360 may be performed by a single processing thread. Alternatively, each of methods 300, 320, 340, and 360 may be performed by two or more processing threads, each thread performing one or more of the individual functions, routines, subroutines, or operations of the method.
[0059] For simplicity of explanation, the methods of the present disclosure are depicted and described as a series of acts. However, acts in accordance with the present disclosure may occur in various orders and / or concurrently, and in conjunction with other acts not presented and described herein. Moreover, not all acts shown may be required to implement the methods in accordance with the disclosed subject matter. Additionally, those skilled in the art will understand and appreciate that these methods may alternatively be represented as a series of related states via a state diagram or events. Additionally, it should be understood that the methods disclosed in this specification are capable of being stored on an article of manufacture to facilitate transporting and transferring such methods to a computing device. As used herein, the term "article of manufacture" is intended to include a computer program accessible from any computer-readable device or storage medium. For example, a non-transitory machine-readable storage medium may store instructions that, when executed, cause a processing device (e.g., prediction system 105, media item server 110, user device 120, prediction server 130, content owner device 140, server machines 170, server machines 180, media item platform, etc.) to perform operations including the methods disclosed therein. In another example, a system includes a memory for storing instructions and a processing device communicatively coupled to the memory, the processing device executing the instructions to perform the methods disclosed therein. In one implementation, methods 300, 320, 340, and 360 may be performed by Figure 1 the prediction system 105.
[0060] Refer to Figure 3A, Method 300 may be executed by one or more processing devices of prediction server 130 for predicting a review decision of a media item. Method 300 may be executed by an application or a background thread executing on one or more processing devices on prediction server 130. In some implementations, one or more portions of Method 300 may be executed by one or more of prediction system 105, media item server 110, prediction server 130 (e.g., prediction manager 132), or the media item platform.
[0061] At block 302, the processing device identifies the current media item 116 to be processed. For example, the current media item 116 may have been uploaded by content owner device 140 to be available for playback via the media item platform. In some implementations, the current media item 116 has been previously reviewed (e.g., associated with label 164), and the current media item 116 is identified as undergoing further processing (e.g., via Method 300) to update or confirm label 164. In some implementations, the media item 116 has been labeled (e.g., by user device 120 during playback, by a machine learning model trained with images and labels 164), and the current media item 116 is identified as undergoing further processing (e.g., via Method 300). In some implementations, the current media item 116 is the next media item in a queue of media items to be processed.
[0062] At block 304, the processing device processes the tagged media item 114 to identify at least one corresponding segment of the tagged media item that includes one of the segments similar to (e.g., matching, substantially the same as) the current media item 116. For example, the frames of the current media item 116 may be compared to the frames of the tagged media item 114 (e.g., the indices of the frames of the current media item 116 are compared to the frames of the tagged media item 114). The processing device may determine that a first segment of the tagged media item 114 and a second segment of the current media item 116 are similar (e.g., matching, substantially similar), although one of the segments has a border or frame, and / or is inverted, and / or is accelerated or decelerated, and / or has static, and / or has different quality, etc. The processing device may determine that at least a portion of the frames or audio segments of the current media item 116 match at least a portion of the frames or audio segments of the tagged media item 114. The processing device may identify segments of the current media item 116, each of which at least partially matches one or more segments of the tagged media item 114. The processing device may divide the current media item 116 (V) into segments (S) (e.g., V = S1 +... + S k ), where adjacent segments belong to different clusters or do not belong to a cluster.
[0063] At block 306, for each segment of the current media item 116, the processing device generates a segment prediction value 168 that indicates a specific attribute associated with the corresponding segment of the current media item 116, based on the attributes associated with the corresponding tagged media item 114. The processing device can combine positive predictions (e.g., the segment is similar to one or more tagged media items with a "positive review" tag) and negative predictions (e.g., the segment is also similar to one or more tagged media items with a "negative review" tag). Determining that a segment is similar to one or more tagged media items with a "positive review" tag (tagged as "good") and one or more media items with a "negative review" tag (tagged as "bad") can be referred to as determining the match graph for the segment. Bayesian inference can be used to predict the review decision for the current media item 116 in the match graph (e.g., the media item prediction value 169 based on the segment prediction value 168).
[0064] For a media item (V), the following values can be used (e.g., based on Figure 4A Table 400A):
[0065] - The prior probability of the hypothesis that video V should have a "negative review" tag ~(20 + 4) / (20 + 4 + 75 + 1) = 24%
[0066] The prior probability of the hypothesis that video V should have a "negative review" tag ~(20 + 4) / (20 + 4 + 75 + 1) = 24%
[0067] - The marginal similarity of the event that video V has a "negative review" tag ~21%
[0068] The marginal similarity of the event that video V has a "negative review" tag ~21%
[0069] - The probability that V has a "negative review" tag under the hypothesis that V should have a "negative review" tag ~20 / 24 ≈ 83%
[0070] - The probability that V should have a "negative review" tag given that V has a "negative review" tag ~20 / 21 ≈ 95%
[0071] These values (e.g., by Bayes' formula) can be related by the following equation:
[0072]
[0073] This equation can be verified by using the values in Table 400A:
[0074] 20 / 21 = 24 / 100 * 20 / 24 * 100 / 21
[0075] Method 300 can predict a media item prediction value 169 (e.g., whether the predicted media item should have a "positive review" label or a "negative review" label) based on a tag 164 (event E) that tags a media item 114 (e.g., a previous reviewer decision represented as event E). If the media item contains disjoint segments S1, ..., S k , the probability that media item V should have a "positive review" label is the product of the probabilities of the segments:
[0076]
[0077] A media item 112 is labeled as not containing a content type only if each segment of the media item 112 does not contain the content type (e.g., a media item 112 has a "positive review" label only if each part of the media item 112 has a "positive review" label). This equation also holds in the absence of evidence:
[0078]
[0079] where is a constant (e.g., parameter 191, the probability that any media item should have a "negative review" label in the absence of information). Under the assumption that the (prior) probability of a media item 112 containing a content type (e.g., a media item 112 with a "negative review" label) is independent of the length of the media item 112, the (prior) probability that a segment of the current media item 116 (e.g., the media item 112 under consideration) contains a content type (e.g., has a "negative review" label) decreases exponentially with the fractional length of the segment. The shorter the segment of the current media item 116 (e.g., the corresponding segment that matches a tagged media item 114 with a "negative review" label), the lower the probability that the current media item 116 should have a "negative review" label.
[0080] A tagged media item 116 that is labeled as containing a content type (e.g., has a "negative review" label) can be referred to as media item A, and a tagged media item 116 that is labeled as not containing a content type (e.g., has a "positive review" label) can be referred to as media item O.
[0081] A segment (S) of the current media item 116 (V) can be similar to (e.g., match) a tagged media item 114 (A) that has been labeled as containing a content type (e.g., both V and A contain S). The relationship between and can be estimated by the influence of on and is represented by the impact of. On The impact of can be represented by the following equation:
[0082]
[0083] On The impact of may be a multiplication factor of approximately 6.3% (using the values in Table 400A). If S covers all of A, this estimate of the impact is likely to be accurate. If S does not cover all of A, then the fraction of A covered by S can be considered as an exponent. Since S is divided into two segments S1 and S2, both of which match different parts of A, the resulting probabilities will be multiplied, and thus this factor can be used as an exponent, as shown in the following equation:
[0084]
[0085] where f is a constant factor In some implementations, f is an impact factor and is specific to the tagged media item (A). For example, an 85% match has an impact of 9.5%, a 50% match has an impact of 25%, a 10% match has an impact of 76%, etc.
[0086] In response to the tagged media item 114 being tagged as not containing a content type (e.g., having a "positive review" tag), all segments of the tagged media item 114 also do not contain the content type (e.g., also having a "positive review" tag). For a long media item 112, it is not possible for a reviewer to consider all of the media item 112 (e.g., in the absence of a signal indicating that the reviewer should consider all of the media item 112). On The impact of can be shown in the following equation:
[0087]
[0088] where
[0089] The constant g can be specific to the tagged media item (O). Based on segments similar to the tagged media item 114 that are tagged as containing a content type, the probability (x) that a segment (S) does not contain the content type (e.g., x is the probability that S should have a "positive review" tag based on a "negative review" tag) can be represented by the following equation:
[0090]
[0091] Based on a segment similar to the tagged media item 114 marked as not containing the content type, the probability (y) that the segment (S) contains the content type (e.g., y is the probability that S should have a "negative review" label given a "positive review" label) can be represented by the following equation:
[0092]
[0093] These probabilities can be combined based on Figure 4B Table 400B of. There are two independent pieces of evidence and four possibilities, two of which (as shown in the shaded cells of Table 400B) can be excluded.
[0094] The resulting probability that the segment (S) does not contain the content type (e.g., the resulting probability that S should have a "positive review" label given a specific attribute associated with label 164) can be shown by the following equation:
[0095]
[0096] If the predictions are the same (i.e., if x = 1 - y), the equation simplifies to x 2 / (x 2 +(1 - x) 2 )(e.g., a sigmoid function). If y = 1 / 2, the equation simplifies to If x = 1 / 2, the equation simplifies to 1 - y.
[0097] In the absence of a segment (S) similar to the tagged media item 114 marked as not containing the content type (e.g., in the absence of a "positive review" label), the prior probability of the hypothesis that the segment (S) contains the content type (e.g., should have a "negative review" label) can be one - half (e.g., ). If and the segment (S) is only similar to the tagged media item 114 marked as not containing the content type (e.g., only a "positive review" label exists), then and and if is close to 100%, then is very close to (e.g., if the error is less than 1%).
[0098] In some implementations, for each segment Si, the calculations of having a review and x and y are as follows:
[0099]
[0100]
[0101] Among them
[0102]
[0103]
[0104]
[0105] (The factor 1 / 2 in y comes from the fact that it is neutral and combines the "positive review" label and the "negative review" label)
[0106] And make
[0107]
[0108] If there is no review, this simplifies to k|S i | / |V|. In other implementations, the x and y probabilities based on the "positive review" label and the "negative review" label can be combined in other ways.
[0109] In block 308, the processing device calculates a media item prediction value 169 (e.g., combined probability) of the current media item 116 based on the generated segment prediction values 168 of each segment of the current media item 116. The segment prediction values 168 can be combined into the media item prediction value 169 of the entire current media item 116 through the following equation:
[0110]
[0111] In some implementations, the media item prediction value 169 can indicate the probability that the current media item 116 does not contain a type of content or attribute (e.g., the probability of having a "positive review" label). For example, the media item prediction value 169 can indicate that there is a 90% probability that the current media item 116 does not contain a type of content or attribute. In some implementations, the segment prediction value 168 and the media item prediction value 169 can be scores indicating the relative probabilities of containing a type of content or attribute. For example, the first media item prediction value of the first current media item can indicate a larger final score, and the second media item prediction value of the second current media item can indicate a smaller final score. A relatively larger final score indicates that the first current media item is more likely to contain the content type than the second current media item. In some implementations, the segment prediction values 168 are multiplied together to calculate the media item prediction value 169. In some implementations, the segment prediction values 168 are combined using one or more other operations (e.g., instead of multiplication, combining multiplication, etc.). In some implementations, the segment prediction values 168 are combined with a previous prediction value of the current media item 116 (e.g., a probability or score generated by a previous review of the current media item 116) via one or more operations (e.g., multiplication, etc.).
[0112] In the example, at block 304, the processing device can determine that a first segment (e.g., 0 - 15 seconds) of the current media item 116 matches a first corresponding segment of the first labeled media item 114A, a second segment (e.g., 0 - 30 seconds) of the current media item 116 matches a second corresponding segment of the second labeled media item 114B, and the first segment is a sub-segment of the second segment. The second segment of the current media item 116 can include the first segment and a third segment of the current media item 116. At block 306, the processing device can process the first segment (e.g., generate a first segment prediction value indicating a first label of the first segment) based on the first corresponding label 164A of the first labeled media 114A item and the second corresponding label 164B of the second labeled media item 114B. At block 306, the processing device can process the second segment (e.g., generate a second segment prediction value indicating a second label of the first segment) based on the second corresponding label 164B of the second labeled media item 114B. At block 308, the processing device can calculate the media item prediction value 169 based on the generated first segment prediction value 168 and the generated second segment prediction value 168.
[0113] At block 310, the processing device causes the current media item 116 to be processed based on the media item prediction value 169. For example, the processing device may provide the media item prediction value 169 (or information about the media item prediction value 169) to the media item platform to initiate the processing of the current media item 116. Alternatively, the processing device itself may perform the processing of the current media item 116 based on the media item prediction value 169. In some implementations, the processing of the current media item 116 includes applying a policy (e.g., blocking playback, applying playback, sending for review, etc.) to the current media item 116 based on the media item prediction value 169. In some implementations, Figure 3B illustrates the processing of the current media item 116 based on the media item prediction value 169. In some implementations, the processing device may cause one or more media items (e.g., advertisements, interstitial media items) to be associated with the playback of the current media item 116 (e.g., causing the playback of the additional media item to be combined with the playback of the current media item 116). In some implementations, the processing device may cause an indication to be transmitted to the content owner device 140 associated with the current media item 116 (e.g., an indication about a problem with the current media item 116) based on the media item prediction value 169. In some implementations, the processing device may modify or cause the current media item 116 to be modified based on the media item prediction value 169. In some implementations, the processing device associates the tag 164 with the current media item 116 based on the media item prediction value 169.
[0114] In some implementations, the processing device may send the current media item 116 for review (e.g., manual review) and may receive one or more review decisions (e.g., manual review decisions). The processing device may tune k, f i and g i (e.g., parameter 191) (e.g., via retraining, via a feedback loop). In some implementations, this tuning optimizes the area under the curve (AUC). In some implementations, this tuning optimizes the precision at a specific recall point.
[0115] Referring Figure 3B , method 320 may be performed by one or more processing devices of the prediction system 105 and / or the media item platform for predicting a review decision. Method 320 may be used to process the current media item 116 based on the media item prediction value 169. In some implementations, Figure 3ABlock 310 includes method 320. Method 320 can be executed by an application or a background thread executed on one or more processing devices of the prediction system 105 and / or the media item platform. As described herein, the threshold condition can be one or more of a threshold media item prediction value, probability, score, confidence level, etc. For example, the first threshold condition can be a media item prediction value of 99% or greater. In some implementations, the threshold condition can be a combination of one or more of a threshold media item prediction value, probability, score, confidence level, etc. For example, the first threshold condition can satisfy the first media item prediction value and the threshold confidence level.
[0116] In block 322, the processing device determines whether the calculated media item prediction value 169 satisfies the first threshold condition. In response to the calculated media item prediction value 169 satisfying the first threshold condition, the process proceeds to block 324. In response to the calculated media item prediction value 169 not satisfying the first threshold condition, the process proceeds to block 326.
[0117] In block 324, the processing device blocks the playback of the current media item 116 via the media item platform. For example, if the media item prediction value 169 satisfies the first probability (e.g., a probability at or exceeding 99%) that the current media item 116 contains a content type (e.g., inappropriate, has technical problems, violates the rights of the content owner, etc.), then the current media item 116 may be blocked by the media item platform. In some implementations, an indication of the content type (e.g., inappropriate, technical problem, infringement, etc.) can be transmitted to the content owner device 140 that uploaded the current media item 116. The content owner device can modify (e.g., or replace) the current media item 116 and upload the modified (or new) media item 112.
[0118] In block 326, the processing device determines whether the calculated media item prediction value 168 satisfies the second threshold condition. In response to the calculated media item prediction value 168 satisfying the second threshold condition, the process proceeds to block 328. In response to the calculated media item prediction value 168 not satisfying the second threshold condition, the process proceeds to block 330.
[0119] At block 328, the processing device allows the media item to be played via the media item platform. For example, if the media item prediction value 169 satisfies a second probability (e.g., a probability less than or equal to 50%) that the current media item 116 contains a content type (e.g., inappropriate, has technical issues, violates the rights of the content owner, etc.), then the playback of the current media item 116 may be allowed (e.g., the current media item 116 may be accessed via the media item platform for playback). In some implementations, an indication to allow the playback of the current media item 116 via the media item platform may be transmitted to the content owner device 140 that uploaded the current media item 116.
[0120] At block 330, the processing device determines whether the calculated media item prediction value 169 satisfies a third threshold condition. In response to the calculated media item prediction value 169 satisfying the third threshold condition, the process continues to block 332. In response to the calculated media item prediction value 169 not satisfying the third threshold condition, the process ends.
[0121] At block 332, the processing device causes the current media item 116 to be reviewed to generate a label 164 indicating whether playback via the media item platform is allowed. In some implementations, the third threshold condition is between the first and second threshold conditions (e.g., a probability lower than the first threshold condition and higher than the second threshold condition, e.g., a probability of 50 - 99%). The processing device may cause the current media item 116 to be manually reviewed.
[0122] At block 334, the processing device receives the generated label 164 for the current media item 116. The processing device may receive the generated label 164 that is generated by a user (e.g., an administrator of the media item platform) who manually reviews the current media item 116 in response to the media item prediction value 169 satisfying the third threshold condition.
[0123] In some implementations, at block 336, the processing device adjusts the weights (e.g., the weights of the parameters 191 of the model 190) for generating the segment prediction value 168 (e.g., retunes the trained machine learning model 190) based on the generated label 164.
[0124] At block 338, the processing device determines whether the generated label 164 indicates allowed playback. In response to the generated label 164 indicating allowed playback, the process continues to block 328. In response to the generated label 164 indicating not allowed playback, the process continues to block 324.
[0125] Reference Figure 3C, Method 340 may be executed by one or more processing devices of prediction system 105 to predict a review decision. According to an implementation of the present disclosure, prediction system 105 may use Method 340 to train a machine learning model. In one implementation, some or all operations of Method 340 may be performed by Figure 1 one or more components of system 100. In other implementations, one or more operations of Method 340 may be performed by referring to Figure 1-2 the tuning set generator 171 of server machine 170 described.
[0126] Method 340 generates tuning data for a machine learning model. In some implementations, at block 342, the processing logic implementing Method 300 initializes the data set (e.g., tuning set) T to an empty set.
[0127] At block 344, the processing logic generates tuning inputs that for each segment 214 of the current media item 116 include the length 212 of the corresponding segment 214, the length 216 of the current media item 116, and the length 218 of the corresponding tagged media item 114 (e.g., having segments similar to the segments of the current media item 116).
[0128] At block 346, the processing logic generates target outputs for one or more of the tuning inputs. The target outputs may include the label 164 of the current media item 116. In some implementations, the label 164 may be received at Figure 3B block 334.
[0129] At block 348, the processing logic optionally generates mapping data indicating an input / output mapping (e.g., information 162 associated with the current media item 116 mapped to the label 164 of the current media item 116). The input / output mapping (or mapping data) may refer to the tuning inputs (e.g., one or more tuning inputs described herein), the target outputs of the tuning inputs (e.g., where the target output identifies an indication of the user's preference to cancel the corresponding transmission), and the association between the tuning inputs and the target outputs.
[0130] At block 350, the processing logic adds the mapping data to the data set T initialized at block 342.
[0131] At block 352, the processing logic branches based on whether the tuning set T is sufficient to train the machine learning model 190. If so, the execution proceeds to block 354; otherwise, the execution continues back to block 344. It should be noted that in some implementations, the sufficiency of the tuning set T can be determined solely based on the number of input / output mappings in the tuning set, while in some other implementations, in addition to or instead of the number of input / output mappings, the sufficiency of the tuning set T can be determined based on one or more other criteria (e.g., a measure of the diversity of the tuning examples, accuracy, etc.).
[0132] At block 354, the processing logic provides the tuning set T to train the machine learning model 190. In one implementation, the tuning set T is provided to the tuning engine 181 of the server machine 180 to perform the training or retraining of the model 190. In some implementations, the training or retraining of the model includes adjusting the weights of the parameters 191 of the model 190 (see block 336 of Figure 3B ). After block 354, the machine learning model 190 can be trained or retrained based on the tuning set T, and the trained machine learning model 190 can be implemented (e.g., by the prediction manager 132) to predict the review decision of the current media item 116.
[0133] Refer to Figure 3D , the method 360 can be executed by one or more processing devices of the prediction system 105 for predicting the review decision. The method 360 can be used to predict the review decision. The method 360 can be executed by an application or a background thread running on one or more processing devices (e.g., the prediction manager 132) of the prediction system 105.
[0134] At block 362, for each segment 214 of the current media item 116, the processing device provides the length 212 of the corresponding segment 214, the length 216 of the current media item 116, and the length 218 of the corresponding tagged media item 114 to the trained machine learning model 190. The trained machine learning model 190 can be trained by the method 340 of Figure 3C . The trained machine learning model 190 can execute one or more of blocks 306 - 308 of the Figure 3A method. For example, the trained machine learning model can determine the segment prediction value 168 based on the tuning inputs and parameters 191 provided in block 362. In some implementations, the trained machine learning model can process the segment prediction values 168 (e.g., multiply the segment prediction values 168, combine the segment prediction values 168) to generate the media item prediction value 169.
[0135] At block 344, the processing device may obtain one or more outputs from the trained machine learning model 190. At block 346, the processing device may determine a media item prediction value 169 for the current media item 116 based on the one or more outputs. In some implementations, the processing device may extract from the one or more outputs the confidence level that the media item prediction value 169 will correspond to (e.g., match) the generated label 164 (e.g., the label 164 received in response to a manual review of the current media item 116).
[0136] Figure 4A -B depicts a table 400 associated with a review decision of a predicted media item 112 according to an implementation of the present disclosure. As used herein, the terms "positive review" (e.g., rated good, actually good, good review, etc.) and "negative review" (rated bad, actually bad, poor review) may indicate an attribute of the label 164 or the media item 112. In some implementations, a "negative review" may be a label indicating that the media item 112 contains a certain type of content or attribute (e.g., inappropriate, having technical problems, infringing on the rights of others, etc.), and a "positive review" may be a label indicating the absence of that type of content or attribute. In some implementations, the media item 112 may be associated with multiple labels simultaneously. For example, a first label of the media item 112 may indicate an age-appropriateness rating (e.g., teenagers and above), a second label of the media item 112 may indicate that the media item 112 is suitable for a specific advertisement, a third label of the media item 112 may indicate a specific technical problem (e.g., subtitle problem), etc. The label may indicate that a specific type of content is inappropriate (e.g., including one or more sexual content, violence or offensive content, hate or abusive content, harmful or dangerous behavior, child abuse, terrorism promotion, spam or misleading, etc.), infringing on rights, having technical problems (e.g., subtitle problems, etc.), having a rating (e.g., age-appropriateness rating, etc.), suitable for advertisement, etc.
[0137] Figure 4A Depicts a table 400A associated with the label 164 of a media item 112 (e.g., labeled media item 114) according to an implementation of the present disclosure. The media item 112 may have been manually reviewed to be assigned the label 164. As depicted in table 400A, the manual review results in a 79% probability that the label 164 indicates a positive review of the media item 112 and a 21% probability that the label 164 indicates a negative review of the media item 112. A positive review may indicate that the media item does not contain inappropriate content or does not have technical problems. A negative review may indicate that the media item contains inappropriate content or has technical problems. Although the terms "positive review" and "negative review" and the term "content type" are used herein, it should be understood that the present disclosure is applicable to any type of label of the media item 112.
[0138] Manual reviews can be partially accurate. For example, during a manual review, one or more portions of media item 112 can be reviewed, and one or more other portions of media item 112 can be not reviewed (e.g., sampling portions of media item 112, skimming media item 112, etc.). Different users can provide different labels for the same media item 112. For example, for a label of a provocative dance, a first user can consider the dance in media item 112 not provocative enough to merit a negative review label, and a second user can consider the dance in media item 112 provocative enough to merit a negative review label.
[0139] Actual accuracy can be determined by a manual review by an administrator (e.g., the supervisor of the user performing the initial manual review). As depicted in Table 400A, the actual values result in a 76% probability that label 164 indicates a positive review of media item 112 (e.g., actually positive, the administrator will assign a positive review label), and a 24% probability of a negative review of media item 112 (e.g., actually negative, the administrator will assign a negative review label). The values from Table 400A can be used to calculate the media item predicted value 169.
[0140] The actual percentages in Table 400A can be replaced with actual data when available. The values in Table 400A can be updated periodically as the distribution changes. The values in Table 400A may have a recency bias (e.g., a more recent review of media item 112 may be weighted more than a less recent review of media item 112).
[0141] Figure 4B Depict Table 400B showing a review decision predicting media item 112 according to an implementation of the present disclosure.
[0142] As discussed herein, equations can be used to calculate the segment predicted value 168. A first segment 214A of the current media item 116 can be a corresponding first segment similar to a first labeled media item 114A having a "positive review" label, and a second segment 214B of the current media item 116 can be a corresponding second segment similar to a second labeled media item 114B having a "negative review" label. Based on a labeled media item 114A similar to having a "positive review" label (e.g., good from a good review), the probability that a segment of the current media item 116 has a "positive review" label can be represented by a first equation:
[0143] x*(1 - y)
[0144] Based on a tagged media item 114A similar to having a "positive review" tag (e.g., good from a good review), the probability that a segment of the current media item 116 has a "negative review" tag can be represented by a second equation:
[0145] (1 - x)*y
[0146] The segment prediction value 168 can be based on Figure 3A the first and second equations described in block 306 of
[0147] Figure 5 FIG. [FIG. number not provided in the original] is a block diagram of one implementation of a computer system according to an implementation of the present disclosure. In a particular implementation, the computer system 500 may be connected (e.g., via a network such as a local area network (LAN), intranet, extranet, or the Internet) to other computer systems. The computer system 500 may operate as a server or a client computer in a client-server environment, or as a peer computer in a peer-to-peer or distributed network environment. The computer system 500 may be provided by a personal computer (PC), tablet PC, set-top box (STB), personal digital assistant (PDA), cellular phone, web device, server, network router, switch, or bridge, or any device capable of executing a set of instructions (sequential or otherwise) that specify actions to be taken by that device. Additionally, the term "computer" shall include any collection of computers that individually or jointly execute a set (or multiple sets) of instructions to perform any one or more of the methods described herein.
[0148] In another aspect, the computer system 500 may include a processing device 502, volatile memory 504 (e.g., random access memory (RAM)), non-volatile memory 506 (e.g., read-only memory (ROM) or electrically erasable programmable ROM (EEPROM)), and a data storage device 516, which may communicate with each other via a bus 508.
[0149] The processing device 502 may be provided by one or more processors, such as a general-purpose processor (e.g., a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a microprocessor implementing other types of instruction sets, or a microprocessor implementing a combination type of instruction set) or a dedicated processor (e.g., an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), or a network processor).
[0150] The computer system 500 may also include a network interface device 522. The computer system 500 may also include a video display unit 510 (e.g., an LCD), an alphanumeric input device 512 (e.g., a keyboard), a cursor control device 514 (e.g., a mouse), and a signal generation device 520.
[0151] In some implementations, the data storage device 516 may include a non-transitory computer-readable storage medium 524, on which instructions 526 encoding any one or more of the methods or functions described herein may be stored, including instructions Figure 1 encoding the prediction manager 132 and used to implement one or more of the methods 300, 320, 340, or 360.
[0152] The instructions 526 may also reside, completely or partially, within the volatile memory 504 and / or the processing device 502 during execution by the computer system 500. Thus, the volatile memory 504 and the processing device 502 may also constitute machine-readable storage media.
[0153] Although the computer-readable storage medium 524 is shown as a single medium in the illustrative example, the term "computer-readable storage medium" shall include a single medium or multiple media (e.g., a centralized or distributed database, and / or an associated cache and server) storing a set or multiple sets of executable instructions. The term "computer-readable storage medium" shall also include any tangible medium that can store or encode a set of instructions for execution by a computer such that the computer performs any one or more of the methods described herein. The term "computer-readable storage medium" shall include, but not be limited to, solid-state memory, optical media, and magnetic media.
[0154] The methods, components, and features described herein may be implemented by discrete hardware components or may be integrated into the functionality of other hardware components such as ASICs, FPGAs, DSPs, or similar devices. Additionally, the methods, components, and features may be implemented by firmware modules or functional circuits within a hardware device. Further, the methods, components, and features may be implemented in any combination of a hardware device and computer program components or in a computer program.
[0155] Unless otherwise specified, terms such as "identifying", "processing", "generating", "calculating", "handling", "determining", "preventing", "allowing", "causing", "adjusting", "training", "tuning", etc. refer to actions and processes performed or implemented by a computer system that manipulates and transforms data represented as physical (electronic) quantities within the computer system registers and memory into other data similarly represented as physical quantities within the computer system memory or registers or other such information storage, transmission, or display devices. Additionally, as used herein, the terms "first", "second", "third", "fourth", etc. are labels used to distinguish different elements and may not have ordinal significance in accordance with their numerical names.
[0156] The examples described herein also relate to apparatus for performing the methods described herein. The apparatus may be specially constructed to perform the methods described herein or it may comprise a general purpose computer system selectively programmed by a computer program stored in the computer system. Such a computer program may be stored in a computer-readable tangible storage medium.
[0157] The methods and illustrative examples described herein are not inherently related to any particular computer or other apparatus. Various general purpose systems may be used in accordance with the teachings described herein, or it may prove convenient to construct a more specialized apparatus to perform methods 300, 320, 340, and 360 and / or their respective functions, routines, subroutines, or operations. Examples of the structure of various such systems are set forth in the foregoing description.
[0158] The foregoing description is intended to be illustrative and not restrictive. Although the present disclosure has been described with reference to specific illustrative examples and implementations, it will be recognized that the present disclosure is not limited to the examples and implementations described. The scope of the present disclosure should be determined with reference to the following claims and the full scope of equivalents authorized by the claims.
Claims
1. A method for predicting a review decision, comprising: Identifying a current media item to be processed, the current media item including a current digital video or a current digital audio; Processing a plurality of tagged media items to identify tagged media items, the tagged media items including at least one corresponding segment similar to one segment of a plurality of segments of the current media item, the tagged media items including tagged digital video or tagged digital audio; For each segment of the plurality of segments of the current media item, generating a segment prediction value indicative of a specific attribute associated with the corresponding segment of the current media item, based on an attribute associated with a corresponding tagged media item, each tagged media item including a corresponding segment similar to the corresponding segment of the current media item; Calculating a media item prediction value for the current media item based on the generated segment prediction values for each segment of the plurality of segments of the current media item; And Causing the current media item to be processed based on the calculated media item prediction value.
2. The method according to claim 1, wherein: Each of the tagged media items is assigned a corresponding label; and Each segment of the plurality of segments of the current media item at least partially matches one or more segments of the tagged media item.
3. The method according to claim 1, wherein, Generating the segment prediction value for the corresponding segment for each segment of the plurality of segments of the current media item is based on a plurality of parameters, the plurality of parameters including at least one of the following: the length of the corresponding segment, the length of the current media item, or the length of the corresponding tagged media item, each tagged media item including the corresponding segment similar to the corresponding segment of the current media item.
4. The method according to claim 1, further comprising: Determining that a first segment of the current media item matches a first corresponding segment of a first tagged media item; Determining that a second segment of the current media item matches a second corresponding segment of a second tagged media item; Determining that the first segment is a sub-segment of the second segment, wherein the second segment includes the first segment and a third segment, and wherein generating the segment prediction value for each segment of the plurality of segments includes: Generating a first segment prediction value indicative of a first label of the first segment based on a first corresponding label of the first tagged media item and a second corresponding label of the second tagged media item; and Generating a second segment prediction value indicative of a second label of the third segment based on the second corresponding label of the second tagged media item, wherein calculating the media item prediction value is based on the generated first segment prediction value and the generated second segment prediction value.
5. The method according to claim 1, wherein, Causing the current media item to be processed includes one of the following: In response to the calculated media item prediction value satisfying a first threshold condition, causing the current media item to be blocked from being played via a media item platform; In response to the calculated media item prediction value satisfying a second threshold condition, causing the current media item to be allowed to be played via the media item platform; Or In response to the calculated media item prediction value satisfying a third threshold condition, causing the current media item to be reviewed to generate a label indicating whether the playback of the current media item via the media item platform is allowed.
6. The method according to claim 5, wherein, The segment prediction value for each segment of the plurality of segments is generated based on a plurality of parameters and one or more weights associated with one or more of the plurality of parameters.
7. The method according to claim 6, further comprising adjusting the one or more weights based on the generated label of the current media item.
8. The method according to claim 7, wherein: Adjusting the one or more weights includes training a machine learning model based on a tuning input and a target tuning output of the tuning input to provide the adjusted one or more weights; For each segment of the plurality of segments of the current media item, the tuning input includes the length of the corresponding segment, the length of the current media item, and the length of the corresponding marked media item, each corresponding marked media item including the corresponding segment similar to the corresponding segment of the current media item; and The tuning target output of the tuning input includes the generated label of the current media item.
9. A non-transitory machine-readable storage medium storing instructions that, when executed, cause a processing device to perform operations including the following: Identifying a current media item to be processed, the current media item including a current digital video or a current digital audio; Processing a plurality of marked media items to identify marked media items, the marked media items including at least one corresponding segment similar to a segment of the plurality of segments of the current media item, the marked media items including marked digital videos or marked digital audios; For each segment of the plurality of segments of the current media item, generating a segment prediction value indicating a specific attribute associated with the corresponding segment of the current media item based on an attribute associated with the corresponding marked media item, each marked media item including the corresponding segment similar to the corresponding segment of the current media item; Calculating a media item prediction value of the current media item based on the generated segment prediction values of each segment of the plurality of segments of the current media item; And Causing the current media item to be processed based on the calculated media item prediction value.
10. The non-transitory machine-readable storage medium according to claim 9, wherein: Each of the marked media items is assigned a corresponding label; Each segment of the plurality of segments of the current media item at least partially matches one or more segments of the marked media item; And Generating the segment prediction value for each segment of the plurality of segments of the current media item is based on a plurality of parameters, the plurality of parameters including at least one of the following: the length of the corresponding segment, the length of the current media item, or the length of the corresponding labeled media item, each labeled media item including the corresponding segment similar to the corresponding segment of the current media item.
11. The non-transitory machine-readable storage medium according to claim 9, wherein, The operation further includes: Determining that a first segment of the current media item matches a first corresponding segment of a first labeled media item; Determining that a second segment of the current media item matches a second corresponding segment of a second labeled media item; Determining that the first segment is a sub-segment of the second segment, where the second segment includes the first segment and a third segment, and where generating the segment prediction value for each segment of the plurality of segments includes: Generating a first segment prediction value indicating a first label of the first segment based on a first corresponding label of the first labeled media item and a second corresponding label of the second labeled media item; and Generating a second segment prediction value indicating a second label of the third segment based on the second corresponding label of the second labeled media item, where calculating the media item prediction value is based on the generated first segment prediction value and the generated second segment prediction value.
12. The non-transitory machine-readable storage medium according to claim 9, wherein, Causing the current media item to be processed includes one of the following: In response to the calculated media item prediction value satisfying a first threshold condition, causing the playback of the current media item via the media item platform to be blocked; In response to the calculated media item prediction value satisfying a second threshold condition, causing the playback of the current media item via the media item platform to be allowed; Or In response to the calculated media item prediction value satisfying a third threshold condition, causing the current media item to be reviewed to generate a label indicating whether the playback of the current media item via the media item platform is allowed.
13. The non-transitory machine-readable storage medium according to claim 12, wherein: The segment prediction value for each segment of the plurality of segments is generated based on a plurality of parameters and one or more weights associated with one or more of the plurality of parameters; and The operation further includes adjusting the one or more weights based on the generated label of the current media item.
14. The non-transitory machine-readable storage medium according to claim 13, wherein: Adjusting the one or more weights includes training a machine learning model based on a tuning input and a target tuning output of the tuning input to provide the adjusted one or more weights; For each segment of the plurality of segments of the current media item, the tuning input includes the length of the corresponding segment, the length of the current media item, and the length of the corresponding labeled media item, each corresponding labeled media item including the corresponding segment similar to the corresponding segment of the current media item; and The target tuning output of the tuning input includes the generated label of the current media item.
15. A system for predicting review decisions, comprising: a memory configured to store instructions; and a processing device communicatively coupled to the memory, the processing device being configured to execute the instructions to: identify a current media item to be processed, the current media item including a current digital video or a current digital audio; process a plurality of tagged media items to identify tagged media items, the tagged media items including at least one corresponding segment similar to one segment of the plurality of segments of the current media item, the tagged media items including tagged digital video or tagged digital audio; for each segment of the plurality of segments of the current media item, generate a segment prediction value indicative of a specific attribute associated with the corresponding segment of the current media item based on an attribute associated with a corresponding tagged media item, each tagged media item including a corresponding segment similar to the corresponding segment of the current media item; calculate a media item prediction value of the current media item based on the generated segment prediction values for each segment of the plurality of segments of the current media item; and cause the current media item to be processed based on the calculated media item prediction value.
16. The system according to claim 15, wherein: each of the tagged media items is assigned a corresponding tag; each segment of the plurality of segments of the current media item at least partially matches one or more segments of the tagged media item; and the processing device generates the segment prediction value for each segment of the plurality of segments of the current media item based on a plurality of parameters, the plurality of parameters including at least one of the following: the length of the corresponding segment, the length of the current media item, or the length of the corresponding tagged media item, each tagged media item including the corresponding segment similar to the corresponding segment of the current media item.
17. The system according to claim 15, wherein, The processing device further: determines that a first segment of the current media item matches a first corresponding segment of a first tagged media item; determines that a second segment of the current media item matches a second corresponding segment of a second tagged media item; determines that the first segment is a sub-segment of the second segment, wherein the second segment includes the first segment and a third segment, and wherein, to generate the segment prediction value for each segment of the plurality of segments: generate a first segment prediction value indicative of a first tag of the first segment based on a first corresponding tag of the first tagged media item and a second corresponding tag of the second tagged media item; and generate a second segment prediction value indicative of a second tag of the third segment based on the second corresponding tag of the second tagged media item, wherein the processing device calculates the media item prediction value based on the generated first segment prediction value and the generated second segment prediction value.
18. The system according to claim 15, wherein To cause the current media item to be processed, the processing device will perform one of the following: In response to the calculated media item prediction value satisfying a first threshold condition, blocking the playback of the current media item via a media item platform; In response to the calculated media item prediction value satisfying a second threshold condition, allowing the playback of the current media item via the media item platform; Or In response to the calculated media item prediction value satisfying a third threshold condition, reviewing the current media item to generate a label indicating whether the playback of the current media item via the media item platform is allowed.
19. The system according to claim 18, wherein: The segment prediction value of each segment of the plurality of segments is generated based on a plurality of parameters and one or more weights associated with one or more of the plurality of parameters; And The processing device further adjusts the one or more weights based on the generated label of the current media item.
20. The system according to claim 19, wherein: To adjust the one or more weights, the processing device trains a machine learning model based on a tuning input and a target tuning output of the tuning input to provide the adjusted one or more weights; For each segment of the plurality of segments of the current media item, the tuning input includes the length of the corresponding segment, the length of the current media item, and the length of the corresponding labeled media item, each corresponding labeled media item including the corresponding segment similar to the corresponding segment of the current media item; and The tuning target output of the tuning input includes the generated label of the current media item.