Multi-stage adaptive system for content moderation
The multi-stage toxicity moderation system efficiently filters and adapts to toxic content in large-scale platforms by using adaptive triage stages, reducing computational load on human moderators and enhancing accuracy over time.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- MODULATE INC
- Filing Date
- 2021-10-08
- Publication Date
- 2026-05-25
AI Technical Summary
Current content moderation systems in large-scale multi-user platforms are either too simple to effectively prevent abusive behavior or too expensive to scale, and they struggle to adapt to changing environments or new platforms, leading to inefficiencies and high latency in addressing toxicity and nuisance behavior.
A multi-stage toxicity moderation system that employs a series of adaptive triage stages, each filtering out non-nuisance content and passing potentially harmful content to subsequent stages with more complex analysis, allowing for efficient and accurate identification of toxic content using a combination of automated and human moderation.
The system reduces the computational burden on human moderators by filtering out non-harmful content, improving accuracy over time through feedback, and enabling real-time moderation of large volumes of content while minimizing false positives and negatives.
Smart Images

Figure 0007864722000001 
Figure 0007864722000002 
Figure 0007864722000003
Abstract
Description
Technical Field
[0001] Priority This patent application claims the priority of U.S. Provisional Patent Application No. 63 / 089226, titled "MULTI-STAGE ADAPTIVE SYSTEM FOR CONTENT MODERATION," filed on October 8, 2020, with William Carter Huffman, Michael Pappas, and Henry Howie as inventors, and the entire disclosure thereof is incorporated herein by reference.
[0002] Field of the Invention Exemplary embodiments of the present invention generally relate to content moderation, and more particularly, various embodiments of the present invention relate to the moderation of audio content in an online environment.
[0003] Background of the Invention In large-scale multi-user platforms that enable communication between users, such as Reddit, Facebook, and video games, some users engage in harassment, discomfort, or humiliation of other users, causing problems of toxicity and nuisance behavior that discourage participation in the platform. Usually, nuisance behavior is carried out through text, speech, or video media. For example, verbally harassing another user in a voice chat or posting an aggressive video or article. Nuisance behavior may also be carried out by intentionally interfering with team-based activities. For example, a player in a team game intentionally performs poorly to shake up teammates. Such behavior affects both users and the platform itself. Users who encounter nuisance behavior may be less likely to participate in the platform or only participate for a short time, and sufficiently harmful behavior may cause users to completely abandon the platform.
[0004] Platforms can directly combat abusive behavior through content moderation, which involves observing their users and addressing any inappropriate content that is discovered. These responses can be direct, such as temporarily or permanently banning users who harass others, or more subtly, such as grouping toxic users together to prevent them from disrupting the rest of the platform. Traditional content moderation systems fall into two categories: those that are highly automated but easily circumvented and only exist in specific domains, and those that are precise but highly manual, time-consuming, and costly.
[0005] Overview of various embodiments According to one embodiment of the present invention, a toxicity moderation system has an input unit configured to receive speech from an utterancer. The system includes a multi-stage toxicity machine learning system having a first stage and a second stage. The first stage is trained to analyze the received utterance and determine whether the toxicity level of the utterance meets a toxicity threshold. The first stage is also configured to filter utterances that meet the toxicity threshold through to the second stage, and is further configured to filter out utterances that do not meet the toxicity threshold.
[0006] In various embodiments, the first stage is trained using a database having training data that includes positive and / or negative examples of the training content for the first stage. The first stage can be trained using a feedback process. The feedback process can receive speech content, analyze the speech content using the first stage, and classify the speech content as having positive and / or negative speech content for the first stage. The feedback process can also analyze the positive speech content for the first stage using the second stage, and classify the positive speech content for the first stage as having positive and / or negative speech content for the second stage. The feedback process can also update the database using positive and / or negative speech content for the second stage.
[0007] To support the overall efficiency of the system, the first stage may discard at least a portion of the negative speech content from the first stage. Furthermore, the first stage may be trained using a feedback process that includes analyzing some of the negative speech content from the first stage using the second stage to classify the negative speech content from the first stage as having positive speech content from the second stage and / or negative speech content from the second stage. The feedback process can update the database with positive speech content from the second stage and / or negative speech content from the second stage.
[0008] In particular, a toxicity moderation system may include a random uploader configured to upload portions of utterances that did not meet the toxicity threshold to a subsequent stage or a human moderator. The system may also include a session context flagger configured to receive instructions that an utterancer has previously met the toxicity threshold within a predetermined time. When such instructions are received, the flagger may (a) adjust the toxicity threshold, or (b) upload portions of utterances that did not meet the toxicity threshold to a subsequent stage or a human moderator.
[0009] The toxicity moderation system may also include a user context analyzer. The user context analyzer is configured to adjust the toxicity threshold and / or toxicity confidence level based on the speaker's age, the receiver's age, the speaker's geographical location, the speaker's friend list, the receiver's recent interaction history, the speaker's gameplay time, the length of the speaker's playtime, the start and end times of the game, and / or gameplay history. The system may also include a sentiment analyzer trained to determine the speaker's emotions. The system may also include an age analyzer trained to determine the speaker's age.
[0010] In various embodiments, the system has a temporal receptive field configured to segment utterances into time segments receivable by at least one stage. The system also has an utterance segmentator configured to segment utterances into time segments analyzable by at least one stage. In various embodiments, the first stage is more efficient than the second stage.
[0011] According to another embodiment, a multi-stage content analysis system includes a first stage trained with a database having training data including positive and / or negative examples of training content from the first stage. The first stage is configured to receive and analyze utterance content to classify the utterance content as having positive utterance content and / or negative utterance content from the first stage. The system includes a second stage configured to receive at least some (but not all) of the negative utterance content from the first stage. The second stage is further configured to analyze the positive utterance content from the first stage to classify the positive utterance content from the first stage as having positive utterance content and / or negative utterance content from the second stage. The second stage is further configured to update the database with the positive utterance content and / or negative utterance content from the second stage.
[0012] In particular, the second stage is configured to analyze the received negative speech content from the first stage and classify the negative speech content from the first stage as having positive speech content and / or negative speech content from the second stage. Furthermore, the second stage is configured to update the database using the positive speech content and / or negative speech content from the second stage.
[0013] In yet another embodiment, the method trains a multi-stage content analysis system. The method provides a multi-stage content analysis system having a first stage and a second stage. The system trains the first stage using a database having training data including positive and / or negative examples of training content for the first stage. The method receives speech content. The speech content is analyzed using the first stage and classified as having positive speech content and / or negative speech content for the first stage. The positive speech content for the first stage is analyzed using the second stage and classified as having positive speech content and / or negative speech content for the second stage. The method updates the database with positive speech content and / or negative speech content for the second stage. The method also discards at least some of the negative speech content for the first stage.
[0014] This method can further analyze some of the negative speech content from Stage 1 using Stage 2, and classify the negative speech content from Stage 1 as having positive speech content and / or negative speech content from Stage 2. This method can further update the database using the positive speech content and / or negative speech content from Stage 2.
[0015] In particular, this method can use a database containing training data that includes positive and / or negative examples of first-stage training content. This method generates first-stage positive judgments ("S1-positive judgments") and / or first-stage negative judgments ("S1-negative judgments") related to a portion of the utterance content. The utterances related to the S1-positive judgments are analyzed. In particular, the positive and / or negative examples are related to specific categories of hazards.
[0016] According to another embodiment, a moderation system for managing content includes a series of consecutive stages arranged in series. Each stage is configured to receive input content, filter the input content, and produce filtered content. Each of the stages is configured to transfer the filtered content to the subsequent stages. The system includes training logic operably coupled with the stages. The training logic is configured to train the processing of the previous stage using information related to processing by a given subsequent stage, the given subsequent stage receiving content obtained directly from the previous stage or content obtained from at least one stage between the given subsequent stage and the previous stage.
[0017] The content may be spoken content. The filtered content of each stage may include a subset of the received input content. Each stage may be configured to generate filtered content from the input content and transfer it to a less efficient stage, where a given less efficient stage is more powerful than a second, more efficient stage.
[0018] An exemplary embodiment of the present invention is implemented as a computer program product having a computer-usable medium containing computer-readable program code. The computer-readable code can be read and used by a computer system according to a conventional process. [Brief explanation of the drawing]
[0019] Those skilled in the art will better understand the advantages of various embodiments of the present invention from the following “Description of Exemplary Embodiments,” which will be discussed with reference to the drawings summarized below. [Figure 1A]A diagram schematically showing a system for content moderation according to an exemplary embodiment of the present invention. [Figure 1B] A diagram schematically showing an alternative configuration of the system for content moderation in FIG. 1A. [Figure 1C] A diagram schematically showing an alternative configuration of the system for content moderation in FIG. 1A. [Figure 2] A diagram schematically showing details of a content moderation system according to an exemplary embodiment of the present invention. [Figure 3A] A diagram showing a process for determining whether a speech is harmful according to an exemplary embodiment of the present invention. [Figure 3B] A diagram showing a process for determining whether a speech is harmful according to an exemplary embodiment of the present invention. [Figure 4] A diagram schematically showing a received speech according to an exemplary embodiment of the present invention. [Figure 5] A diagram schematically showing a speech chunk segmented by a segmenter according to an exemplary embodiment of the present invention. [Figure 6] A diagram schematically showing details of a system that can be used with the processes of FIGS. 3A - 3B according to an exemplary embodiment. [Figure 7] A diagram schematically showing a four - stage system according to an exemplary embodiment of the present invention. [Figure 8A] A diagram schematically showing a process for training machine learning according to an exemplary embodiment of the present invention. [Figure 8B] A diagram schematically showing a system for training the machine learning in FIG. 8A according to an exemplary embodiment of the present invention.
[0020] Note that the foregoing figures and the elements depicted therein are not necessarily drawn to consistent or any scale. Unless the context suggests otherwise, like elements are denoted by like numerals. The drawings are primarily for illustrative purposes and are not intended to limit the scope of the inventive subject matter described herein.
[0021] Description of Exemplary Embodiments In an exemplary embodiment, a content moderation system analyzes an utterance or its characteristics to determine the likelihood that the utterance is harmful. The system uses multi-stage analysis to increase cost efficiency and reduce computing requirements. A series of stages communicate with each other. Each stage filters out non-harmful utterances and sends potentially harmful utterances or data representing them to subsequent stages. Subsequent stages use more reliable (e.g., computationally expensive) analysis techniques than the previous stage. Thus, the multi-stage system can filter the utterances most likely to be harmful to more reliable and computationally expensive stages. The results of subsequent stages can be used to retrain the previous stage. Thus, the exemplary embodiment provides triage of input utterances, filtering out non-harmful utterances so that later, more complex stages do not have to operate on as many input utterances.
[0022] Furthermore, in various embodiments, the stages are adaptive, receiving feedback regarding the correctness of filtering decisions from subsequent stages or external judgments, updating the filtering process each time more data passes through the system, and more appropriately separating utterances likely to be harmful from utterances likely not to be harmful. Such tuning may be done automatically or manually by a trigger; continuously or periodically (often training a batch of feedback at a time).
[0023] For clarity, various embodiments can refer to user utterances or their analysis. While the term "utterance" is used, it should be understood that the system does not necessarily receive or "hear" utterances directly in real time. When a particular stage receives an "utterance," that utterance may contain some or all of a previous "utterance" and / or data representing that utterance or part thereof. The data representing an utterance can be encoded in various ways, such as being a raw speech sample represented by pulse code modulation (PCM), e.g., linear pulse code modulation, or encoded through A-law or u-law quantization. Utterances may also be in forms other than raw speech, such as being represented as a spectrogram, Mel-frequency sepstrum coefficients, a cochleogram, or other representations of the utterance generated by signal processing. Utterances can be filtered (e.g., band-pass or compression). Utterance data can be represented in additional forms of data derived from the utterance, such as frequency peaks and amplitudes, distributions on phonemes, or abstract vector representations generated by neural networks. Data can be uncompressed, or input in various lossless formats (e.g., FLAC or WAVE) or lossy formats (e.g., MP3 or Opus), or, in the case of other representations of speech, input as image data (PNG, JPEG, etc.) or encoded in a custom binary format. Therefore, although the term "speech" is used, it should be understood that this is not limited to human-audible audio files. Furthermore, in some embodiments, other types of media such as images or videos may be used.
[0024] Automated moderation is primarily used in text-based media, such as social media posts or text chats in multiplayer video games. Its basic form typically involves a blacklist of prohibited words or phrases that are matched against the text content of the media. If matches are found, the matching words may be censored, or the writer may be restricted. The system employs fuzzy matching techniques to circumvent simple circumvention methods, such as users replacing letters with similarly shaped numbers or omitting vowels. While scalable and cost-effective, traditional automated moderation is generally considered relatively easy to bypass with minimal creativity, lacks sophistication to detect inappropriate behavior beyond the use of simple keywords or short phrases, and struggles to adapt to new communities or platforms, or to the changing terminology and communication styles of existing communities. Some examples of traditional automated moderation include moderation against illegal videos and images, or the illegal use of copyrighted works. In such cases, media outlets are often hashed to provide a more compact representation of their content, a blacklist of hashes is created, and then new content is hashed and matched against the blacklist.
[0025] In contrast, manual moderation generally employs a team of humans who consume a portion of the content transmitted on the platform and determine whether that content violates the platform's policies. Typically, the team can only monitor a fraction of the content transmitted on the platform. Therefore, a selection mechanism is employed to determine which content the team should inspect. This is usually done through user reports, where users consuming the content can flag other users participating in offensive behavior. Content communicated between users is queued and inspected by human moderators, who make judgments based on the context of the communication and apply penalties.
[0026] Manual moderation presents additional challenges. Hiring humans is costly, and moderation teams are small, meaning only a small fraction of platform content can be manually determined to be safe to consume, forcing the platform to default to allowing most content without moderation. The queue of reported content can easily overflow, especially with hostile behavior, sometimes overwhelming the moderation team with all users simultaneously participating in the abusive behavior, or conversely, all users reporting harmless content, making the selection process inefficient. Human moderation is also time-consuming; humans must receive, understand, and then respond to the content, making low-latency actions like censorship impossible on content-heavy platforms. This problem is exacerbated by the potential for selection queues to become saturated, leading to delays in content while queues are processed. Moderation also burdens human teams; team members are directly exposed to large volumes of offensive content, potentially leading to emotional impact. The high cost of maintaining such teams can lead to long working hours for team members, potentially leaving them with little access to resources that could help address the issue.
[0027] Current content moderation systems known to the inventors are either too simple to effectively prevent abusive behavior or too expensive to scale to large volumes of content. These systems take a long time to adapt to changing environments or new platforms. Sophisticated systems, in addition to being expensive, typically have significant latency between content transmission and moderation, making real-time response or censorship extremely difficult at scale.
[0028] In an exemplary embodiment, the improved moderation platform is implemented as a series of adaptive triage stages, each stage filtering out content that it can reliably determine to be non-nuisance from the stages in the series, and passing on content that cannot be filtered to subsequent stages. Upon receiving information that the filtered content was deemed nuisance or not at a later stage, the stage can update itself to perform filtering more effectively on subsequent content. By chaining several of these stages in sequence, content is triaged down to a manageable level that can be handled by a human team or a further autonomous system. Each stage filters a portion of the incoming content, achieving a reduction in the amount of content that needs to be moderated at subsequent stages (e.g., exponentially).
[0029] Figure 1A schematically illustrates a system 100 for content moderation according to an exemplary embodiment of the present invention. While the system 100 described with reference to Figure 1A moderates audio content, those skilled in the art will understand that various embodiments can be modified to moderate other types of content (e.g., media, text, etc.) in a similar manner. Additionally or alternatively, the system 100 can also assist a human moderator 106 in identifying utterances 110 that are most likely to be harmful. While the system 100 can be applied in a variety of settings, it may be particularly useful in video games. Global revenues in the video game industry are strong, projected to increase by 20% annually in 2020. This projected growth is partly due to new gamers (i.e., users) joining video games and the increasing availability of voice chat as a game option. Many other voice chat options exist outside of games. While voice chat is a desirable feature in many online platforms and video games, user safety is a critical consideration. The prevalence of online toxicity through harassment, racism, sexism, and other types of toxicity is detrimental to the user's online experience and can lead to decreased use of voice chat and / or safety concerns. Therefore, there is a need for a system 100 that can efficiently (i.e., cost and time) identify toxic content (e.g., racism, sexism, and other forms of bullying) from a very large volume of content (e.g., all voice chat communications in video games).
[0030] To this end, system 100 provides an interface between multiple users, such as a speaker 102, a listener 104, and a moderator 106. The speaker 102, listener 104, and moderator 106 can communicate via a network 122 provided by predetermined platforms such as Fortnite, Call of Duty, Roblox, and Halo, streaming platforms such as YouTube and Twitch, and other social apps such as Discord, WhatsApp, Clubhouse, and dating platforms.
[0031] To facilitate discussion, Figure 1A shows an utterance 110 flowing in one direction (i.e., toward the receiver 104 and the moderator 106). In practice, the receiver 104 and / or the moderator 106 may communicate bidirectionally (i.e., the receiver 104 and / or the moderator 106 may be speaking with the utterance 102). However, a single utterance 102 is used as an example to illustrate the operation of system 100. Furthermore, there may be multiple receivers 104, some or all of whom may be utterance 102 (for example, in the context of video game voice chat, all participants are both utterance 102 and receiver 104). In various embodiments, system 100 operates similarly with each utterance 102.
[0032] Additionally, when determining the harmfulness of an utterance from a given speaker, information from other speakers can be combined and used. For example, participant (A) might insult another participant (B), and B might use vulgar language to defend themselves. The system can determine that B's words are not harmful because they are used for self-defense, but A's words are harmful. Alternatively, system 100 may determine that both are harmful. This information is consumed by inputting it into one or more stages of the system (usually later stages that perform more complex processing), which could be any stage or all stages.
[0033] System 100 includes several stages 112-118, each configured to determine whether an utterance 110, or its representation, may be considered harmful (for example, according to a company's policy defining "harmfulness"). In various embodiments, a stage is a logical or abstract entity defined by its interface. It has an input (several utterances) and two outputs (filtered utterances and discarded utterances) (however, it may or may not have additional inputs—such as session context—or additional outputs—such as an estimate of the speaker's age), and it receives feedback from later stages (and may also provide feedback to earlier stages). These stages are, of course, physically implemented and therefore typically run on hardware such as a general-purpose computer (CPU or GPU), and are software / code (individual programs implementing logic such as digital signal processing, neural networks, or a combination thereof). However, they can also be implemented as FPGAs, ASICs, analog circuits, etc. Typically, a stage has one or more algorithms that run on the same or adjacent hardware. For example, one stage may be a keyword detector running on the speaker's computer. Another stage could be a speech recognition engine (transcription engine) running on a GPU, followed by speech recognition interpretation logic running on the CPU of the same computer. Alternatively, the stage could consist of multiple neural networks whose outputs are combined and filtered at the end, running on different computers but within the same cloud (e.g., AWS).
[0034] Figure 1A shows four stages 112-118. However, it should be understood that fewer or more stages may be used. Some embodiments may have only a single stage 112, but preferred embodiments have two or more stages for efficiency purposes, as will be discussed later. Furthermore, stages 112-118 may be entirely on the user device 120, on the cloud server 122, and / or distributed across the user device 120 and the cloud 122, as shown in Figure 1A. In various embodiments, stages 112-118 may be on a server of the platform 122 (e.g., the game network 122).
[0035] The first stage 112 may be located on the speaker device 120 and receives the utterance 110. The speaker device 120 may be, in particular, a mobile phone (e.g., an iPhone), a video game system (e.g., PlayStation, Xbox), and / or a computer (e.g., a laptop or desktop computer). The speaker device 120 may have an integrated microphone (e.g., the iPhone microphone) or may be coupled to a microphone (e.g., a headset with a USB or AUX microphone). The listener device may be identical or similar to the speaker device 120. By providing one or more stages on the speaker device 120, it becomes possible to perform the processing of implementing one or more stages on hardware owned by the speaker 102. Typically, this means that the software that implements the stages runs on the speaker 102's hardware (CPU or GPU), but in some embodiments, the speaker 102 may have a dedicated hardware unit (e.g., a dongle) that attaches to its device. In some embodiments, one or more stages may be located on the listener device.
[0036] As will be explained in more detail below, the first stage 112 receives a large volume of utterances 110. For example, the first stage 112 may be configured to receive all utterances 110 made by speaker 102 and received by device 120 (e.g., a continuous stream during a call). Alternatively, the first stage 112 may be configured to receive utterances 110 when a specific trigger is met (e.g., a video game application is active and / or the user presses a chat button). In a use case scenario, the utterances 110 may be utterances intended to be received by receiver 104, such as team voice communication in a video game.
[0037] The first stage 112 is trained to determine whether any of the utterances 110 have a high probability of being harmful (i.e., contain a harmful utterance). In an exemplary embodiment, the first stage 112 analyzes the utterances 110 using an efficient method (i.e., a computationally efficient method and / or a low-cost method) compared to subsequent stages. The efficient method used by the first stage 112 may not be as accurate in detecting harmful utterances as subsequent stages (e.g., stages 114-118), but the first stage 112 generally receives more utterances 110 than subsequent stages 114-118.
[0038] If the first stage 112 does not detect a likelihood of a harmful utterance, the utterance 110 is discarded (shown as discarded utterance 111). However, if the first stage 112 determines that there is a likelihood that a portion of the utterance 110 is harmful, a subset of the utterance 110 is sent to a subsequent stage (e.g., second stage 114). In Figure 1A, the transferred / uploaded subset is a filtered utterance 124, which includes at least a portion of the utterance 110 that is considered to have a likelihood of containing a harmful utterance. In exemplary embodiments, the filtered utterance 124 is preferably a subset of the utterance 110 and is therefore represented by a smaller arrow. However, in some other embodiments, the first stage 112 may transfer all of the utterance 110.
[0039] Furthermore, when describing an utterance 110, it should be made clear that the utterance 110 may refer to a specific parsing chunk. For example, the first stage 112 may receive a 60-second utterance 110, and the first stage 112 may be configured to parse the utterance at 20-second intervals. Thus, there are three 20-second chunks of utterance 110 to be parsed. Each utterance chunk may be parsed independently. For example, the first 20-second chunk may not have a high probability of being harmful and may be discarded. The second 20-second chunk may meet the threshold for high probability of being harmful and may therefore be passed on to a subsequent stage. The third 20-second chunk may not have a high probability of being harmful and may also be discarded. Thus, references to discarding and / or passing on an utterance 110 relate to a specific segment of utterance 110 parsed by a given stage 112-118, as opposed to a universal determination of all utterances 110 from speaker 102.
[0040] The filtered utterances 124 are received by the second stage 114. The second stage 114 is trained to determine whether any of the utterances 110 have a high probability of being harmful. However, the second stage 114 generally uses a different analysis method than the first stage 112. In exemplary embodiments, the second stage 114 analyzes the filtered utterances 124 using a computationally more computationally intensive method than the previous stage 112. Therefore, the second stage 114 may be considered less efficient than the first stage 112 (i.e., a computationally less efficient and / or more expensive method compared to the previous stage 112). However, the second stage 114 is considerably more likely to be accurate than the first stage 112 in performing rigorous detection of harmful utterances 110. Furthermore, while the subsequent stage 114 may be less efficient than the previous stage 112, this does not necessarily mean that the time it takes for the second stage 114 to analyze the filtered utterance 124 is longer than the time it takes for the first stage 112 to analyze the initial utterance 110. This is partly because the filtered utterance 124 is a subsegment of the initial utterance 110.
[0041] Similar to the process described above with reference to the first stage 112, the second stage 114 analyzes the filtered utterance 124 and determines whether the filtered utterance 124 has a high probability of being harmful. If not, the filtered utterance 124 is discarded. If there is a high probability of being harmful (for example, if the probability is determined to be above a predetermined harmfulness likelihood threshold), the filtered utterance 126 is passed on to the third stage 116. It should be understood that the filtered utterance 126 may be the entire filtered utterance 124, a chunk 110A, and / or subsegments. However, the filtered utterance 126 is represented by a smaller arrow than the filtered utterance 124 because, generally, a portion of the filtered utterance 124 is discarded by the second stage 114, and therefore less filtered utterance 126 is passed on to the subsequent third stage 116.
[0042] This process, which involves analyzing utterances in subsequent stages using more computationally intensive analysis methods, can be repeated for as many stages as needed. In Figure 1A, this process is repeated in the third stage 116 and the fourth stage 118. Similar to the previous stage, the third stage 116 filters out utterances that are less likely to be harmful and sends the filtered utterances 128 that are more likely to be harmful to the fourth stage 118. The fourth stage 118 uses an analysis method to determine whether the filtered utterances 128 contain harmful utterances 130. The fourth stage 118 may discard those that are less likely to be harmful, or it may take over the utterances 130 that are more likely to be harmful. This process may end in the fourth stage 118 (or other stages depending on the desired number of stages).
[0043] System 100 can perform an automatic determination of the harmfulness of an utterance after the final stage 118 (i.e., whether the utterance is harmful and, if necessary, what action would be appropriate). However, in other embodiments, and as shown in Figure 1A, the final stage 118 (i.e., the least computationally efficient but most accurate stage) may provide a human moderator 106 with what is considered to be a harmful utterance 130. The human moderator can listen to the harmful utterance 130 and determine whether the utterance 130, which System 100 has determined to be harmful, is actually a harmful utterance (for example, according to the company's policy on harmful utterances).
[0044] In some embodiments, one or more non-final stages 112-116 may determine that an utterance is "definitely harmful" (e.g., 100% certain that the utterance is harmful) and make a determination that bypasses subsequent and / or final stages 118 collectively (e.g., by forwarding the utterance to a human moderator or other system). Furthermore, the final stage 118 may provide what is considered a harmful utterance to an external processing system, which itself may determine whether the utterance is harmful or not (thus functioning like a human moderator, but being automated). For example, some platforms may have a reputation system configured to receive harmful utterances and further process them automatically using the history of the speaker 102 (e.g., a video game player).
[0045] Moderator 106 makes a determination as to whether the harmful utterance 130 is harmful or not, and provides moderator feedback 132 back to the fourth stage 118. The feedback 132 may be received directly by the fourth stage 118 and / or by a database containing the training data of the fourth stage 118, and this feedback is then used to train the fourth stage 118. Thus, the feedback can instruct the final stage 118 on whether the harmful utterance 130 was correctly or incorrectly determined (i.e., whether a true positive or false positive determination was made). Thus, the final stage 118 can be trained to improve its accuracy over time using human moderator feedback 132. Generally, the resources (i.e., effort) of the human moderator 106 available to review harmful utterances are considerably less than the throughput processed by the various stages 112-118. By filtering the initial utterance 110 through a series of stages 112-118, the human moderator 106 sees only a small portion of the initial utterance 110 and, moreover, favorably receives the utterances 110 that are most likely to be harmful. As an additional benefit, the final stage 118 is trained to more accurately identify harmful utterances using human moderator feedback 132.
[0046] Each stage may process all of the information within the filtered audio clip, or only a portion of the information within that clip. For example, for computational efficiency, stages 112–118 may process only small windows of utterances looking for individual words or phrases, requiring only a small amount of context (e.g., 4 seconds of utterance instead of the entire 15-second clip). Stages 112–118 can also use additional information from previous stages (e.g., calculation of perceptual loudness over the duration of the clip) to determine which regions of utterance 110 clip may contain utterances, and thus dynamically determine which parts of utterance 110 clip to process.
[0047] Similarly, a subsequent stage (e.g., Stage 4 118) can provide feedback 134-138 to a previous stage (e.g., Stage 3 116) regarding whether the previous stage accurately determined that the utterance was harmful. Those skilled in the art will understand that the term "accurately" is used here, relating to the probability that the utterance determined by the stage is harmful, and not necessarily true accuracy. Of course, the system is configured to be trained to become more truly accurate in accordance with the harmfulness policy. Thus, Stage 4 118 can train Stage 3 116, Stage 3 116 can train Stage 2 114, and Stage 2 112 can train Stage 1 112. As mentioned above, feedback 132-138 may be received directly by the previous stages 112-118, or provided to a training database used to train each of the stages 112-118.
[0048] Figures 1B to 1C schematically show a system 100 for content moderation in an alternative configuration according to an exemplary embodiment of the present invention. As illustrated and described, various stages 112 to 118 may be on the speaker device 120 and / or on the platform server 122. However, in some embodiments, the system 100 may be configured so that a user utterance 110 reaches the receiver 104 without passing through the system 100, or only by passing through one or more stages 112 to 114 on the user device 120 (for example, as shown in Figure 1B). However, in some other embodiments, the system 100 may be configured so that a user utterance 110 reaches the receiver 104 after passing through various stages 112 to 118 of the system 100 (as shown in Figure 1C).
[0049] The inventors hypothesize that the configuration shown in Figure 1B may increase the waiting time for receiving the utterance 110. However, by passing through stages 112-114 on the user device 120, it may be possible to take corrective actions and moderate the content before it reaches the intended recipient (e.g., the listener 104). This is also true for the configuration shown in Figure 1C, and the waiting time would be further increased, especially considering that the utterance information passes through the cloud server before reaching the listener 104.
[0050] Figure 2 schematically illustrates the details of a speech moderation system 100 according to an exemplary embodiment of the present invention. The system 100 has an input unit 208 configured to receive utterances 110 from a speaker 102 and / or a speaker device 120 (e.g., as an audio file). It should be understood that the reference to utterances 110 includes not only audio files but also other digital representations of utterances 110. The input includes a temporal receptive field 209 configured to divide utterances 110 into utterance chunks. In various embodiments, machine learning 215 determines whether the entire utterance 110 and / or utterance chunks contain harmful utterances.
[0051] The system also includes a stage converter 214 configured to receive an utterance 110 and convert the utterance in a meaningful manner that can be interpreted by stages 112-118. Furthermore, the stage converter 214 enables communication between stages 112-118 by receiving filtered utterances 124, 126, or 128 and converting the filtered utterances 124, 126, and 128 so that each of the stages 114, 116, and 118 can receive and analyze the utterance.
[0052] System 100 includes a user interface server 210 configured to provide a user interface that allows a moderator 106 to communicate with System 100. In various embodiments, the moderator 106 can listen to (or read a transcript of) utterances 130 that have been determined to be harmful by System 100. Furthermore, the moderator 106 can provide feedback through the user interface regarding whether or not the harmful utterances 130 are actually harmful. The moderator 106 may access the user interface via an electronic device (e.g., a computer, smartphone, etc.) and use the electronic device to provide feedback to the final stage 118. In some embodiments, the electronic device may be a network device such as a smartphone or desktop computer connected to the Internet.
[0053] The input unit 208 is also configured to receive the voice of speaker 102 and map the voice of speaker 102 to a voice database 212, also referred to as the timbre vector space 212. In various embodiments, the timbre vector space 212 may also include a voice mapping system 212. The timbre vector space 212 and the voice mapping system 212 have been previously invented by the inventors and, in particular, described in U.S. Patent No. 1,0861,476, which is incorporated herein by reference in its entirety. The timbre vector space 212 is a multidimensional discrete or continuous vector space that represents encoded voice data. This representation is referred to as a voice "mapping". When encoded voice data is mapped, the vector space 212 characterizes the voices and arranges them relative to each other based on these characterizations. For example, part of the representation may relate to pitch or the gender of the speaker. The timbre vector space 212 maps speeches relative to each other so that mathematical operations can be performed on speech coding and qualitative and / or quantitative information (e.g., the identity, gender, race, and age of the speaker 102) can be obtained from the speech. However, it should be understood that various embodiments do not require the entire timbre mapping component / timbre vector space 112. Instead, information such as gender / race / age can be extracted independently via a separate neural network or other system.
[0054] System 100 also includes a toxicity machine learning 215 configured to determine the likelihood (i.e., confidence interval) that an utterance 110 contains toxicity for each stage. The toxicity machine learning 215 operates for each stage 112-118. For example, for a given amount of utterances 110, the toxicity machine learning 215 may determine that the confidence level of toxicity is 60% in the first stage 112 and 30% in the second stage 114. An exemplary embodiment may include a separate toxicity machine learning 215 for each of the stages 112-118. However, for convenience, various components of the toxicity machine learning 215 that may be distributed across various stages 112-118 are shown as being within a single toxicity machine learning component 215. In various embodiments, the toxicity machine learning 215 may be one or more neural networks.
[0055] In each stage 112-118, the harmfulness machine learning 215 is trained to detect harmful utterances 110. To do this, the machine learning 215 communicates with a training database 216 that contains relevant training data. The training data in the database 216 may include a library of utterances classified as harmful and / or non-harmful by a trained human operator.
[0056] The toxic machine learning system 215 has an utterance segmentator 234 configured to segment the received utterance 110 and / or chunk 110A into segments, which are then analyzed. These segments are referred to as analysis segments and are considered to be part of the utterance 110. For example, speaker 102 may provide utterance 110 totaling 1 minute. The segmentator 234 may segment the utterance 110 into three 20-second intervals, each of which is analyzed independently by stages 112-118. Furthermore, the segmentator 234 may be configured to segment the utterance 110 into segments of different lengths for different stages 112-118 (e.g., two 30-second segments for the first stage, three 20-second segments for the second stage, four 15-second segments for the third stage, and five 10-second segments for the fifth stage). Furthermore, the segmentator 234 can segment the utterance 110 into overlapping intervals. For example, a 30-second segment of the utterance 110 may be segmented into five segments (e.g., 0-10 seconds, 5-15 seconds, 10-20 seconds, 15-25 seconds, and 20-30 seconds).
[0057] In some embodiments, the segmentator 234 can segment later stages into longer segments than the previous stage. For example, a subsequent stage 112 might attempt to combine previous clips to obtain a broader context. The segmentator 234 can accumulate multiple clips to obtain additional context and then pass through the entire clip. This can also be done dynamically, for example, by accumulating utterances within a clip up to a silent region (e.g., 2 seconds or more) and then sending the accumulated clip all at once. In this case, even if the clips were input as separate individual clips, the system will subsequently treat the accumulated clip as a single clip (and thus make a single decision, for example, about filtering or discarding utterances).
[0058] The machine learning 215 may include an uploader 218 (which may be a random uploader) configured to upload or pass through a small percentage of discarded utterances 111 from each stage 112-118. Thus, the random uploader module 218 is configured to assist in determining the false negative rate. In other words, if the first stage 112 discards utterance 111A, a small portion of that utterance 111A is taken up by the random uploader module 218 and sent to the second stage 114 for analysis. Thus, the second stage 114 can determine whether the discarded utterance 111A was correctly identified as not actually harmful or incorrectly identified (i.e., a false negative, or a true negative of likely harmful). This process can be repeated for each stage (for example, discarded utterance 111B is analyzed by the third stage 116, discarded utterance 111C is analyzed by the fourth stage, discarded utterance 111D is analyzed by the moderator 106).
[0059] Various embodiments aim to improve efficiency by minimizing the amount of utterances uploaded / analyzed by the upper stages 114-118 or moderator 106. However, various embodiments sample only a small percentage of discarded utterances 111, such as less than 1% of discarded utterances 111, preferably less than 0.1% of discarded utterances 111. The inventors believe that this small sampling rate of discarded utterances 111 advantageously trains system 100 to reduce false negatives without placing an excessive burden on system 100. Thus, system 100 efficiently checks for false negatives (by minimizing the amount of information to check) and improves the false negative rate over time. This is important because it is advantageous to correctly identify harmful speech, but it is also advantageous not to misidentify harmful speech.
[0060] The toxicity threshold setter 230 is configured to set toxicity likelihood thresholds for each stage 112-118. As previously described, each stage 112-118 is configured to determine / output toxicity confidence. This confidence is used to determine whether a segment of utterance 110 should be discarded 111 or filtered and passed on to a subsequent stage. In various embodiments, confidence is compared to a threshold that can be adjusted by the toxicity threshold setter 230. The toxicity threshold setter 230 can be automatically adjusted by training a neural network over time to increase the threshold as false negatives and / or false positives decrease. Alternatively, or additionally, the toxicity threshold setter 230 can also be adjusted by a moderator 106 via a user interface 210.
[0061] The machine learning 215 may also include a session context flagger 220. The session context flagger 220 communicates with various stages 112-118 and is configured to provide one or more stages 112-118 with an indication (session context flag) that a previous harmful utterance was determined by another stage 112-118. In various embodiments, the previous indication may be session or time-limited (e.g., a harmful utterance 130 determined by the final stage 118 within the last 15 minutes). In some embodiments, the session context flagger 220 may be configured to receive flags only from subsequent stages or specific stages (e.g., the final stage 118).
[0062] The machine learning 215 may also include an age analyzer 222 configured to determine the age of the speaker 102. The age analyzer 222 may be provided with a training dataset of various speakers paired with the speaker's age. Thus, the age analyzer 222 can analyze the utterance 110 to determine the speaker's approximate age. Using the speaker 102's approximate age, the toxicity threshold for a particular stage can be adjusted by communicating with the toxicity threshold setter 230 (for example, the threshold may be lowered for teenagers, as they are considered more likely to be toxic). Additionally or alternatively, the speech of the speaker 102 may be mapped to a speech timbre vector space 212, from which their age may be estimated.
[0063] The emotion analyzer 224 may be configured to determine the emotional state of the speaker 102. The emotion analyzer 224 may be provided with training datasets of various speakers paired with emotions. Thus, the emotion analyzer 224 can analyze the utterance 110 to determine the speaker's emotion. The harm threshold for a particular stage can be adjusted by communicating with the harm threshold setter using the emotion of the speaker 102. For example, an angry speaker is considered more likely to be harmful, so the threshold may be lowered.
[0064] The user context analyzer 226 may be configured to determine the context in which speaker 102 provided utterance 110. The context analyzer 226 may be provided with access to account information for a particular speaker 102 (e.g., by the platform or video game subscribed to by speaker 102). This account information may include, among other things, the user's age, geographical location, friend list, history of recently interacted users, and other activity history. Furthermore, where applicable in the context of video games, it may also include the user's game history, including gameplay time, game length, start and end times of the game, and, where applicable, recent user-to-user activity such as death or killing (e.g., in a shooting game).
[0065] For example, the user's geographical location can be used to assist in language analysis to avoid confusing harmless language in one language with harmful speech in another. Furthermore, the user context analyzer 226 can adjust the harmfulness threshold by communicating with the threshold setter 230. For example, the harmfulness threshold may be increased for utterance 110 in communication with someone on the user's friend list (e.g., aggressive speech may be made as a joke with a friend). Another example is adjusting the harmfulness threshold downward using recent deaths in video games or a low overall team score (e.g., if speaker 102 is losing the game, it is more likely to be harmful). Yet another example is adjusting the harmfulness threshold using the time of day of utterance 110 (e.g., utterance 110 at 3 AM is more likely to be harmful than utterance 110 at 5 PM, and therefore the harmful speech threshold is reduced).
[0066] In various embodiments, the hazard machine learning 215 may include a speech recognition engine 228. The speech recognition engine 228 is configured to transcribe the utterance 110 into text. The text may then be used by one or more stages 112-118 to analyze the utterance 110, or it may be provided to a moderator 106.
[0067] The feedback module 232 receives feedback from each of the subsequent stages 114-118 and / or moderator 106 regarding whether the filtered utterances 124, 126, 128 and / or 130 were deemed harmful. The feedback module 232 can provide this feedback to the preceding stages 112-118 to update their training data (for example, directly or by communicating with the training database 216). For example, the training data for the fourth stage 118 may include negative examples, such as the display of a harmful utterance 130 that was escalated to a human moderator 106 that was not deemed harmful. The training data for the fourth stage 118 may also include positive examples, such as the display of a harmful utterance 130 that was escalated to a human moderator 106 that was deemed harmful.
[0068] Each of the above components of system 100 can operate in multiple stages 112-118. Additionally or alternatively, each of stages 112-118 may have one or all of the components as dedicated components. For example, each stage 112-118 may have a stage converter 214, or system 100 may have a single stage converter 214. Furthermore, various machine learning components, such as a random uploader 218 or a speech recognition engine 228, can operate in one or more of stages 112-118. For example, all stages 112-118 may use the random uploader 218, but only the final stage may use the speech recognition engine 228.
[0069] Each of the above components is operablely connected by any conventional interconnection mechanism. Figure 2 shows a simplified bus 50 communicating the components. Those skilled in the art will understand that this generalized representation can be modified to include other conventional direct or indirect connections. Therefore, the discussion of bus 50 is not intended to limit the various embodiments.
[0070] It should be noted that Figure 2 is only a schematic representation of each of these components. Those skilled in the art will understand that each of these components can be implemented in various conventional ways, such as using hardware, software, or a combination of hardware and software, across one or more other functional components. For example, the speech recognition engine 228 can be implemented using multiple microprocessors running firmware. Another example is the speech segmentator 234, which can be implemented using one or more application-specific integrated circuits (i.e., "ASICs") and associated software, or a combination of ASICs, discrete electronic components (e.g., transistors), and microprocessors. Therefore, the representation of the segmentator 234, speech recognition engine 228, and other components in a single box in Figure 2 is for simplification purposes only. In practice, in some embodiments, the speech segmentator 234 may be distributed across multiple different machines and / or servers, not necessarily within the same housing or chassis. Of course, the machine learning 215 and other components of system 100 can also have implementations similar to those described above for the speech recognition engine 228.
[0071] Additionally, in some embodiments, components shown separately (e.g., the age analyzer 222 and the user context analyzer 226) may be replaced by a single component (e.g., the user context analyzer 226 for the entire machine learning system 215). Furthermore, certain components and sub-components in Figure 2 are optional. For example, in some embodiments, the sentiment analyzer 224 may not be used. As another example, in some embodiments, the input 108 may not have a temporal receptive field 109.
[0072] It should be reiterated that the representation in Figure 2 is a simplified representation. Those skilled in the art should understand that such a system is likely to have many other physical and functional components, such as a central processing unit, other packet processing modules, and short-term memory. Therefore, this discussion is not intended to suggest that Figure 2 represents all elements of various embodiments of the voice moderation system 100.
[0073] Figures 3A and 3B illustrate a process 300 for determining whether an utterance 110 is harmful, according to an exemplary embodiment of the present invention. It should be noted that this process simplifies a longer process typically used to determine whether an utterance 110 is harmful. Therefore, the process for determining whether an utterance 110 is harmful is likely to have many steps that a person skilled in the art would likely use. Furthermore, some of the steps may be performed in a different order than those shown, or may be skipped entirely. Additionally or alternatively, some of the steps may be performed simultaneously. Therefore, a person skilled in the art can modify this process as appropriate.
[0074] Furthermore, the discussion of specific implementation examples of the stages with reference to Figures 3A and 3B is for illustrative purposes only and is not intended to limit the various embodiments. Those skilled in the art will understand that the training of the stages and the various components and interactions of the stages can be adjusted, removed, and / or added while developing a toxicity moderation system 100 operating according to the exemplary embodiments.
[0075] Figures 1A to 1C show four stages 112 to 118 as an example, so each stage 112 to 118 is referred to with a separate reference number. However, from now on, when referring to any stage 115, one or more stages will be referred to with a single reference number 115. It should be understood that a reference to stage 115 does not mean that stage 115 is identical or that stage 115 is limited to any particular order of system 100 or the previously described stages 112 to 118, unless the context specifically requests otherwise. The reference number for stage 112 can be used to refer to a preceding or preceding stage 112 of system 100, regardless of the actual number of stages (e.g., 2 stages, 5 stages, 10 stages, etc.), and the reference number for stage 118 can be used to refer to a subsequent or later stage 112 of system 100. Thus, a stage referred to as stage 115 is similar to or the same as stages 112 to 118, and vice versa.
[0076] Process 300 begins in step 302 by setting the toxicity threshold for stage 115 of system 100. The toxicity threshold for each stage 115 of system 100 may be set automatically by system 100, by moderator 106 (e.g., via a user interface), manually by a developer, community manager, or by another third party. For example, the first stage 115 may have a toxicity threshold of 60% for any given utterance 110 being analyzed, where the likelihood of it being toxicity is 60%. If the machine learning 215 of the first stage 115 determines that there is a 60% or greater probability that the utterance 110 is toxicity, the utterance 110 is determined to be toxicity and is passed on to subsequent stages 115 or "filtered through". Those skilled in the art will understand that when it is mentioned that an utterance has been determined to be toxicity by stage 115, this does not necessarily mean that the utterance is actually toxicity according to company policy, nor does it necessarily mean that subsequent stages 115 (if any) agree that the utterance is toxicity. If the likelihood of an utterance being harmful is less than 60%, the utterance 110 is discarded or "filtered out" and not sent to the subsequent stage 115. However, as described below, in some embodiments, a random uploader 218 can be used to analyze a portion of the filtered out utterance 111.
[0077] In the example above, the toxicity threshold is described as an inclusive range (i.e., the 60% threshold is achieved up to 60%). In some embodiments, the toxicity threshold may be an exclusive range (i.e., the 60% threshold is achieved only by likelihoods greater than 60%). Furthermore, in various embodiments, the threshold does not necessarily have to be expressed as a percentage, but can be expressed in some other form that represents the likelihood of toxicity (e.g., a representation that is not understandable to humans but understandable to the neural network 215).
[0078] The second stage 115 may have its own toxicity threshold such that any utterance that does not meet the toxicity threshold analyzed by the second stage 115 is discarded. For example, the second stage may have a threshold of toxicity of 80% or higher. If an utterance has a toxicity likelihood greater than the toxicity threshold, the utterance is transferred to the subsequent third stage 115. Transferring an utterance 110 to the next stage can also be referred to as "uploading" the utterance 110 (for example, to a server that the subsequent stage 115 can access the uploaded utterance 110). If an utterance does not meet the threshold of the second stage 115, the utterance is discarded. This process of setting toxicity thresholds can be repeated for each stage 115 of the system 100. Therefore, each stage may have its own toxicity threshold.
[0079] Next, the process proceeds to step 304, which receives an utterance 110 from the speaker 102. The utterance 110 is first received by the input unit 208 and then by the first stage 112. Figure 4 schematically shows the received utterance 110 according to an exemplary embodiment of the present invention. For illustrative purposes, it is assumed that the first stage 112 is configured to receive 10 seconds of audio input at a time, which is segmented into 2-second 50% overlap sliding windows.
[0080] The time-receptive field 209 decomposes the utterance 110 into utterance chunks 110A and 110B (e.g., 10 seconds) that can be received by the input of the first stage 112. The utterance 110 and / or utterance chunks 110A and 110B can then be processed by the segmentator 234 (e.g., of the first stage 112). For example, as shown in Figure 4, a 20-second utterance 110 can be received by the input 208 and filtered by the time-receptive field 209 into 10-second chunks 110A and 110B.
[0081] Next, the process proceeds to step 306, which segments the utterance 110 into analytical segments. Figure 5 schematically shows the utterance chunk 110A segmented by the segmentator 234 according to an exemplary embodiment of the present invention. As previously described, the utterance segmentator 234 is configured to segment the received utterance 110 into segments 140 which are analyzed by each stage 115. These segments 140 are referred to as analytical segments 140 and are considered to be part of the utterance 110. In the present example, the first stage 112 is configured to analyze segments 140 that are in a 2-second 50% overlap sliding window. Thus, the utterance chunk 110A is broken down into analytical segments 140A to 140I.
[0082] The various analysis segments are executed at 2-second intervals with 50% overlap. Thus, as shown in Figure 5, segment 140A covers the 0:00-0:02 second of chunk 110A, segment 140B covers the 0:01-0:03 second of chunk 110A, segment 140C covers the 0:02-0:04 second of chunk 110A, and so on, for each segment 140 until chunk 110A is completely covered. This process is similarly repeated for subsequent chunks (e.g., 110B). In some embodiments, stage 115 can analyze the entire chunk 110A or all utterances 110, depending on the machine learning model 215 in stage 115. Thus, in some embodiments, all utterances 110 and / or chunks 110A, 110B may be analysis segments 140.
[0083] The short segments 140 (e.g., 2 seconds) analyzed by the first stage 115 make it possible to detect, among other things, whether the speaker 102 is talking, shouting, crying, silent, or saying a specific word. The length of the analyzed segment 140 is preferably long enough to detect some or all of these features. While a short segment 140 may contain several words, it is difficult to detect entire words with high accuracy without more context (e.g., longer segments 140).
[0084] Next, the process proceeds to step 308, which asks whether a session context flag has been received from the context flagger 220. To this end, the context flagger 220 queries the server to determine whether any harmfulness determination has occurred within a predefined period of previous utterances 110 from speaker 102. For example, if utterance 110 from speaker 102 has been determined to be harmful by the final stage 115 within the last two minutes, the session context flag may be received. The session context flag provides context to the stage 115 that received the flag (for example, dirty words detected by another stage 115 may mean that the utterance may be escalating to something harmful). Therefore, if the session context flag is received, the process can proceed to step 310, which reduces the harmfulness threshold of the stage 115 that received the flag. Alternatively, in some embodiments, if the session context flag is received, utterance 110 may be automatically uploaded to the subsequent stage 115. Next, the process proceeds to step 312. If no flag is received, this process proceeds directly to step 312 without adjusting the toxicity threshold.
[0085] In step 312, the process analyzes the utterance 110 (e.g., utterance chunk 110A) using the first stage 115. In this example, the first stage 115 performs machine learning 215 (e.g., a neural network on the speaker device 120) to analyze a 2-second segment 140 and determines individual confidence levels for each segment 140 input. The confidence level may be expressed as a percentage.
[0086] To determine the confidence interval, Stage 115 (e.g., neural network 215) may have been previously trained using a set of training data in the training database 216. The training data for the first Stage 115 may include multiple negative examples of toxicity, i.e., utterances that do not contain toxicity and can be discarded. The training data for the first Stage 115 may also include multiple positive examples of toxicity, i.e., utterances that contain toxicity and should be passed on to the next Stage 115. The training data may be obtained, for example, from professional voice actors. Additionally or alternatively, the training data may be actual utterances pre-classified by a human moderator 106.
[0087] In step 314, the first stage 115 determines the confidence interval for the harm of each of the speech chunks 110A and / or segments 140. The confidence intervals for chunks 110A and 110B can be based on the analysis of various segments 140 from trained machine learning.
[0088] In various embodiments, the first stage 115 provides a toxicity threshold for each segment 140A to 140I. However, step 316 determines whether the utterance 110 and / or utterance chunk 110A meet the toxicity threshold to be carried over to the next stage 115. In various embodiments, the first stage 115 uses a different method to determine the toxicity confidence of utterance chunk 110A based on the various toxicity confidences of segments 140A to 140I.
[0089] The first option is to use the maximum confidence from any given segment as the confidence interval for the entire utterance chunk 110A. For example, if segment 140A is silent, the confidence in its harmfulness is 0%. However, if segment 140B contains offensive language, the confidence in its harmfulness could be 80%. If the harmfulness threshold is 60%, then at least one segment 140B will meet the threshold, and the entire utterance chunk 110A will be passed on to the next stage.
[0090] Another option is to use the average confidence from all segments within utterance chunk 110A as the confidence level for utterance chunk 110A. Therefore, if the average confidence level does not exceed the toxicity threshold, utterance chunk 110A is not forwarded to the subsequent stage 115. A further option is to use the minimum toxicity from any segment 140 as the confidence level for utterance chunk 110A. In the current example provided, using the minimum is undesirable because a large number of potentially harmful utterances would likely be discarded due to silent periods in one of the segments 140. However, in other implementations of stage 115, this may be desirable. A further approach is to use another neural network to learn a function that combines various confidence levels from segment 140 to determine the overall toxicity threshold for utterance chunk 110A.
[0091] Next, the process proceeds to step 316, where it checks whether the toxicity threshold for the first stage 115 is met. If the toxicity threshold for the first stage is met, the process proceeds to step 324, where the filtered harmful utterances 124 are forwarded to the second stage 115. Returning to Figure 1A, it is clear that not all utterances 110 pass through the first stage 115. Therefore, utterances 110 that pass through the first stage 115 are considered to be filtered harmful utterances 124.
[0092] Steps 312-316 are repeated for all remaining chunks 110B.
[0093] If the first stage toxicity threshold is not met in step 316, the process proceeds to step 318, where non-toxic utterances are filtered out. Next, the non-toxic utterances are discarded in step 320, becoming discarded utterance 111.
[0094] In some embodiments, the process proceeds to step 322, where the random uploader 218 sends a small percentage of filtered utterances to the second stage 115 (even though the filtered utterances do not meet the toxicity threshold of the first stage 115). The random uploader 218 sends a small percentage of all filtered utterances (also referred to as negatives) to a subsequent stage 115, which samples a subset of the filtered utterances 124. In various embodiments, a more advanced second stage 115 analyzes the negatives from the first stage 115 in a random proportion.
[0095] As explained earlier, the first stage 115 is generally more computationally efficient than subsequent stages 115. Therefore, the first stage 115 filters out utterances that are less likely to be harmful and leaves utterances that are more likely to be harmful for analysis by more advanced stages 115. It may seem counterintuitive for subsequent stages 115 to analyze the filtered utterances. However, analyzing a small portion of the filtered utterances offers two advantages. First, the second stage 115 detects false negatives (i.e., filtered utterances 111 that should have been forwarded to the second stage 115). These false negatives can be added to the training database 216 to help further train the first stage 115 and reduce the likelihood of further false negatives. Furthermore, the proportion of filtered utterances 111 sampled is small (e.g., 1% to 0.1%), which does not excessively waste resources from the second stage 115.
[0096] An example of the analysis that can be performed by the second stage in step 324 is described below. In various embodiments, the second stage 115 may be a cloud-based stage. The second stage 115 receives the utterance chunk 110A as input, if it has been uploaded by the first stage 115 and / or the random uploader 218. Thus, continuing the previous example, the second stage 115 may receive a 20-second chunk 110A.
[0097] The second stage 115 can be trained using a training dataset that includes, for example, a human moderator 106 that has determined age and emotion category labels corresponding to a dataset of clips of human speakers 102 (e.g., adult and child speakers 102). In an exemplary embodiment, the set of content moderators may manually label data obtained from various sources (e.g., voice actors, Twitch streams, video game voice chats, etc.).
[0098] The second stage 115 can analyze the utterance chunk 110A by running a machine learning / neural network 215 on the 20-second input utterance chunk 110A and generate a harmfulness confidence output. In contrast to the first stage 115, the second stage 115 may analyze the 20-second utterance chunk 110A as a whole unit, unlike the divided segments 240. For example, the second stage 115 may determine that an utterance 110 expressing anger is likely to be harmful. Similarly, the second stage 115 may determine that a teenage speaker 102 is likely to be harmful. Furthermore, the second stage 115 can learn some of the prominent features of a speaker 102 of a specific age (e.g., vocabulary and phrases added to the confidence score).
[0099] Furthermore, Stage 2 115 can be trained using negative and positive examples of speech hazards from subsequent Stages 115 (e.g., Stage 3 115). For example, speech 110 analyzed by Stage 3 115 and found to be non-hazardous can be incorporated into Stage 2 training. Similarly, speech analyzed by Stage 3 115 and found to be hazardous can be incorporated into Stage 2 training.
[0100] Next, the process proceeds to step 326, which outputs a confidence interval for the harmfulness of utterance 110 and / or utterance chunk 110A. The second stage 115, in this example, analyzes the entire utterance chunk 110A, so a single confidence interval is output for the entire chunk 110A. Furthermore, the second stage 115 can also output an estimate of emotion and speaker age based on the timbre in utterance 110.
[0101] Next, the process proceeds to step 328, which asks whether the toxicity threshold of the second stage is met. The second stage 115 has a preset toxicity threshold (e.g., 80%). If the toxicity threshold is met by the confidence interval provided by step 326, the process proceeds to step 336 (shown in Figure 3B). If the toxicity threshold is not met, the process proceeds to step 330. Steps 330-334 operate in a similar manner to steps 318-322. Therefore, the discussion of these steps will not be repeated in much detail here. However, it is worth mentioning again that a small percentage (e.g., less than 2%) of negative (i.e., non-toxic) utterances determined by the second stage 115 are passed through to the third stage 115, which helps to retrain the second stage 115 and reduce false negatives. This process provides similar benefits to those described earlier.
[0102] As shown in Figure 3B, the process proceeds to step 336, which analyzes the harmful utterances filtered out using the third stage 115. The third stage 115 can receive 20 seconds of audio that has been filtered through by the second stage 115. The third stage 115 can also receive an estimate of the speaker's age 102, or the most common age category of speaker 102, from the second stage 115. The age category of speaker 102 may be determined by the age analyzer 222. For example, the age analyzer 222 may analyze multiple parts of utterance 110 and determine that speaker 102 is an adult 10 times and a child 1 time. The most common age category of the speaker is adult. Furthermore, the third stage 115 can receive a transcript of the previous utterance 110 in the conversation that reached the third stage 115. The transcript can be prepared by the speech recognition engine 228.
[0103] The third stage 115 can be initially trained with human-created speech recognition labels corresponding to separate data from the audio clip. For example, a human can rewrite various different utterances 110 and classify their transcripts as either harmful or harmless. Thus, the speech recognition engine 228 can also be trained to rewrite and analyze utterances 110.
[0104] The speech recognition engine 228 analyzes the filtered utterances and rewrites them. Some utterances are determined to be harmful by the third stage 115 and forwarded to the moderator 106. The moderator 106 can then provide feedback 132 regarding whether the forwarded harmful utterances were true positives or false positives. Furthermore, steps 342-346, similar to steps 330-334, upload random negative samples from the third stage using a random uploader. The moderator 106 can then provide further feedback 132 regarding whether the uploaded random utterances were true negatives or false negatives. Thus, the stage 115 is further trained using the positive and negative feedback from the moderator 106.
[0105] When analyzing filtered utterances, the third stage 115 can rewrite 20 seconds of utterances into text. Generally, machine learning-based rewriting is very expensive and time-consuming. Therefore, it is used in the third stage 115 of this system. The third stage 115 analyzes the 20 seconds of rewritten text and generates estimates of clip-separated harmfulness categories (e.g., sexual harassment, racial hate speech, etc.) with a predetermined confidence level.
[0106] The probability of a currently rewritten category is updated based on the previous clip, using clips that were previously rewritten and reached stage 3 (115) of the conversation. Therefore, the confidence level for a given toxicity category increases if previous instances of that category have been detected.
[0107] In various embodiments, the user context analyzer 226 may receive information as to whether any member of the conversation (e.g., speaker 102 and / or receiver 104) is presumed to be a child (e.g., determined by the second stage 115). If any member of the conversation is considered to be a child, the confidence level may increase and / or the threshold may decrease. Thus, in some embodiments, the third stage 115 is trained to be more likely to forward the utterance 110 to the moderator if a child is involved.
[0108] Next, the process proceeds to step 338, where the third stage 115 outputs a confidence interval for the utterance hazardity of the filtered utterances. It should be understood that the confidence output is training-dependent. For example, if a particular hazard policy is indifferent to general swear words and only pays attention to harassment, the training will take that into account. Therefore, stage 115 can be adapted to take into account the type of hazardity as needed.
[0109] Next, the process proceeds to step 340, which asks whether the third stage toxicity threshold is met. If yes, the process proceeds to step 348, which forwards the filtered utterance to the moderator. In various embodiments, the third stage 115 also outputs a transcript of the utterance 110 to the human moderator 106. If no, the utterance is filtered out in step 342 and then discarded in step 344. However, the random uploader 218 may pass on some of the filtered utterances to the human moderator, as previously described with reference to other stages 115.
[0110] In step 350, moderator 106 receives filtered harmful utterances through the multi-stage system 100. Therefore, moderator 106 should recognize a significant amount of filtered utterances. This helps resolve the issue of the moderator being manually called by the player / user.
[0111] If the moderator determines that the filtered utterance is harmful in accordance with the harmfulness policy, the process proceeds to step 352 to take corrective action. The moderator's "harmful" or "not harmful" rating 106 may also be forwarded to another system that determines what corrective action (if any) should be taken, including the possibility of taking no action for a first-time offender. Corrective action may include, among other options, warning the speaker 102, banning the speaker 102, muting the speaker 102, and / or altering the speaker's voice. The process then proceeds to step 354.
[0112] In step 354, the training data for various stages 115 is updated. Specifically, the training data for stage 1 115 is updated using positive and negative toxicity judgments from stage 2 115. The training data for stage 2 115 is updated using positive and negative toxicity judgments from stage 3 115. The training data for stage 3 115 is updated using positive and negative toxicity judgments from moderator 106. Thus, each subsequent stage 115 (or moderator) trains the preceding stage 115 on whether its judgment of a harmful utterance was accurate (as determined by the subsequent stage 115 or moderator 106).
[0113] In various embodiments, the preceding stage 115 is trained by the succeeding stage 115 to better detect false positives (i.e., non-harmful utterances that are considered harmful). This is because the preceding stage 115 takes over utterances that it considers harmful (i.e., those that meet the harmfulness threshold of a given stage 115). Furthermore, steps 322, 334, and 346 are used to train the succeeding stage to better detect false negatives (i.e., harmful utterances that are considered non-harmful). This is because a random sample of discarded utterances 111 is analyzed by the succeeding stage 115.
[0114] In addition, this training data makes the system 100 more robust as a whole and improves over time. Step 354 can be performed at various points in time. For example, step 354 may be performed adaptively in real time after each stage 115 has completed its analysis. Additionally or alternatively, the training data can also be batched at different time intervals (e.g., daily or weekly) and used to retrain the model on a cyclical schedule.
[0115] Next, the process proceeds to step 356, where it asks whether there are any more utterances 110 to analyze. If there are, the process returns to step 304, and process 300 is restarted. If there are no more utterances to analyze, the process can be terminated.
[0116] Therefore, the content moderation system is trained to reduce the rate of false negatives and false positives over time. For example, depending on the implementation or type of the system in the stage, it can be trained via gradient descent, Bayesian optimization, evolutionary methods, other optimization methods, or a combination of multiple optimization methods. If there are multiple distinct components in stage 115, they can be trained via different methods.
[0117] It should be noted that this process simplifies a longer process typically used to determine whether an utterance is harmful according to exemplary embodiments of the present invention. Therefore, the process for determining whether an utterance is harmful has many steps that a person skilled in the art is likely to use. Furthermore, some of the steps may be performed in a different order than those illustrated, or may be skipped entirely. Additionally or alternatively, some of the steps may be performed simultaneously. Therefore, a person skilled in the art can modify this process as appropriate.
[0118] While various embodiments refer to the “discarding” of utterances, it should be understood that this term does not necessarily mean that the utterance data is deleted or discarded. Instead, discarded utterances may be stored. Discarded utterances are simply intended to illustrate that the utterances are not transferred to subsequent stages 115 and / or moderator 106.
[0119] Figure 6 schematically shows details of System 100 that can be used with the processes of Figures 3A-3B according to an exemplary embodiment. Figure 6 is not intended to limit the use of the processes of Figures 3A-3B. For example, the processes of Figures 3A-3B can be used with various moderation content systems 100, including the system 100 shown in Figures 1A-1C.
[0120] In various embodiments, stage 115 may receive additional input (e.g., information about other speakers 102 in the session, such as the geographical location, IP address, or session context of speaker 102) and generate additional output that is stored in a database or input to a later stage 115 (e.g., player age estimation).
[0121] Throughout the operation of System 100, additional data is extracted and used by various stages 115 to support decision-making or provide additional context around clips. This data can be stored in a database and, combined with historical data, can provide a comprehensive understanding of a particular player. The additional data can also be aggregated across time periods, geographical regions, game modes, etc., to provide a high-level view of the state of in-game content (in this case, chat). For example, transcripts can be aggregated into a comprehensive picture of the frequency of use of various words and phrases, which can be charted to show how they change over time. Certain phrases whose frequency of use changes over time may attract the attention of platform administrators, who can use their deep contextual knowledge of the game to update the configuration of the multi-stage triage system to account for these changes (for example, when evaluating chat transcripts, giving a stronger rating to a keyword if its meaning changes from positive to negative). This can also be linked with other data; for example, even if the frequency of a word remains constant, if the sentiment of the phrase in which that word is used changes from positive to negative, that word can be highlighted. The aggregated data is displayed to platform administrators via a dashboard, which can show charts, statistics, and changes over time for various extracted data.
[0122] Figure 6 shows various segments of system 100 as separate entities (e.g., the first stage 115 and the random uploader 218), but this is not intended to limit the various embodiments. The random uploader 218 and other components of the system can be considered as part of various stages 115 or as separate from stage 115.
[0123] As generally explained in Figures 3A and 3B, the speaker 102 provides an utterance 110. The utterance 110 is received via the input unit 208, which decomposes the utterance 110 into chunks 110A and 110B that can be processed by the stage 115. In some embodiments, the utterance 110 may not be decomposed into chunks 110A and 110B, but may be received by the stage as is. The segmentator 234 may further decompose chunks 110A and 110B into analysis segments 240. However, in some embodiments, chunks 110A and 110B may be analyzed as a whole unit and can therefore be considered as an analysis segment 240. Furthermore, in some embodiments, the entire utterance 110 may be analyzed as a unit and can therefore be considered as an analysis segment 140.
[0124] The first stage 115 determines that a portion of the utterance 110 is potentially harmful and sends that portion of the utterance 110 (i.e., the filtered utterance 124) to the subsequent stage 115. However, a portion of the utterance 110 is deemed not harmful and is therefore discarded. As previously stated, to assist in the detection of false negatives (i.e., to detect utterances that are harmful but deemed not harmful), the uploader 218 uploads a certain percentage of the utterances to the subsequent stage 115 for analysis. If the subsequent stage 115 determines that the uploaded utterances were indeed false negatives, it may communicate directly with the first stage 115 (e.g., feedback 136A) and / or update the training database of the first stage (feedback 136B). The first stage 115 can be retrained actively adaptively or at pre-scheduled times. Thus, the first stage 115 is trained to reduce false negatives.
[0125] The filtered harmful utterance 124 is received and analyzed by a second stage 115, which determines whether the utterance 124 is likely to be harmful. The filtered harmful utterance 124 is found to be harmful by the first stage 115. The second stage 115 further analyzes the filtered harmful utterance 124. If the second stage 115 determines that the filtered utterance 124 is not harmful, the second stage 115 discards the utterance 124. However, the second stage 115 also provides feedback to the first stage 115 that the filtered utterance 124 was a false positive (either directly via feedback 136A or by updating the training database via feedback 136B). False positives can be included in the database 216 as false positives. Thus, the first stage 115 can be trained to reduce false positives.
[0126] Furthermore, Stage 2 115 passes utterances 124 that are considered likely to be harmful as harmful utterances 126. Utterances 124 that are considered less likely to be harmful become discarded utterances 111B. However, some of these discarded utterances 111B are uploaded by random upload 218 (to reduce false negatives in Stage 2 115).
[0127] The third stage 115 receives the further filtered harmful utterance 126 and analyzes it to determine whether it is likely to be harmful. The filtered harmful utterance 126 has been found to be harmful by the second stage 115. The third stage 115 further analyzes the filtered harmful utterance 126. If the third stage 115 determines that the filtered utterance 126 is not harmful, the third stage 115 discards the utterance 126. However, the third stage 115 also provides feedback to the second stage 115 that the filtered utterance 126 was a false positive (either directly via feedback 134A or by updating the training database via feedback 134B). False positives can be included in the training database 216 as false positives. Thus, the second stage 115 can be trained to reduce false positives.
[0128] Stage 3, 115, passes through utterances 126 that are considered likely to be harmful as harmful utterances 128. Utterances 126 that are considered less likely to be harmful become discarded utterances 111C. However, some of these discarded utterances 111C are uploaded by random upload 218 (to reduce false negatives in Stage 3, 115).
[0129] Moderator 106 receives the further filtered harmful utterance 128 and analyzes it to determine whether it is likely to be harmful. The filtered harmful utterance 128 has been found to be harmful by the third stage 115. Moderator 106 further analyzes the filtered harmful utterance 128. If moderator 106 determines that the filtered utterance 128 is not harmful, moderator 106 discards the utterance 128. However, moderator 106 also provides feedback to the third stage 115 that the filtered utterance 128 was a false positive (either directly via feedback 132A or by updating the training database via feedback 132B) (e.g., through the user interface). False positives can be included in the training database 216 as false positives. Thus, the third stage 115 can be trained to reduce false positives.
[0130] It is clear that various embodiments may have one or more stages 115 (e.g., 2 stages 115, 3 stages 115, 4 stages 115, 5 stages 115, etc.) distributed across multiple devices and / or cloud servers. Each stage may operate using a different machine learning approach. Preferably, earlier stages 115 use less computational power than later stages 115 in their utterance length-by-utterance analysis. However, by filtering out utterances 110 using a multi-stage process, fewer and fewer utterances are received at higher stages. Ultimately, the moderator receives a very small amount of utterance. Thus, the exemplary embodiments solve the problem of efficiently moderating audio content on large-scale platforms.
[0131] For example, let's assume that Stage 115 is low-cost (theoretically) enough to analyze 100,000 hours of audio for $10,000. Let's assume that Stage 215 is too expensive to process all 100,000 hours of audio, but can process 10,000 hours for $10,000. Let's assume that Stage 315 can perform even more computations and can analyze 1,000 hours for $10,000. Therefore, it is desirable to optimize the efficiency of the system so that potentially harmful utterances are progressively analyzed by more advanced (in this example, more expensive) stages, and non-harmful utterances are filtered out by more efficient and less advanced stages.
[0132] While various embodiments refer to speech modulation, it should be understood that similar processes can be used for other types of content such as images, text, and video. Generally, text does not have the same high-throughput problems as speech. However, video and images can also encounter similar throughput analysis problems.
[0133] The multi-stage triage system 100 can also be used for other purposes (for example, within the scope of gaming). For instance, the first two stages 115 may remain the same, but the output of the second stage 115 can be additionally sent to a separate system.
[0134] Furthermore, while various embodiments refer to the moderation of harmful content, it should be understood that the systems and methods described herein can be used to moderate any kind of utterance (or other content). For example, instead of monitoring harmful behavior, system 100 could monitor any specific content (e.g., product references or discussions about recent game changes "patches") to discover player sentiment on these topics. Similar to moderation systems, these stages can aggregate their findings, along with the extracted data, into a database and present them to administrators via a dashboard. Similarly, vocabulary and associated sentiments can be tracked and change over time. Stage 115 could output what is likely to be a product reference to a human moderation team to verify and determine the sentiment, or, if stage 115 is confident about the topic of the discussion and the associated sentiment, it could store its findings in the database to filter out content from subsequent stages, making the system more computationally efficient.
[0135] The same thing could happen with other enforcement topics, such as cheating or "selling gold coins" (selling in-game currency for real money). Similarly, there could be a Stage 115 that triages potential violations (for example, looking for mentions of popular cheating software, though the name may change over time), and a human moderation team that can make enforcement decisions on clips that have been passed over from Stage 115.
[0136] Therefore, using artificial intelligence or other known techniques, the exemplary embodiment allows later stages to improve upon the processing of earlier stages, bringing as much intelligence as possible closer to or effectively onto the user device. This reduces the need for later, slower stages (e.g., outside the device), enabling faster and more effective moderation.
[0137] Furthermore, while various embodiments mention that Stage 115 outputs a confidence interval, in some embodiments, Stage 115 may output its confidence in a different format (e.g., as yes or no, as a percentage, as a range, etc.). Additionally, instead of completely filtering out content from the moderation pipeline, a stage may prioritize content for a later stage or as output from the system, rather than explicitly rejecting any of them. For example, instead of rejecting content as unlikely to be abusive, a stage may assign abusiveness score to the content and insert it into a prioritized list of content for later stages to moderate. Later stages can then take the highest-scoring content from the list and filter it (or even later stages may prioritize it into a new list). Thus, by adjusting later stages to use a certain amount of computational power and prioritizing moderating the content that is most likely to be abusive, a certain amount of computation can be used efficiently.
[0138] Figure 7 schematically illustrates a four-stage system according to an exemplary embodiment of the present invention. The multi-stage adaptive triage system is computationally efficient, cost-effective, and scalable. Earlier stages of the system can be configured / designed to run more efficiently (e.g., faster) than later stages, keeping costs low by filtering out the majority of the content before less efficient, slower, but more powerful later stages are used. The earliest stages can even run locally on the user's device, further reducing platform costs. These initial stages adapt to filtering out contextually identifiable content by updating themselves with feedback from later stages. Later stages, having dramatically less content to look at overall, are given larger models and more computational resources, allowing them to be more accurate and improve upon the filtering done in earlier stages. By filtering out simpler content using computationally efficient initial stages, the system maintains high accuracy with efficient resource use, and the more powerful later-stage models are primarily employed for more complex moderation decisions that require those models. Furthermore, multiple options are available for different stages in the later stages of the system, and the previous stage or other monitoring system selects which next stage is appropriate based on content, extracted or historical data, or cost / accuracy trade-offs considering the stage options.
[0139] In addition to filtering out content that is unlikely to be offensive, the stages can also individually filter out easily identifiable offensive content and take autonomous action against it. For example, an initial stage that filters on a device might censor detected keywords indicating offensive behavior, and if no keywords are found, it can pass the task on to a later stage. Another example is an intermediate stage that might detect offensive words or phrases that were overlooked in previous stages, immediately warn the offending user, and discourage them from engaging in offensive behavior in the rest of their communication. Such determinations can also be reported to later stages.
[0140] The earlier stages in the system can perform other operations that assist in filtering later stages, thereby distributing some of the later stages' calculations to the earlier stages in the pipeline. This is particularly important when the earlier stages generate useful data or summaries of content that can also be used by later stages, thus avoiding repetitive calculations. The operations may be content summaries or semantically important compressions, which are passed to later stages instead of the content itself (which also reduces bandwidth between stages) or performed in addition to the content. The operations may also be extracting certain properties of the content that may be useful for purposes other than moderation tasks, and may be sent as metadata. The extracted properties may be stored by themselves or combined with past values to create more accurate averaged property values or a history of values over time, which can be used in filtering decisions in later stages.
[0141] The system may be configured to weight different elements of moderation, with more or less prioritization based on the preferences or needs of the platform employing moderation. The final stage of the system can output filtered content and / or extracted data to a team of human moderators or a configurable automated system, which can then return the decision results to itself, update the system, and make decisions that are more consistent with that team or system in the future. Individual stages can also be configured directly or indirectly updated based on feedback from external teams or systems, allowing the platform to control how the system uses various aspects of content to make moderation decisions. For example, in a voice chat moderation system, an intermediate stage could extract text from spoken content, compare that text to a (potentially weighted) word blacklist, and use the results to notify a moderation decision. Human teams can directly improve the speech-to-text engine used by providing manually annotated data, manually adapt the speech-to-text engine to new domains (new languages or accents), or manually adjust word blacklists (or potentially their severity weights) to prioritize more aggressive moderation of certain types of content.
[0142] The stages preferably update themselves based on feedback information from later stages, so that the entire system, or at least a part of it, can easily adapt to a new or changing environment. Updates may be performed online while the system is running, or they may be performed in batches later, such as in a bulk update or by waiting until the system has free resources to update. In the case of online updates, the system adapts to changing types of content by making initial filtering decisions on the content and then receiving feedback from a final team of human moderators or other external automated systems. The system can also inform manual configuration of the system by tracking the extracted properties of the content over time and indicating changes in those properties.
[0143] For example, in a chat moderation use case, the system might highlight changes in language distribution over time. For instance, if a new word (e.g., slang) suddenly becomes frequently used, this new word could be identified and displayed in a dashboard or summary, at which point the system administrator could configure the system to adapt to the changing language distribution. This also addresses cases where certain extracted properties change the impact on the moderation system's decisions. For example, when a chat moderation system is deployed, the word "sick" might have negative connotations, but over time, "sick" might take on positive connotations, and the context of its usage might change. The chat moderation system could highlight this change (e.g., report that "the word 'sick' was previously used in sentences with negative sentiment, but recently it has started to be used as a short, positive interjection"), provide administrators with a clear decision (e.g., "Is the word 'sick' offensive in this context?"), and allow itself to update to suit platform preferences.
[0144] An additional issue in content moderation relates to protecting the privacy of users whose content is being moderated, as content may contain identifying information. Moderation systems can use individual personally identified information (PII) filtering components to remove or censor ("scrub") PII from content before processing. In an exemplary multi-stage triage system, this PII scrubbing can be a pre-processing step before system execution, or it can be performed after several stages to assist in PII identification using extracted properties of the content.
[0145] This is achievable with pattern matching in text-based systems, but PII scrubbing is more difficult in video, images, and audio. One approach is to use a content identification system, such as a speech-to-text engine or optical character recognition engine, in combination with a text-based rule system to backtrack the locations of offensive words in audio, images, or videos, and then censor those areas of the content. This can also be done using a facial recognition engine to censor faces in images and videos to protect privacy during the moderation process. There is an additional technique of using style transfer systems to conceal the identity of subjects in content. For example, style transfer or “deepfake” systems for images or videos can anonymize faces present in content, allowing for effective moderation while preserving the rest of the content. In the audio domain, some embodiments may include an anonymizer, such as a speech skin or timbre transfer system, configured to convert speech into a new timbre, anonymizing the speaker's identifiable vocal characteristics without altering the content and emotion of the speech for the moderation process.
[0146] The multi-stage adaptive triage system is applicable to a variety of content moderation tasks. For example, it can moderate user postings of images, audio, video, text, or mixed media to social media sites (or parts of such sites, such as separate moderation criteria for a platform's "kids-friendly" section). It can also monitor user-to-user chats on platforms that allow audio, video, or text. For instance, in a multiplayer video game, it could monitor live audio chat between players, or on a video streaming site channel, it could manage text comments or chat. The system can also monitor more abstract properties, such as gameplay. For example, by tracking a player's past play style in a video game or the state of a particular game (e.g., score), the system can detect players who are behaving abnormally (e.g., intentionally losing or making mistakes to harass teammates) or various play styles that should be suppressed (e.g., "camping," where a player attacks before other players can react as soon as they spawn in the game, or when a player targets only other players in the game).
[0147] Beyond content moderation for inappropriate behavior, multi-stage adaptive triage systems in various embodiments can be used in other contexts to process large volumes of content. For example, the system can be used to monitor internal employee chats discussing confidential information. The system can also be used for behavioral analysis or tracking advertising sentiment by, for example, listening for mentions of a product or brand in voice or text chat and analyzing whether there are any associated positive or negative sentiments, or by monitoring players' reactions to new changes introduced by a game. The system can also be employed to detect illegal activities such as sharing illegal or copyrighted images, or activities prohibited by the platform, such as cheating or selling in-game currency for real money during gameplay.
[0148] As an example, consider the potential use of a multi-stage adaptive triage system in the context of moderating voice chat in a multiplayer game. The first stage of this system can be an audio segment detection system that filters out when a user is not speaking, and can operate on speech windows of several hundred milliseconds or one second at a time. The first stage can use an efficient parameterized model to determine whether a particular speaker is speaking or not, which can be adapted or calibrated based on additional information such as the game or region, and / or the user's voice settings or past volume levels. Furthermore, the various stages can classify what kind of harmful or noisy the user is making (e.g., blowing an air horn into voice chat). An exemplary embodiment can classify sounds (e.g., screaming, crying, air horning, groaning, etc.) to help classify harmfulness for moderator 106.
[0149] In addition to filtering out audio segments where the user is not speaking or making noise, the first stage can also identify properties of spoken content, such as normal volume level, current volume level, and background noise level, which can be used by itself or by subsequent stages to make filtering decisions (for example, louder speech is more likely to be disruptive). The first stage receives more useful and up-to-date information from the second stage and estimates its own performance by sending audio segments that are likely to contain speech, as well as a small number of segments that are unlikely to contain speech, to the second stage. The second stage returns information about the segments it determined are unlikely to be moderated, and the first stage updates itself to better mimic its reasoning in the future.
[0150] Stage 1 manipulates only short audio segments, while Stage 2 manipulates 15-second clips that may contain multiple sentences. Stage 2 can analyze voice tone and basic audio content, and can make more accurate judgments by utilizing past information about the player (for example, is a sudden change in tone of voice usually correlated with disruptive behavior?). Stage 2 can also make more informed judgments about spoken and non-spoken segments than Stage 1, taking into account its much larger temporal context, and can feed those judgments back to Stage 1 for optimization. However, Stage 2 requires significantly more computing power than Stage 1 to perform its filtering, so Stage 1 maintains the efficiency of Stage 2 by triaging silent segments. Both Stage 1 and Stage 2 can run locally on the user's device and do not require any direct computing costs from the game's centralized infrastructure.
[0151] As an extension of this example, the first stage can further detect sequences of phonemes in an utterance that are likely to be associated with swear words or other insults. The first stage can make autonomous decisions to censor swear words or other words / phrases, and during that time, it may mute or replace the speech with a tone. A more advanced first stage can replace the phonemes of the original utterance to generate a non-offensive word or phrase in either a standard voice or the player's own voice (with a voice skin or a special speech synthesis engine tailored to the vocal cords) (e.g., changing "f**k" to "fork").
[0152] After the second stage, any clips that were not filtered out are passed to the third stage, which runs on a cloud platform rather than locally on the device (although in some embodiments, more than three stages can run locally). The third stage has access to more context and more computing power. For example, a received 15-second audio clip can be analyzed in relation to the past two minutes of audio in the game, or it can be analyzed in addition to additional game data (e.g., "Is the player currently losing?"). In the third stage, a rough transcript can be created using an efficient speech-to-text engine, and the direct audio content of the utterance can be analyzed in addition to the tonal metadata sent from the second stage. If a clip is deemed potentially offensive, it is sent to the fourth stage, where additional information that may be part of a conversation can be incorporated, such as clips or transcripts from other players in the target player's party or game instance. Clips from conversations and other relevant clips may have transcripts from the third stage, refined by a more sophisticated but expensive speech recognition engine. Stage 4 may include game-specific vocabulary or phrases to aid in understanding the conversation, and sentiment analysis or other language comprehension may be performed to distinguish difficult cases (for example, is one player casually teasing another player who has been friends with them for a long time (e.g., they have played many games together)? Or are the two players exchanging angry insults, with the conversation becoming more intense over time?).
[0153] As another extension of this example, the third or fourth stage may detect abrupt changes in a player's sentiment, tone, or language that could indicate a serious change in the player's mental state. This can be automatically addressed with visual or auditory warnings to the player, an automatic change in their voice (e.g., to a high-pitched chipmunk voice), or muting the chat stream. In contrast, if such abrupt changes occur regularly with a particular player and indicate no connection to the game state or in-game behavior, it can be determined that the player has a periodic health problem, allowing for the avoidance of penalties while mitigating the impact on other players.
[0154] Stage 4 could include even more additional data, such as similar analysis of text chat (which may be performed by another multi-stage triage system), game status, and in-game images (e.g., screenshots).
[0155] Clips deemed potentially offensive in Stage 4 may be sent to the final human moderation team, along with context or other data. This team uses deep contextual knowledge of the game, along with metadata, properties, transcripts, and context surrounding the clip presented by the multi-stage triage system, to make the final moderation decision. This decision triggers a message to the game studio, which can then take action (e.g., warning or banning the players involved). The moderation decision information, along with potential additional data (e.g., "Why did the moderator make this decision?"), returns to Stage 4 and serves as training data to help update and improve Stage 4 itself.
[0156] Figure 8A schematically illustrates a process for training a machine learning model according to an exemplary embodiment of the present invention. Note that this process simplifies a longer process typically used to train stages of a system. Therefore, the process for training a machine learning model is likely to have many steps that a person skilled in the art would likely use. Furthermore, some of the steps may be performed in a different order than those illustrated, or may be skipped entirely. Additionally or alternatively, some of the steps may be performed simultaneously. Therefore, a person skilled in the art can modify this process as appropriate. In fact, it will be obvious to a person skilled in the art that the process described herein can be repeated for two or more stages (e.g., three or four stages).
[0157] Figure 8B schematically illustrates a system for training the machine learning model shown in Figure 8A, according to an exemplary embodiment of the present invention. Furthermore, the discussion of specific implementation examples of stage training with reference to Figure 8B is for illustrative purposes only and is not intended to limit the various embodiments. Those skilled in the art will understand that the stage training, and the various components and interactions of the stages, can be modified, removed, and / or added while developing a working toxicity moderation system 100 according to the exemplary embodiment.
[0158] Process 800 begins in step 802, preparing a multi-stage content analysis system such as system 100 in Figure 8B. In step 804, machine learning training is performed using a database 216 that has training data including positive and negative examples of training content. For example, in a hazard moderation system, positive examples may include audio clips that are hazardous, and negative examples may include audio clips that are not hazardous.
[0159] In step 806, the first stage analyzes the received content and generates a positive (S1-positive) or negative (S1-negative) judgment for the received speech content. Therefore, based on the first stage training received in step 804, it can be determined whether the received content is likely to be positive (e.g., contains harmful speech) or negative (e.g., does not contain harmful speech). The relevant S1-positive content is transferred to the subsequent stage. The relevant S1-negative content may be partially discarded and partially transferred to the subsequent stage (e.g., using the uploader described above).
[0160] In step 808, S1-positive content is analyzed using the second stage, which generates its own second-stage positive (S2-positive) and second-stage negative (S2-negative) judgments. Because the second stage is trained in a different way than the first stage, not all S1-positive content becomes S2-positive, and vice versa.
[0161] In step 810, S2-positive and S2-negative content are used to update the first stage training (e.g., in database 216). In an exemplary embodiment, the updated training provides a reduction in false positives from the first stage. In some embodiments, false negatives can also be reduced as a result of step 810. For example, if it is assumed that the classification of S2-positive and S2-negative is much easier to determine than existing training examples (if starting with some low-quality training examples) - this leads to the first stage 115 having a more learning-friendly time overall, and false negatives can be reduced as well.
[0162] In step 812, the transferred portion of the S1-negative content is analyzed using the second stage, which again generates both second-stage positive (S2-positive) and second-stage negative (S2-negative) judgments. In step 814, the S2-positive and S2-negative content are used to update the first-stage training (e.g., in database 216). In exemplary embodiments, the updated training provides a reduction in false negatives from the first stage. Similarly, in some embodiments, a reduction in false positives is also provided as a result of step 812.
[0163] Next, the process moves to step 816, where it is asked whether the training should be updated by discarding old training data. Periodically, by discarding old training data and retraining the first stage 115, the performance difference between the old and new data can be observed, and an increase in accuracy can be determined by removing old training data with low accuracy. Those skilled in the art will understand that in various embodiments, various stages can be retrained actively and adaptively or at pre-scheduled times. Furthermore, the training data in database 216 may be refreshed, updated, and / or discarded from time to time to allow for a shift in the input distribution of subsequent stages 115, taking into account that the output distribution of the previous stage 115 changes due to training. In some embodiments, changes in the previous stage 115 may have an undesirable effect on the type of input recognized by subsequent stages 115, potentially adversely impacting the training of subsequent stages 115. Therefore, in exemplary embodiments, some or all of the training data may be periodically updated and / or discarded.
[0164] If there are no training updates in step 816, the training process ends.
[0165] Various embodiments of the present invention may be implemented at least partially in any conventional computer programming language. For example, some embodiments may be implemented in a procedural programming language (e.g., "C") as a visual programming process, or in an object-oriented programming language (e.g., "C++"). Other embodiments of the present invention may be implemented as pre-configured standalone hardware elements and / or pre-programmed hardware elements (e.g., application-specific integrated circuits, FPGAs, and digital signal processors), or other related components.
[0166] In alternative embodiments, the disclosed apparatus and methods (e.g., as in any of the methods, flowcharts, or logical flows described above) may be implemented as computer program products used with a computer system. Such implementations may include a set of computer instructions fixed on any tangible, non-temporary, or non-temporary medium, such as a computer-readable medium (e.g., a diskette, CD-ROM, ROM, or fixed disk). The set of computer instructions may embody all or part of the functions previously described herein with respect to the system.
[0167] Those skilled in the art will understand that such computer instructions can be written in many programming languages for use with many computer architectures or operating systems. Furthermore, such instructions can be stored in any memory device, such as tangible, non-temporary semiconductor, magnetic, optical, or other memory devices, and can be transmitted over any suitable medium, such as wired (e.g., wires, coaxial cables, fiber optic cables, etc.) or wireless (e.g., through air or space), using any communication technology, such as optical, infrared, RF / microwave, or other transmission technologies.
[0168] In particular, such computer program products may be distributed as removable media accompanied by printed or electronic documentation (e.g., shrink-wrapped software), pre-loaded onto a computer system (e.g., on system ROM or a fixed disk), or distributed from a server or electronic bulletin board on a network (e.g., the Internet or the World Wide Web). In fact, some embodiments may be implemented as software as a service model ("SAAS") or as a cloud computing model. Of course, some embodiments of the present invention may be implemented as a combination of both software (e.g., computer program products) and hardware. Yet another embodiment of the present invention may be implemented entirely as hardware or entirely as software.
[0169] Computer program logic that implements all or part of the functions described herein may run on a single processor at different times (e.g., simultaneously), on multiple processors at the same time or at different times, under a single operating system process / thread, or under different operating system processes / threads. Therefore, the term “computer process” generally refers to the execution of a set of computer program instructions, regardless of whether different computer processes run on the same or different processors, and whether different computer processes run under the same or different operating system processes / threads. Software systems may be implemented using various architectures, such as monolithic architecture or microservices architecture.
[0170] Exemplary embodiments of the present invention may employ conventional components such as conventional computers (e.g., off-the-shelf PCs, mainframes, microprocessors), conventional programmable logic devices (e.g., off-the-shelf FPGAs or PLDs), or conventional hardware components (e.g., off-the-shelf ASICs or discrete hardware components), and when programmed or configured to perform the non-conventional methods described herein, non-conventional devices or systems are produced. Therefore, nothing is conventional in the inventions described herein. This is because, even when embodiments are implemented using conventional components, without special programming or configuration, the conventional components do not essentially perform the non-conventional functions described, and the resulting devices and systems are necessarily non-conventional.
[0171] While various inventive embodiments have been described and illustrated herein, those skilled in the art will readily conceive of various other means and / or structures to perform the functions described herein and / or to obtain the results and / or one or more advantages, and each of such variations and / or modifications will be considered to fall within the scope of the inventive embodiments described herein. More generally, those skilled in the art will readily understand that all parameters, dimensions, materials and configurations described herein are intended to be illustrative, and that actual parameters, dimensions, materials and / or configurations will depend on the specific use or application in which the teachings of the present invention are used. Those skilled in the art will recognize many equivalents to the specific inventive embodiments described herein, or can verify them by routine experimentation alone. Thus, the embodiments described herein are presented only as examples, and it will be understood that, within the scope of the appended claims and their equivalents, inventive embodiments can be implemented in ways other than those specifically described and claimed. The inventive embodiments of this disclosure are directed to each individual feature, system, article, material, kit and / or method described herein. Furthermore, any combination of two or more such features, systems, articles, materials, kits, and / or methods is included in the inventive scope of this disclosure, provided that they are not mutually inconsistent.
[0172] Various inventive concepts may be embodied in one or more methods, and examples of such methods are provided. The actions performed as part of a method can be ordered in any suitable manner. Thus, embodiments in which the actions are performed in a different order than those shown may be constructed, and this may include performing several actions simultaneously, even if they are shown as sequential actions in the exemplary embodiments.
[0173] While the above discussion discloses various exemplary embodiments of the present invention, it will be apparent to those skilled in the art that various modifications can be made to achieve some of the advantages of the present invention without departing from the true scope of the invention.
Claims
1. A hazard moderation system, wherein the system is An input unit configured to receive speech from the speaker, A multi-stage toxicity machine learning system comprising a first stage and a second stage, wherein the first stage is trained to analyze received utterances and determine whether the toxicity level of the utterances meets the toxicity threshold. Includes, The first stage is configured to filter and pass utterances that satisfy the harmfulness threshold to the second stage, and is further configured to filter and remove utterances that do not satisfy the harmfulness threshold. The first stage is computationally more efficient than the second stage. Hazardous substance moderation system.
2. The toxicity moderation system according to claim 1, wherein the first stage is trained using a database having training data including positive and / or negative examples of the training content of the first stage.
3. The first stage is trained using a feedback process, and the feedback process is Steps to receive spoken content, The steps include analyzing the speech content using the first stage and classifying the speech content as having positive speech content and / or negative speech content of the first stage, The steps include: analyzing the positive speech content of the first stage using the second stage, and classifying the positive speech content of the first stage as having positive speech content of the second stage and / or negative speech content of the second stage; The steps include updating the database using the positive speech content from the second stage and / or the negative speech content from the second stage, and A toxicity moderation system according to claim 2, comprising:
4. The toxicity moderation system according to claim 3, wherein the first stage discards at least a portion of the negative speech content of the first stage.
5. The first stage is trained using the feedback process, and the feedback process is The steps include: analyzing the content of the negative speech content in the first stage, but not all of it, using the second stage, and classifying the negative speech content in the first stage as having positive speech content and / or negative speech content in the second stage; The steps include further updating the database using the positive speech content from the second stage and / or the negative speech content from the second stage. The toxicity moderation system according to claim 3, further comprising:
6. The harmful moderation system according to claim 1, further comprising a random uploader configured to upload portions of speech that do not meet the harmfulness threshold to a subsequent stage or a human moderator.
7. The toxicity moderation system further includes a session context flagger configured to receive instructions that the speaker has previously met the toxicity threshold within a predetermined time, wherein the session context flagger is configured to (a) adjust the toxicity threshold, or (b) upload the portion of the utterance that did not meet the toxicity threshold to a subsequent stage or a human moderator, according to claim 1.
8. The toxicity moderation system according to claim 1, further comprising a user context analyzer, the user context analyzer configured to adjust the toxicity threshold and / or toxicity confidence level based on the speaker's age, the receiver's age, the speaker's geographical location, the speaker's friend list, the receiver's recent interaction history, the speaker's gameplay time, the length of the speaker's game, the start and end times of the game, and / or gameplay history.
9. The toxicity moderation system according to claim 1, further comprising an emotion analyzer trained to determine the emotions of the speaker.
10. The toxicity moderation system according to claim 1, further comprising an age analyzer trained to determine the age of the speaker.
11. The toxicity moderation system according to claim 1, further comprising a temporal receptive field configured to divide an utterance into time segments receivable by at least one stage.
12. The toxicity moderation system according to claim 1, further comprising a speech segmentator configured to divide an utterance into time segments analyzable by at least one stage.