Artificial wait time for moderating voice communications

A machine learning model identifies harmful content in audio streams and introduces time delays to moderate online communications, ensuring a safe and respectful environment without disrupting conversations.

JP2025532007AActive Publication Date: 2025-09-29ROBLOX CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025514438
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-09-08
Filing Date
2023-09-06
Publication Date
2025-09-29
Estimated Expiration
2043-09-06

AI Technical Summary

Technical Problem

Moderating audio streams in online platforms is challenging due to variations in accent, tone, volume, and sarcasm, making it difficult to analyze and introduce delays without disrupting conversations.

Method used

Introduce artificial latency into audio streams by using a trained machine learning model to identify harmful content and insert time delays or gaps, while synchronizing visual signals, to moderate audio and video communications effectively.

Benefits of technology

Effectively moderates harmful content in real-time without perceptible delays, maintaining conversation flow and providing a safe communication environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025532007000001_ABST
    Figure 2025532007000001_ABST
Patent Text Reader

Abstract

A computer-implemented method for determining whether to introduce latency into an audio stream from a particular speaker includes receiving an audio stream from a sending device. The method further includes providing the audio stream and speech analysis scores, information about one or more vocal emotion parameters, and one or more vocal emotion scores for a first user associated with the sending device as inputs to a trained machine learning model, where the trained machine learning model is iteratively applied to the audio stream, each iteration corresponding to a respective portion of the audio stream. The method further includes using the trained machine learning model to generate as output a level of harmfulness in the audio stream. The method further includes transmitting the audio stream to a receiving device, where the transmitting step is performed to introduce a time delay into the audio stream based on the level of harmfulness.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is an international application and claims the benefit of priority under 35 U.S.C. § 119(e) to U.S. patent application Ser. No. 17 / 940,749, filed Sep. 8, 2022, entitled ARTIFICIAL LATENCY FOR MODERATING VOICE COMMUNICATION, the entire contents of which are incorporated herein by reference. [Background technology]

[0002] Online platforms need ways to provide a safe and respectful environment for communication between user devices. Text communication is easier to moderate than audio communication because users are more tolerant of delays when sending text messages. In addition, moderating text communication is easier than audio streams because the text can be compared to a list of banned or problematic words. Conversely, moderating audio streams is more difficult to analyze due to variations in accent, tone, volume, use of sarcasm, etc.

[0003] The background discussion provided herein is for the purpose of presenting a context for the present disclosure. To the extent described in this background section, the work of the presently named inventors, as well as aspects of the body of the specification that may not qualify as prior art at the time of filing, are not admitted expressly or impliedly as prior art to the present disclosure. Summary of the Invention [Means for solving the problem]

[0004] Embodiments generally relate to systems and methods for introducing artificial latency into an audio stream for moderation. According to one aspect, a computer-implemented method includes receiving an audio stream from a sending device. The method further includes providing the audio stream and a speech analysis score, information about one or more vocal emotion parameters, and one or more vocal emotion scores for a first user associated with the sending device as inputs to a trained machine learning model, the trained machine learning model being iteratively applied to the audio stream, each iteration corresponding to a respective portion of the audio stream. The method further includes using the trained machine learning model to generate as output a level of harmfulness in the audio stream. The method further includes transmitting the audio stream to a receiving device, the transmitting step being performed to introduce a time delay into the audio stream based on the level of harmfulness.

[0005] In some embodiments, the method further includes identifying instances of harmfulness in the audio stream and replacing the instances of harmfulness in the audio stream with noise or silence before transmitting the audio stream to the receiving device. In some embodiments, the method further includes identifying inter-word silences or pauses in the audio stream, where the silences or pauses correspond to specific timestamps in the audio stream, and a time delay is introduced as a gap in the audio stream at the specific timestamp of the inter-word silences or pauses. In some embodiments, the method further includes updating a speech analysis score based on identifying the instances of harmfulness in the audio stream. In some embodiments, the method further includes receiving text from a text channel associated with the sending device, the text channel being separate from the audio stream, and generating a text score indicative of a harmfulness rating for the text, where the input to the trained machine learning model further includes the text score. In some embodiments, the input to the trained machine learning model further includes a harm history of a first user, a speaker history and metadata associated with the first user, and a listener history and metadata associated with a second user associated with the receiving device. In some embodiments, the one or more vocal emotion parameters include tone, pitch, and vocal effort level determined based on one or more previous audio streams from the sending device. In some embodiments, the audio stream is provided with a visual signal, and the method further includes synchronizing the visual signal with the audio stream by introducing a time delay in the visual signal that is the same as the time delay of the audio stream.In some embodiments, the audio stream is part of the video stream, and the method further includes analyzing the audio stream to identify instances of harm; detecting, in response to identifying the instances of harm, portions of the video stream depicting an objectionable gesture, where the objectionable gesture occurs within a predetermined time period of the instance of harm; and modifying, in response to detecting the objectionable gesture, at least the portions of the video stream by one or more of blurring the portions or replacing the portions with pixels that correspond to a background region. In some embodiments, the audio stream is part of the video stream, and the method further includes performing motion detection on the video stream to detect objectionable gestures; and modifying, in response to detecting the objectionable action, at least the portions of the video stream by one or more of blurring the portions or replacing the portions with pixels that correspond to a background region. In some embodiments, if the level of harm is below a minimum threshold, the time delay is zero seconds.

[0006] According to one aspect, a device includes a processor and a memory coupled to the processor having instructions that, when executed by the processor, cause the processor to perform operations including receiving an audio stream from a sending device; providing the audio stream and a voice analysis score, information regarding one or more vocal emotion parameters, and one or more vocal emotion scores for a first user associated with the sending device as inputs to a trained machine learning model, wherein the trained machine learning model is iteratively applied to the audio stream, with each iteration corresponding to a respective portion of the audio stream; generating, using the trained machine learning model, a level of harmfulness in the audio stream as an output; and transmitting the audio stream to a receiving device, wherein the transmitting is performed to introduce a time delay in the audio stream based on the level of harmfulness.

[0007] In some embodiments, the operations further include identifying instances of harmfulness in the audio stream and replacing the instances of harmfulness in the audio stream with noise or silence before transmitting the audio stream to a receiving device. In some embodiments, the operations further include identifying inter-word silences or pauses in the audio stream, the silences or pauses corresponding to particular timestamps in the audio stream, and a time delay is introduced as a gap in the audio stream at the particular timestamp of the inter-word silences or pauses. In some embodiments, the operations further include updating a speech analysis score based on identifying instances of harmfulness in the audio stream.

[0008] According to one aspect, a non-transitory computer-readable medium having stored thereon instructions that, when executed by one or more computers, cause the one or more computers to perform operations including receiving an audio stream from a sending device; providing the audio stream and a voice analysis score, information regarding one or more vocal emotion parameters, and one or more vocal emotion scores for a first user associated with the sending device as inputs to a trained machine learning model, wherein the trained machine learning model is iteratively applied to the audio stream, each iteration corresponding to a respective portion of the audio stream; generating, using the trained machine learning model, as output, a level of harmfulness in the audio stream; and transmitting the audio stream to a receiving device, wherein the transmitting is performed to introduce a time delay in the audio stream based on the level of harmfulness.

[0009] In some embodiments, the operations further include identifying instances of harm in the audio stream and replacing the instances of harm in the audio stream with noise or silence before transmitting the audio stream to the receiving device. In some embodiments, the operations further include identifying inter-word silences or pauses in the audio stream, where the silences or pauses correspond to particular timestamps in the audio stream, and a time delay is introduced as a gap in the audio stream at the particular timestamp of the inter-word silences or pauses. In some embodiments, the operations further include updating a speech analysis score based on identifying the instances of harm in the audio stream. In some embodiments, the operations further include receiving text from a text channel associated with the sending device, where the text channel is separate from the audio stream, and generating a text score indicative of a harm rating for the text, where the input to the trained machine learning model further includes the text score.

[0010] One method for preventing harm in an audio stream is to buffer the audio stream, identify instances of harm, and remove them from the audio stream before it is transmitted from the sending device to the receiving device. However, moderating the audio stream introduces delays of several seconds into the audio stream. Audio delays of more than 50 milliseconds introduce unnatural delays that disrupt conversation, and audio delays of more than 250 milliseconds disrupt conversation.

[0011] The present application advantageously describes a metaverse engine and / or metaverse application that provides a method for identifying instances of harmfulness while selectively inserting gaps or pauses in an audio stream to perform moderation without perceptible delay. [Brief explanation of the drawings]

[0012] [Figure 1] FIG. 1 is a block diagram of an example network environment for identifying instances of harmfulness in communications, according to some embodiments described herein. [Figure 2] FIG. 1 is a block diagram of an example computing device for identifying instances of harmfulness in communications, according to some embodiments described herein. [Figure 3] 1 is a diagram of an example user interface of a video stream in which instances of harmfulness are identified, according to certain embodiments described herein. [Figure 4] 1 is a diagram of an example user interface of a video stream in which objectionable actions are identified, according to some embodiments described herein. [Figure 5] FIG. 1 is an example flow diagram for identifying instances of harmfulness in communications, according to some embodiments described herein. [Figure 6] FIG. 10 is another example flow diagram for identifying instances of harmfulness in communications, according to certain embodiments described herein. DETAILED DESCRIPTION OF THE INVENTION

[0013] Network Environment 100 FIG. 1 shows a block diagram of an exemplary network environment 100 for identifying instances of harmfulness in communications. In some embodiments, environment 100 includes server 101, user devices 115a...n, and network 105. Users 125a...n may be associated with respective user devices 115a...n. In FIG. 1 and the remaining figures, a letter following a reference number, e.g., "115a," represents a reference to the element with that specific reference number. A reference number in text without a following letter, e.g., "115," represents a general reference to the element with that reference number. In some embodiments, environment 100 may include other servers or devices not shown in FIG. 1. For example, server 101 may be multiple servers 101.

[0014] Server 101 includes one or more servers, each including a processor, memory, and network communication hardware. In some embodiments, server 101 is a hardware server. Server 101 is communicatively coupled to network 105. In some embodiments, server 101 transmits data to and receives data from user devices 115. Server 101 may include metaverse engine 103 and database 199.

[0015] In some embodiments, the metaverse engine 103 includes code and routines operable to receive communications between two or more users within a virtual metaverse, such as communications at the same location within the metaverse, communications within the same metaverse experience, or communications between friends within a metaverse application. Users interact within the metaverse across a variety of attributes (e.g., different ages, regions, languages, etc.).

[0016] In some embodiments, the metaverse engine 103 receives an audio stream intended for user device 115a through user device 115n. The metaverse engine 103 provides the audio stream and speech analysis scores, information about one or more vocal emotion parameters, and one or more vocal emotion scores for user 125a associated with device 115a as inputs to a trained machine learning model. The trained machine learning model is iteratively applied to each portion of the audio stream, such as a few seconds of the received audio stream.

[0017] The metaverse engine 103 uses the trained machine learning model to generate as output a level of harmfulness in the audio stream. The metaverse engine 103 transmits the audio stream to one or more other user devices 115n, the transmitting being performed to introduce a time delay in the audio stream based on the level of harmfulness. In some embodiments, the metaverse engine 103 uses the time delay to identify instances of harmfulness and replaces the instances of harmfulness with noise or silence before transmitting the audio stream to one or more other user devices 115n.

[0018] In some embodiments, the metaverse engine 103 is implemented using hardware including a central processing unit (CPU), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), any other type of processor, or a combination thereof. In some embodiments, the metaverse engine 103 is implemented using a combination of hardware and software.

[0019] The database 199 may be a non-transitory computer-readable memory (e.g., random access memory), a cache, a drive (e.g., a hard drive), a flash drive, a database system, or another type of component or device capable of storing data. The database 199 may also include multiple storage components (e.g., multiple drives or multiple databases) that may span multiple computing devices (e.g., multiple server computers). The database 199 may store data associated with the metaverse engine 103, such as training datasets for trained machine learning models, history and metadata associated with each user 125.

[0020] The user device 115 may be a computing device that includes a memory, a hardware processor, and a camera. For example, the user device 115 may include a mobile device, a tablet computer, a mobile phone, a wearable device, a head-mounted display, a mobile email device, a portable game player, a portable music player, a reader device, or another electronic device that can access the network 105 and capture images using a camera.

[0021] User device 115a includes metaverse application 104a, and user device 115n includes metaverse application 104b. In some embodiments, user device 115a is the sending device, and user device 115n is the receiving device. In some embodiments, user 125a uses metaverse application 104a on the sending device to generate a communication, such as an audio or video stream, which is sent to metaverse engine 103. Once the communication is approved for transmission, metaverse engine 103 sends the communication to metaverse application 104b on the receiving device for access by user 125n.

[0022] In the illustrated embodiment, the entities of environment 100 are communicatively coupled via network 105. Network 105 may include a public network (e.g., the Internet), a private network (e.g., a local area network (LAN) or wide area network (WAN)), a wired network (e.g., an Ethernet network), a wireless network (e.g., an 802.11 network, Wi-Fi, or wireless LAN (WLAN)), a cellular network (e.g., a long-term evolution (LTE) network), a router, a hub, a switch, a server computer, or a combination thereof. While FIG. 1 shows one network 105 coupled to server 101 and user device 115, in practice, one or more networks 105 may be coupled to these entities.

[0023] 200 Examples of Computing Devices 2 is a block diagram of an example computing device 200 that may be used to implement one or more features described herein. Computing device 200 may be any suitable computer system, server, or other electronic or hardware device. In some embodiments, computing device 200 is server 101. In some embodiments, computing device 200 is user device 115.

[0024] In some embodiments, computing device 200 includes a processor 235, memory 237, an input / output (I / O) interface 239, a microphone 241, a speaker 243, a display 245, and a storage device 247. Depending on whether computing device 200 is a server 101 or a user device 115, some components of computing device 200 may not be present. For example, if computing device 200 is a server 101, the computing device may not include microphone 241 and speaker 243. In some embodiments, computing device 200 includes additional components not shown in FIG. 2 .

[0025] Processor 235 may be coupled to bus 218 via signal line 222, memory 237 may be coupled to bus 218 via signal line 224, I / O interface 239 may be coupled to bus 218 via signal line 226, microphone 241 may be coupled to bus 218 via signal line 228, speaker 243 may be coupled to bus 218 via signal line 230, display 245 may be coupled to bus 218 via signal line 232, and storage device 247 may be coupled to bus 218 via signal line 234.

[0026] Processor 235 may include an arithmetic logic unit, microprocessor, general-purpose controller, or some other processor array to perform calculations and provide instructions to a display device. Processor 235 processes data and may include various computing architectures, including complex instruction set computer (CISC) architecture, reduced instruction set computer (RISC) architecture, or architectures implementing a combination of instruction sets. While FIG. 2 shows a single processor 235, multiple processors 235 may be included. In different embodiments, processor 235 may be a single-core processor or a multi-core processor. Other processors (e.g., graphics processing units), operating systems, sensors, displays, and / or physical configurations may be part of computing device 200.

[0027] Memory 237 stores instructions and / or data that may be executed by processor 235. The instructions may include code and / or routines for performing the techniques described herein. Memory 237 may be a dynamic random access memory (DRAM) device, static RAM, or some other memory device. In some embodiments, memory 237 also includes non-volatile memory such as static random access memory (SRAM) or flash memory, or similar persistent storage devices and media, including hard disk drives, compact disc read-only memory (CD-ROM) devices, DVD-ROM devices, DVD-RAM devices, DVD-RW devices, flash memory devices, or some other mass storage device for more permanently storing information. Memory 237 includes code and routines operable to execute metaverse engine 103, described in more detail below.

[0028] I / O interface 239 may provide functionality that allows computing device 200 to interface with other systems and devices. The interfaced devices may be included as part of computing device 200 or may be separate and communicate with computing device 200. For example, network communication devices, storage devices (e.g., memory 237 and / or storage device 247), and input / output devices may communicate through I / O interface 239. In another example, I / O interface 239 may receive data from server 101 and provide the data to components of metaverse engine 103, such as metaverse engine 103 or machine learning module 210. In some embodiments, I / O interface 239 may connect to interface devices such as input devices (keyboard, pointing device, touchscreen, microphone 241, sensor, etc.) and / or output devices (display device, speaker 243, monitor, etc.).

[0029] Some examples of interfaced devices that may be connected to I / O interface 239 include display 245, which may be used to display content, e.g., images, video, and / or user interfaces of output applications described herein, and to receive touch (or gesture) input from a user. Display 245 may include any suitable display device, such as a liquid crystal display (LCD), a light emitting diode (LED) or plasma display screen, a cathode ray tube (CRT), a television, a monitor, a touchscreen, a three-dimensional display screen, or other visual display device.

[0030] Microphone 241 includes hardware for detecting speech spoken by a person and may transmit the speech to metaverse engine 103 via I / O interface 239.

[0031] Speaker 243 includes hardware for generating audio for playback. For example, speaker 243 receives instructions from metaverse engine 103 to generate audio from another user after it has been determined that the audio stream does not contain instances of harm. Speaker 243 converts the instructions into audio and generates the audio for the user.

[0032] Storage device 247 stores data related to metaverse engine 103. For example, storage device 247 may store training datasets for trained machine learning models, history and metadata associated with each user 125, etc. In embodiments in which computing device 200 is server 101, storage device 247 is the same as database 199 in FIG. 1.

[0033] Exemplary Metaverse Engine 103 or Metaverse Application 104 2 illustrates a computing device 200 executing an example metaverse engine 103 or metaverse application 104 that includes a history module 202, a speech analyzer 204, a vocal sentiment analyzer 206, a text module 208, a machine learning module 210, a harmfulness module 212, and a user interface module 214. While the modules are shown as being part of the same metaverse engine 103 or metaverse application 104, those skilled in the art will recognize that the modules may be implemented by the computing device 200. For example, the text module 208 may be part of the user device 115 that provides analysis of text communications before the text communications are sent to the metaverse engine 103, which is part of the server 101, in order to reduce the computational requirements of the server 101.

[0034] The history module 202 generates historical information and metadata about users participating in communications. In some embodiments, the history module 202 includes a set of instructions executable by the processor 235 to generate the historical information and metadata. In some embodiments, the history module 202 may be stored in memory 237 of the computing device 200 and accessible and executable by the processor 235.

[0035] In some embodiments, after obtaining a user's permission, the history module 202 stores information about each communication session in the metaverse associated with the user and metadata associated with the user. The communication session may include an audio stream, a video stream, a text communication, etc. After obtaining a user's permission, the history module 202 may store information about instances of toxicity associated with the user. For example, the history module 202 may identify when the user participated in an instance of toxicity, what harmful behavior the user performed (e.g., spoke foul language, performed offensive actions, bullied another user, threatened another user, etc.), the specific users targeted by the instance of toxicity, etc. In all cases where information about a user is stored, the history module 202 has obtained permission from the user, the user has been made aware that the user can delete the information, and the information is stored securely and in compliance with applicable regulations. Further details are discussed below with reference to the user interface module 214.

[0036] In some embodiments, after obtaining user permission, the history module 202 may store information about the context of an instance of harm, such as a particular experience. For example, a user may use a lot of swear words while playing a violent shooter game, but may not exhibit harmful behavior in a non-violent role-playing game. The history module 202 may receive information about instances of harm from other modules, such as the speech analyzer 204 and the harm module 212.

[0037] In some embodiments, after obtaining a user's permission, the history module 202 stores the listening history and metadata regarding the user's reactions to instances of harm directed at the user by other users. For example, the history module 202 may update the listening history and metadata to indicate whether the user is insensitive to slurs, whether the user responds to abusive language with an abusive language, whether the user reports users in response to seeing offensive behavior, etc. In another example, the history module 202 may track how many times a particular user has been blocked and whether the particular user participates in an event where another user who blocked the particular user is also present. In some embodiments, the history module 202 generates a sensitivity score that reflects the user's sensitivity to instances of harm on a scale (e.g., 3 out of 10, 0.9 out of 1, etc.).

[0038] In some embodiments, after obtaining the user's permission, the history module 202 stores metadata associated with the user. For example, the metadata may include the region in which the user lives, other demographic information (such as sex, gender, age, race, preferred pronouns, orientation, etc.), one or more Internet Protocol (IP) addresses associated with the user, one or more languages ​​spoken by the user, etc. In some embodiments, the history module 202 may characterize the user's reactions in combination with the metadata. For example, the history module 202 may identify that the user is generally not sensitive to instances of harm in the metaverse unless the user is called a slur that matches the user's religious affiliation, gender, race, etc.

[0039] The history module 202 may provide historical information and metadata to the machine learning module 210 as input to a trained machine learning model. The history module 202 may also provide historical information and metadata to other modules to provide context that influences how instances of harm are calculated. For example, the speech analyzer 204 receives historical information and metadata because it uses different rules to identify instances of harm among users aged 13-16, 16-18, or 19+.

[0040] In some embodiments, the speech analyzer 204 analyzes speech during a communication session. In some embodiments, the speech analyzer 204 includes a set of instructions executable by the processor 235 to analyze speech during a communication session. In some embodiments, the speech analyzer 204 may be stored in the memory 237 of the computing device 200 and accessible and executable by the processor 235.

[0041] The voice analyzer 204 receives the audio stream from the sending device. Because the voice analysis may take several seconds, the voice analyzer 204 performs a continuous analysis of the voice in the audio stream. The analysis may be retroactive, in that the voice analyzer 204 performs the analysis after the audio stream has already been sent to the receiving device, or the analysis may occur each time an instance of harmfulness is identified, regardless of whether the audio stream has been sent to the receiving device.

[0042] In some embodiments, the speech analyzer 204 includes machine learning models trained to predict various attributes of an audio stream, such as vocal effort, speaking style, language, and speech activity.

[0043] The speech analyzer 204 may perform automatic speech recognition (ASR), such as speech-to-text translation, and compare the translated text to a list of harmful words to identify instances of harm in the audio stream. The speech analyzer 204 may generate a speech analysis score for a user associated with the sending device. For example, the speech analyzer 204 may generate a speech analysis score based on identifying instances of harm associated with the audio stream.

[0044] In some embodiments, the voice analyzer 204 generates a voice analysis score based on demographics related to a particular user. For example, the voice analyzer 204 applies different rubrics for what constitutes an instance of harm based on whether the user is 13-16 years old, 16-18 years old, 19+ years old (or 12-15 years old, 15-18 years old, 19+ years old, etc.), based on the user's location, based on whether the audio stream is being sent to users with different demographics (e.g., an audio stream may be identified as containing instances of harm if sent to a 13-year-old user), or based on the type of game (e.g., shooter game vs. puzzle game).

[0045] In some embodiments, the speech analyzer 204 performs speech-to-text translation based on one or more languages ​​spoken by the user. For example, the aquatic mammal "seal" in English is called "phoque" in French, which should not be confused with an instance of harm, i.e., the obscene word "fuck" in English. In some embodiments, an identification of the one or more languages ​​spoken by the user is received as part of the metadata determined by the history module 202.

[0046] In some embodiments, the voice analyzer 204 periodically provides voice analyses, such as voice analysis scores, to the machine learning module 210 as input to the trained machine learning model. The voice analysis scores may be associated with timestamps such that the voice analysis scores are aligned with positions within the audio stream. In some embodiments, the voice analyzer 204 sends the voice analysis scores to the machine learning module 210 each time the voice analyzer 204 identifies an instance of harm in the audio stream and updates the voice analysis scores to reflect the identified instance of harm.

[0047] The vocal emotion analyzer 206 analyzes emotions in the audio stream. In some embodiments, the vocal emotion analyzer 206 includes a set of instructions executable by the processor 235 to analyze emotions in the audio stream. In some embodiments, the vocal emotion analyzer 206 may be stored in the memory 237 of the computing device 200 and accessible and executable by the processor 235.

[0048] In some embodiments, the vocal emotion analyzer 206 identifies different speakers in the audio stream and associates the different speakers with corresponding audio identifiers. The vocal emotion analyzer 206 analyzes multiple vocal parameters, such as one or more of tone, pitch, and vocal effort level, of each user in the audio stream. In some embodiments, the vocal emotion analyzer 206 analyzes tone by determining the positivity and energy of the user's voice. For example, the vocal emotion analyzer 206 detects whether the user appears excited, irritated, neutral, sad, etc. In some embodiments, the vocal emotion analyzer 206 determines the speaker's emotional state based on an emotion quadrant. The emotion quadrant includes four states: nervous, happy, angry, and sad. The vocal emotion analyzer 206 may use a transformer-based technique for emotion detection, such as using wav2vec2.0 as part of its front end.

[0049] In some embodiments, the vocal emotion analyzer 206 analyzes pitches within a range of 60 Hertz (Hz) to 2 kHz. In some embodiments, the vocal emotion analyzer 206 establishes the fundamental frequency of a sound and determines the range of pitches occurring in the audio stream.

[0050] In some embodiments, the vocal emotion analyzer 206 analyzes the vocal effort level by determining the level of noise and comparing that level to a predetermined description of vocal effort. For example, the vocal emotion analyzer 206 may determine that a whisperer contributes a vocal effort of 20-30 decibels (dB), a soft-speaker contributes a vocal effort of 30-55 dB, a person speaking at an average level contributes a vocal effort of 55-65 dB, a loud-speaker or shouter contributes a vocal effort of 65-80 dB, and a screamer contributes a vocal effort of 80-120 dB. In some embodiments, the vocal emotion analyzer 206 also identifies whether vocal effort is increasing as a function of time, which may indicate a conversation escalating into an argument where harassment may occur.

[0051] In some embodiments, the vocal emotion analyzer 206 generates a vocal emotion score for the user associated with the audio stream. In some embodiments, the vocal emotion analyzer 206 generates a separate score for each of tone, pitch, and vocal effort level, which are determined based on one or more previous audio streams from the sending device. In some embodiments, the vocal emotion analyzer 206 generates a vocal emotion score that is a combination of tone, pitch, and vocal effort. For example, because users do not shout when they are angry, it may not be clear that a user is angry unless there is a combination of an angry tone, a large variation in pitch, and a low vocal effort level.

[0052] The vocal emotion analyzer 206 may analyze emotion in the audio stream regardless of whether the user used vocal modulation software. For example, if the user selected vocal modulation software that sounds like a popular cartoon character, the vocal emotion analyzer 206 would detect that modulation was occurring and perform emotion analysis regardless of the modulation.

[0053] The vocal emotion analyzer 206 may periodically generate one or more vocal emotion scores and send information about the vocal emotion parameters and the one or more vocal emotion scores to the machine learning module 210. For example, the vocal emotion analyzer 206 sends information about the vocal emotion parameters and the one or more vocal emotion scores to the machine learning module 210 each time the speech analyzer 204 identifies an instance of harmfulness in the audio stream and updates the speech analysis scores to reflect the identified instance of harmfulness. In another example, the vocal emotion analyzer 206 sends information about the vocal emotion parameters and the one or more vocal emotion scores each time there is a change in the vocal emotion parameters, such as when a change in tone, pitch, or vocal effort level occurs, or when a change in tone, pitch, or vocal effort level exceeds one or more predetermined thresholds, such as when a user transitions from a vocal effort level associated with normal speaking to a vocal effort level associated with shouting.

[0054] In some embodiments, the vocal emotion analyzer 206 is stored on the sending device and may perform a quick analysis of the vocal emotion. For example, the sending device may include a vocal emotion analyzer 206 with a 60% accuracy in detecting tone of voice. The vocal emotion analyzer 206 on the sending device may send information to a vocal emotion analyzer 206 on the server 101 that performs a more detailed analysis of the vocal emotion.

[0055] The text module 208 analyzes text from the text channel. In some embodiments, the text module 208 includes a set of instructions executable by the processor 235 to analyze the text. In some embodiments, the text module 208 may be stored in the memory 237 of the computing device 200 and accessible and executable by the processor 235.

[0056] In some embodiments, a user may participate in an audio stream and use a separate text channel to send text messages. For example, during a video call, a user may give a presentation and also add additional information to a chat box associated with the video call. In another example, a user may participate in a game using an audio stream while sending text messages directly to specific users through the game software. Because some users may be polite on the audio stream but directly abuse specific users using direct messages during the game, analyzing both the audio stream and the text messages can be useful.

[0057] In some embodiments, the text module 208 compares the text to a list of harmful words to identify instances of harm in the text. The text module 208 may generate a text score indicating a harm rating for the text based on the text message. The text module 208 may periodically generate the text score and send the text score to the machine learning module 210. In some embodiments, the text module 208 sends the text score each time an instance of harm is identified in a text message, and the text module 208 updates the text score accordingly.

[0058] In some embodiments, the text module 208 removes instances of harm from the text before sending the text to the receiving device. The text module 208 may simply remove instances of harm, such as a first user threatening another user. Or, the text module 208 may include a warning and an explanation as to why the instance of harm was removed from the text.

[0059] In some embodiments, the text module 208 is stored on the sending device as part of the metaverse application 104a, and the text is analyzed on the sending device to conserve computational resources on the server 101. The text module 208 on the sending device may send the text score to the metaverse engine 103 as input to the machine learning module 210.

[0060] The machine learning module 210 trains a machine learning model (or models) to output a level of harmfulness in the audio stream. In some embodiments, the machine learning module 210 includes a set of instructions executable by the processor 235 to train the machine learning model to output a level of harmfulness in the audio stream. In some embodiments, the machine learning module 210 may be stored in the memory 237 of the computing device 200 and accessible and executable by the processor 235.

[0061] In some embodiments, the machine learning module 210 obtains a training dataset having manually labeled audio streams paired with output from one or more of the history module 202, the speech analyzer 204, the vocal sentiment analyzer 206, and the text module 208. In some embodiments, the manual labels include instances of harmfulness in the audio streams and meta-information such as language, emotional state, vocal effort, etc. For example, the training dataset may include an audio stream paired with an output from the history module 202 including one or more of: a first user's (i.e., speaker's) harmfulness history, speaker history and metadata associated with the first user, and listener history and metadata associated with a second user; an output from the speech analyzer 204 including periodically transmitted speech analysis scores for the first user; an output from the vocal sentiment analyzer 206 including periodically transmitted speech analysis scores for the first user, and in some embodiments, an output from the text module 208 including individual speech analysis scores for tone, pitch, and vocal effort level; and an output from the text module 208 including text scores.

[0062] In some embodiments, the training dataset further includes automatically labeled audio streams preprocessed for offline harm detection, where the labels may indicate timestamps within the audio stream at which instances of harm occurred.

[0063] In some embodiments, the training dataset further includes a synthetic audio stream that has been translated from speech to text, the synthetic audio stream including a large corpus of harmful and non-harmful speech, and may be labeled as including harmful and non-harmful speech and may include timestamps detailing where instances of harm occurred.

[0064] In some embodiments, the training dataset further includes audio streams augmented with different parameters to aid in training the machine learning model to output levels of harmfulness in the audio stream in different situations. For example, the training dataset may be augmented with audio streams that include varying pitch of speakers in the audio stream, noise, codecs, echo, background distractions (e.g., traffic, nature sounds, people speaking unclearly, etc.), music, and playback speed.

[0065] The machine learning module 210 trains a machine learning model using a training dataset in a supervised learning manner. The training dataset includes examples of audio streams with non-harmful content and audio streams with one or more instances of harm, which allows the machine learning module 210 to train the machine learning model to classify input audio streams as containing harmful and non-harmful activity using the distinctions between harmful and non-harmful activity as labels during the supervised learning process.

[0066] In some embodiments, the training data used for the machine learning model includes audio streams collected with user permission for training purposes and labeled by human reviewers. For example, the human reviewers listen to the audio streams in the training data and identify whether each audio stream contains instances of harm, and if so, timestamp the locations within the audio stream where the instances of harm occur. The human-generated data are referred to as ground truth labels. Such training data is then used to train a model; for example, the model during training generates a label for each audio stream in the training data, which is compared to the ground truth label, and a feedback function based on the comparison is used to update one or more model parameters.

[0067] In some embodiments, the training data used for the machine learning model also includes video streams. The training dataset may be labeled to include examples of video streams with non-offensive actions and examples of video streams with one or more offensive actions, allowing the machine learning module 210 to train the machine learning model to classify input video streams as including offensive and non-offensive actions using the distinctions between the offensive and non-offensive actions as labels during a supervised learning process. For example, a training dataset with offensive actions may include a set of actions corresponding to lips moving to form a curse word, arms moving to form movements associated with a curse word for a threshold time period, etc., as well as movements that are precursors to offensive actions, such as arms beginning to move in a particular manner that may result in an offensive action. The video stream may include a video of a user or a video of a user's avatar.

[0068] In some embodiments, the machine learning module 210 is a deep neural network. Types of deep neural networks include convolutional neural networks, deep belief networks, stacked autoencoders, generative adversarial networks, variational autoencoders, flow models, recurrent neural networks, and attention-based models. Deep neural networks use multiple layers to progressively extract higher-level features from raw inputs, where the inputs to the layers are various types of features extracted from other modules, and the output is a decision on whether to perform moderation.

[0069] The machine learning module 210 may generate layers that identify increasingly detailed features and patterns in the speech for the audio stream, with the output of one layer serving as input to subsequent, more detailed layers until the final output is the level of harmfulness in the audio stream. Examples of various layers in a deep neural network may include token embedding, segment embedding, and position embedding.

[0070] In some embodiments, the machine learning module 210 trains the machine learning model using a backpropagation algorithm. The backpropagation algorithm modifies the internal weights of input signals (e.g., at each node / layer of a multi-layer neural network) based on feedback that may be a function of the output labels (e.g., "this part of the audio stream has level 1 toxicity") generated by the model during training with the ground truth labels (e.g., "this part of the audio stream has level 2 toxicity") contained in the training data. Such weight adjustments can improve the accuracy of the model during training.

[0071] After the machine learning module 210 trains the machine learning model, the trained machine learning model receives the voice analysis score for the first user associated with the sending device from the voice analyzer 204, information about the one or more vocal emotion parameters and the one or more vocal emotion scores for the first user from the vocal emotion analyzer 206, and the audio stream from the sending device. The information about the one or more vocal emotion parameters may include information about the user's tone, pitch, and vocal effort level in the audio stream.

[0072] In some embodiments, the machine learning module 210 also receives from the history module 202 the toxicity history of a first user, speaker history and metadata associated with the first user, and listener history and metadata associated with a second user associated with the receiving device. The metadata may be utilized in identifying whether a speaker is likely to behave in a manner that violates community guidelines. For example, a banned user may create a new user profile, but the metadata includes indicators of the same user, such as an IP address and demographic information. In some embodiments, the machine learning module 210 also receives from the text module 208 a text score indicating a toxicity rating for the text. In some embodiments, the trained machine learning model periodically receives the speech analysis score, information regarding one or more vocal emotion parameters, and one or more vocal emotion scores. In some embodiments, the trained machine learning model continuously receives an audio stream, and the trained machine learning model is iteratively applied to the audio stream, with each iteration corresponding to a different portion of the audio stream. For example, the trained machine learning model may be applied to the audio stream every 1 second, every 2 seconds, every 0.5 seconds, etc.

[0073] The trained machine learning model produces as output a level of harmfulness in the audio stream, which is a reflection of how harmful the audio stream currently is and a prediction of when the audio stream could become more harmful.

[0074] In some embodiments, the machine learning module 210 sends the level of harmfulness in the audio stream to the harmfulness module 212. The machine learning module 210 may provide the level of harmfulness each time it is generated by a trained machine learning model.

[0075] The harmfulness module 212 introduces a time delay in the audio stream based on the level of harmfulness and analyzes the audio stream for instances of harmfulness. In some embodiments, the harmfulness module 212 includes a set of instructions executable by the processor 235 to introduce the time delay and analyze the audio stream for instances of harmfulness. In some embodiments, the user interface module 214 may be stored in the memory 237 of the computing device 200 and accessible and executable by the processor 235.

[0076] In some embodiments, the harmfulness module 212 receives a level of harmfulness for the audio stream from the machine learning module 210. The level of harmfulness may correspond to a time delay of the transmission of the audio stream to a receiving device. For example, a level of 0 may indicate that there is no likelihood of harmfulness in the audio stream and that the audio stream can be transmitted without a time delay.

[0077] In some embodiments, the toxicity module 212 determines whether to introduce a time delay in transmitting the audio stream to have sufficient time to analyze the audio stream for instances of toxicity. For example, the time delay can be 0 to 5 seconds. In some embodiments, if the level of toxicity is below a minimum threshold, the time delay is zero seconds, and the toxicity module 212 transmits the audio stream to the receiving device. In some embodiments, if the level of toxicity exceeds the minimum threshold, the toxicity module 212 determines the amount of delay to apply based on the increasing level of toxicity. In some embodiments, the toxicity module 212 also applies a stricter level of scrutiny if the level of toxicity indicates that the audio stream is more likely to be harmful. A greater delay in transmission is a negative feedback mechanism that can discourage harmful behavior without additional moderation by slowing down interaction or preventing effective communication. In some embodiments, the level of toxicity may be high enough that the toxicity module 212 mutes the user. For example, if a user uses abusive language word for word, it may be easier to simply mute the audio stream until the user's abusive behavior is over.

[0078] In some embodiments, the harmfulness module 212 performs a speech-to-text translation of the audio stream or receives a translation from one of the other modules. In some embodiments, the harmfulness module 212 performs a speech-to-text translation after a speaker completes a sentence. In some embodiments, the harmfulness module 212 identifies instances of harmfulness in the audio stream without first converting the audio stream to text.

[0079] The harmfulness module 212 identifies instances of harmfulness within the audio stream. In some embodiments, the harmfulness module 212 gradually adjusts time delays in the transmission of the audio stream to avoid creating perceptible audio artifacts, targeting changes in gaps between words or sentences. For example, the harmfulness module 212 introduces gaps where silence or pauses are present to make the gaps less noticeable. Because pauses and silence are perceptible in less than 100 milliseconds, it takes the harmfulness module 212 to identify pauses and silence in less time than it takes to wait for a sentence to end. In some embodiments, the harmfulness module 212 tracks timestamps of all audio signals / packets to add gaps in appropriate places and help make transitions between speech and silence more seamless.

[0080] In some embodiments, the harmfulness module 212 censors instances of harmfulness in the audio stream by replacing the instances of harmfulness with noise or by replacing the instances of harmfulness with silence.

[0081] In some embodiments, the audio stream is part of the visual signal. The visual signal can be an animation, such as an avatar animation or a physics animation, or the visual signal can be a video stream in which the audio stream is part of the video stream. The visual signal is presented to one or more other users, all of whom participate in a metaverse in which their representations (e.g., their avatars) are in the same area of ​​the metaverse so that each avatar can see the other avatars while the users interact with each other. If the harmfulness module 212 introduces a time delay in the transmission of audio, the harmfulness module 212 synchronizes the time delay with the video signal so that the visual signal has the same time delay as the audio stream.

[0082] In some embodiments, the toxicity module 212 analyzes the video stream for instances of toxicity. In response to the toxicity module 212 identifying an instance of toxicity in the audio stream, the toxicity module 212 may analyze the video stream for objectionable actions occurring within a predetermined time period of the instance of toxicity. For example, turning to FIG. 3 , an exemplary user interface 300 for a video stream is shown, in which an instance of toxicity in the audio stream is identified by the toxicity module 212. The toxicity module 212 performs image recognition on the video stream of a user's avatar to identify locations in the video stream where a speaker's mouth moves to form words corresponding to the instance of toxicity in the audio stream. The toxicity module 212 instructs the user interface module 214 to overlay a graphic 305 over the speaker's mouth while the speaker is performing the objectionable action. Because the graphic 305 draws attention to the objectionable action, other mitigation actions are possible, such as adding blur to the mouth or replacing the mouth with pixels that match the background.

[0083] In some embodiments, the harmfulness module 212 performs motion detection and / or object detection on the video to identify objectionable actions. In response to detecting an objectionable action, the harmfulness module 212 blurs the objectionable action in the video stream or replaces the objectionable action with pixels that match the background. In some embodiments, the harmfulness module 212 may analyze the user's avatar while performing motion detection and / or object detection to determine if the avatar appears agitated and use the avatar as a cue that the user may be performing an objectionable action.

[0084] 4, another exemplary user interface 400 is shown for a video stream in which an offensive action is identified. In this example, the harmful module 212 determines that the user is about to perform an offensive action and that the user appears angry. The harmful module 212 generates a mask 405 that replaces pixels associated with the hand with pixels from the background so that the offensive action is not visible.

[0085] The user interface module 214 generates the user interface. In some embodiments, the user interface module 214 includes a set of instructions executable by the processor 235 to generate the user interface. In some embodiments, the user interface module 214 may be stored in the memory 237 of the computing device 200 and accessible and executable by the processor 235.

[0086] The user interface module 214 generates a user interface for a user 125 associated with the user device 115. The user interface may be used to initiate audio communications with other users, join games in the metaverse, send texts to other users, initiate video communications with other users, etc. In some embodiments, the user interface includes an option for adding user preferences, such as the ability to block other users 125.

[0087] In some embodiments, before a user joins the metaverse, the user interface module 214 generates a user interface that includes information about how the user's information will be collected, stored, and analyzed. For example, the user interface requests the user's permission to use any information associated with the user. The user is notified that the user's information may be deleted by the user and that the user may have the option to select which types of information are provided for different uses. Use of the information complies with applicable regulations, and the data is securely stored. Data collection is not performed in specific locations and for specific user categories (e.g., based on age or other demographics), data collection is temporary (i.e., the data is destroyed after a certain period of time), and the data is not shared with third parties. Some of the data may be anonymized, aggregated across users, or otherwise altered so that the identity of a particular user cannot be identified.

[0088] In some embodiments, the user interface module 214 provides a user interface that explains to the user that the metaverse engine 103 may automatically detect instances of harm and store audio or video of the instances of harm in association with the user's account.

[0089] Exemplary Methods 5 is an exemplary flow diagram 500 for identifying instances of harmfulness in communications. In flow diagram 500, thick lines represent data flows including audio streams, and thin lines represent information data flows.

[0090] The flow diagram 500 includes audio and text communication from a sending device 505 to a receiving device 510. The sending device 505 receives audio input via a microphone, performs analog-to-digital conversion of the audio stream, compresses the audio stream, and sends the audio stream to a real-time server 515. The real-time server 515 sends the audio stream to a fixed-length multi-second buffer 520, which sends the audio stream to a module that performs continuous retrospective speech analysis 525. The continuous retrospective speech analysis 525 is not performed in real time to allow sufficient time to improve the accuracy of the analysis. The module that performs continuous retrospective speech analysis 525 sends the speech analysis to a machine learning module 530. If the machine learning module 530 determines that the audio stream does not need to be delayed, the real-time server 515 sends the audio stream to a stream selector / mute / noise module 555 to be forwarded to the receiving device 510. The real-time server 515 sends the audio stream to an adjustable multi-second buffer 545 if the machine learning module 530 determines that the audio stream needs to be delayed.

[0091] The audio stream is also sent to a module 535 that performs vocal sentiment analysis. The vocal sentiment analysis is sent as input to the machine learning module 530.

[0092] The sending device 505 also receives text input via a keyboard. The sending device 505 performs text encoding and sends the text to a module that performs text moderation 540. The module that performs text moderation 540 sends the text to the receiving device 510 as input to the machine learning module 530.

[0093] The machine learning module 530 also receives as input the toxicity history for each game, the speaker history and metadata, and the listener history and metadata.

[0094] The machine learning module 530 predicts the likelihood of an imminent harm and determines how long to buffer the audio. If there is no likelihood of harm, there is no delay in the audio stream, and the machine learning module 530 sends the audio stream directly to the receiving device 510. If there is a likelihood of harm, the machine learning module 530 sends the audio stream to an adjustable multi-second buffer 545. The adjustable multi-second buffer 545 sends the audio stream to a module for detecting actual harm 550. The module for detecting actual harm 550 sends instances of harm to a module for determining whether to select a stream, mute a stream, or add noise to a stream 555. The stream selector / mute / noise module 555 sends the audio stream to the receiving device 510.

[0095] 6 is another example flow diagram 600 for identifying instances of toxicity in communications, according to some embodiments described herein. In some embodiments, the metaverse engine 103 is stored on the server 101. In some embodiments, the metaverse engine 103 is stored on the user device 115. In some embodiments, the metaverse engine 103 is stored partially on the server 101 and partially on the user device 115.

[0096] The method 600 may begin at block 602. At block 602, an audio stream is received from a sending device. Block 602 may be followed by block 604.

[0097] In block 604, inputs to a trained machine learning model are provided, including the audio stream and speech analysis scores, information about one or more vocal emotion parameters, and one or more vocal emotion scores for a first user associated with the sending device. The trained machine learning model is iteratively applied to portions of the audio stream, with each iteration corresponding to a respective portion of the audio stream. Block 604 may be followed by block 606.

[0098] The trained machine learning model generates as output the level of harmfulness in the audio stream in block 606. Block 606 may be followed by block 608.

[0099] The audio stream is transmitted to the receiving device in block 608. The transmission is performed to introduce a time delay in the audio stream based on the level of harmfulness.

[0100] The methods, blocks, and / or operations described herein may be performed in a different order than illustrated or described, and / or may be performed (partially or completely) concurrently with other blocks or operations as appropriate. Some blocks or operations may be performed on a portion of the data and then performed again later on another portion of the data, for example. Not all of the blocks and operations described need be performed in various implementations. In some implementations, blocks and operations may be performed multiple times in a method, in a different order, and / or at different times.

[0101] Various embodiments described herein involve acquiring data from various sensors in a physical environment, analyzing such data, generating recommendations, and providing a user interface. Data collection is performed only with specific user permission and in accordance with applicable regulations. Data is stored in accordance with applicable regulations, including anonymizing or otherwise modifying the data to protect the user's privacy. Users are provided with clear information regarding data collection, storage, and use and are given the option to select the types of data that may be collected, stored, and utilized. Furthermore, users control the devices on which data may be stored (e.g., user device only, client device + server device, etc.) and the devices on which data analysis is performed (e.g., user device only, client device + server device, etc.). Data is utilized for the specific purposes described herein. Data is not shared with third parties without explicit user permission.

[0102] In the above description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the specification. However, it will be apparent to those skilled in the art that the present disclosure may be practiced without these specific details. In some instances, structures and devices are shown in block diagram form to avoid obscuring the description. For example, embodiments may be described above primarily with reference to a user interface and specific hardware. However, embodiments may apply to any computing device capable of receiving data and commands and any peripheral device that provides services.

[0103] References herein to "some embodiments" or "some examples" mean that a particular feature, structure, or characteristic described in connection with the embodiments or examples may be included in at least one implementation of the description. The appearances of the phrase "in some embodiments" in various places in this specification do not necessarily all refer to the same embodiments.

[0104] Some portions of the above detailed descriptions are presented in terms of algorithms and symbolic representations of operations on data bits within a computer memory. These algorithmic descriptions and representations are the means used by those skilled in the data processing arts to most effectively convey the substance of their work to others skilled in the art. An algorithm is here, and generally, conceived to be a self-consistent sequence of steps leading to a desired result. The steps require physical manipulations of physical quantities. Usually, though not necessarily, these quantities take the form of electrical or magnetic data capable of being stored, transferred, combined, compared, and otherwise manipulated. It has proven convenient at times, principally for reasons of common usage, to refer to these data as bits, values, elements, symbols, characters, terms, numbers, or the like.

[0105] It should be borne in mind, however, that all of these and similar terms are to be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities. As will become apparent from the discussion that follows, unless specifically stated otherwise, throughout the description, discussions utilizing terms including "processing" or "computing" or "calculating" or "determining" or "displaying," etc., will be understood to refer to the actions and processing of a computer system, or similar electronic computing device, that manipulates and transforms data represented as physical (electronic) quantities in the computer system's registers and memory into other data similarly represented as physical quantities in the computer system's memory or registers, or other such information storage, transmission, or display device.

[0106]

[0013] Embodiments herein may also relate to a processor for performing one or more steps of the methods described above. The processor may be a special-purpose processor selectively activated or reconfigured by a computer program stored in the computer. Such a computer program may be stored in a non-transitory computer-readable storage medium, including, but not limited to, any type of disk, including an optical disk, a ROM, a CD-ROM, a magnetic disk, a RAM, an EPROM, an EEPROM, a magnetic or optical card, a flash memory, including a USB key with non-volatile memory, or any type of medium suitable for storing electronic instructions, each coupled to a computer system bus.

[0107] This specification may take the form of some entirely hardware embodiments, some entirely software embodiments, or some embodiments containing both hardware and software elements, hi some embodiments, this specification is implemented in software, including but not limited to firmware, resident software, microcode, etc.

[0108] Furthermore, the descriptions may take the form of a computer program product accessible from a computer-usable or computer-readable medium providing program code for use by or in connection with a computer or any instruction execution system. For purposes of this description, a computer-usable or computer-readable medium may be any device that contains, stores, communicates, propagates, or transfers a program for use by or in connection with an instruction execution system, device, or drive.

[0109] A data processing system suitable for storing or executing program code includes at least one processor coupled directly or indirectly via a system bus to memory elements that may include local memory used during the actual execution of the program code, bulk storage, and cache memory that provides temporary storage of at least some of the program code to reduce the number of times the code must be retrieved from bulk storage during execution. [Explanation of symbols]

[0110] 100 Network environment, environment 101 Server 103 Metaverse Engine 104, 104a, 104b Metaverse Applications 105 Network 115, 115a...n user devices 125, 125a...n users 199 databases 200 computing devices 202 History Module 204 Voice Analyzer 206 Vocal Emotion Analyzer 208 Text Module 210 Machine Learning Module 212 Harmfulness Module 214 User Interface Module 218 Bus 222 signal line 224 signal line 226 Signal Line 228 Signal Line 230 Signal Line 232 signal line 234 Signal Line 235 processor 237 memory 239 Input / Output (I / O) Interface, I / O Interface 241 Microphone 243 Speaker 245 Display 247 Storage Devices 300 User Interface 305 Graphics 400 User Interface 405 Mask 500 Flow Diagram 505 sending device 510 receiving device 515 Real-time Server 520 Multi-Second Buffer 525 Module for performing continuous retrospective speech analysis, Continuous Retrospective Speech Analysis 530 Machine Learning Module 535 Module for performing vocal emotion analysis 540 Text Moderation 545 Adjustable multi-second buffer 550 Actual Harm Detection Module 555 Stream Selector / Mute / Noise Module, a module that determines whether to select a stream, mute a stream, or add noise to a stream, Stream Selector / Mute / Noise Module

Claims

1. 1. A computer-implemented method for determining whether to introduce latency into an audio stream from a particular speaker, comprising: receiving an audio stream from a sending device; providing the audio stream and speech analysis scores, information regarding one or more vocal emotion parameters, and one or more vocal emotion scores for a first user associated with the sending device as inputs to a trained machine learning model, wherein the trained machine learning model is applied to the audio stream iteratively, each iteration corresponding to a respective portion of the audio stream; using the trained machine learning model to generate an output representing a level of harmfulness in the audio stream; transmitting the audio stream to a receiving device, the step being performed to introduce a time delay into the audio stream based on the level of harmfulness; A method comprising:

2. identifying instances of harmfulness within the audio stream; replacing instances of the harmfulness in the audio stream with noise or silence before transmitting the audio stream to the receiving device; The method of claim 1 further comprising:

3. identifying silences or pauses between words in the audio stream, the silence or pause corresponds to a particular timestamp within the audio stream; The method of claim 1 , wherein the time delay is introduced as a gap in the audio stream at the particular timestamp of the silence or pause between words.

4. The method of claim 2 , further comprising updating the audio analysis score based on identifying the instance of harmfulness in the audio stream.

5. receiving text from a text channel associated with the sending device, the text channel being separate from the audio stream; generating a text score indicative of a harmfulness assessment for the text; further comprising The method of claim 1 , wherein the input to the trained machine learning model further includes the text score.

6. 10. The method of claim 1, wherein the input to the trained machine learning model further includes an adverse event history of the first user, a speaker history and metadata associated with the first user, and a listener history and metadata associated with a second user associated with the receiving device.

7. The method of claim 1 , wherein the one or more vocal emotion parameters include tone, pitch, and vocal effort level determined based on one or more previous audio streams from the sending device.

8. 2. The method of claim 1, wherein the audio stream is provided with a visual signal, the method further comprising the step of synchronizing the visual signal with the audio stream by introducing a time delay in the visual signal that is the same as the time delay of the audio stream.

9. the audio stream is part of a video stream, and the method comprises: analyzing the audio stream to identify instances of harmfulness; responsive to identifying the instance of harm, detecting a portion of the video stream depicting an objectionable action, the objectionable action occurring within a predetermined time period of the instance of harm; modifying at least the portion of the video stream in response to detecting the objectionable action by one or more of blurring the portion or replacing the portion with pixels corresponding to a background region; The method of claim 1 further comprising:

10. the audio stream is part of a video stream, and the method comprises: performing motion detection on the video stream to detect objectionable gestures; modifying at least the portion of the video stream in response to detecting the objectionable action by one or more of blurring the portion or replacing the portion with pixels corresponding to a background region; The method of claim 1 further comprising:

11. 2. The method of claim 1, wherein the time delay is zero seconds if the level of harmfulness is below a minimum threshold.

12. a processor; a memory coupled to the processor, the memory, when executed by the processor, causing the processor to: receiving an audio stream from a sending device; providing the audio stream and speech analysis scores, information regarding one or more vocal emotion parameters, and one or more vocal emotion scores for a first user associated with the sending device as inputs to a trained machine learning model, wherein the trained machine learning model is applied to the audio stream iteratively, each iteration corresponding to a respective portion of the audio stream; using the trained machine learning model to generate an output indicating a level of harmfulness in the audio stream; transmitting the audio stream to a receiving device, the transmitting being performed to introduce a time delay into the audio stream based on the level of harmfulness; a memory having instructions for performing operations including 1. A device comprising:

13. The operation is identifying instances of harmfulness within the audio stream; replacing instances of the harmfulness in the audio stream with noise or silence before transmitting the audio stream to the receiving device; 13. The device of claim 12, further comprising:

14. The operation is Identifying silences or pauses between words in the audio stream, the silences or pauses corresponding to particular timestamps in the audio stream. further comprising The device of claim 12 , wherein the time delay is introduced as a gap in the audio stream at the particular timestamp of the silence or pause between words.

15. The operation is updating the audio analysis score based on identifying the instances of harmfulness in the audio stream; 13. The device of claim 12, further comprising:

16. A non-transitory computer-readable medium having stored thereon instructions that, when executed by one or more computers, cause the one or more computers to perform operations, the operations including: receiving an audio stream from a sending device; providing the audio stream and speech analysis scores, information regarding one or more vocal emotion parameters, and one or more vocal emotion scores for a first user associated with the sending device as inputs to a trained machine learning model, wherein the trained machine learning model is applied to the audio stream iteratively, each iteration corresponding to a respective portion of the audio stream; using the trained machine learning model to generate an output indicating a level of harmfulness in the audio stream; transmitting the audio stream to a receiving device, the transmitting being performed to introduce a time delay into the audio stream based on the level of harmfulness; 1. A computer-readable medium comprising:

17. The operation is identifying instances of harmfulness within the audio stream; replacing instances of harmfulness in the audio stream with noise or silence before transmitting the audio stream to the receiving device; 17. The computer-readable medium of claim 16, further comprising:

18. The operation is Identifying silences or pauses between words in the audio stream, the silences or pauses corresponding to particular timestamps in the audio stream. further comprising 17. The computer-readable medium of claim 16, wherein the time delay is introduced as a gap in the audio stream at the particular timestamp of the silence or pause between words.

19. The operation is updating the audio analysis score based on identifying the instances of harmfulness in the audio stream; 17. The computer-readable medium of claim 16, further comprising:

20. The operation is receiving text from a text channel associated with the sending device, the text channel being separate from the audio stream; generating a text score indicative of a harmfulness assessment for the text; further comprising 17. The computer-readable medium of claim 16, wherein the input to the trained machine learning model further comprises the text score.

Citation Information

Patent Citations

  • Method and device for screening audio visual material

    JP1997238321A

  • Program, device, and system

    JP2019211776A

  • Multi-stage adaptive system for content moderation

    US20220115033A1

  • Method and apparatus for screening audio-visual materials presented to a subscriber

    US5757417A