Artificial latency for moderating voice communications
A machine learning model moderates audio streams by identifying and replacing toxic content with noise or silence, introducing time delays to prevent disruption, effectively managing audio stream moderation in online platforms.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- ROBLOX CORP
- Filing Date
- 2023-09-06
- Publication Date
- 2026-07-16
AI Technical Summary
Moderating audio streams in online platforms is challenging due to variations in accent, tone, and volume, making it difficult to analyze and introduce delays without disrupting conversations.
Introduce artificial latency into audio streams by using a trained machine learning model to identify and replace instances of toxicity with noise or silence, synchronized with visual signals, and introduce time delays based on toxicity levels.
Effectively moderates audio streams without perceptible delay, preventing harmful content while maintaining natural conversation flow.
Smart Images

Figure 0007891195000001 
Figure 0007891195000002 
Figure 0007891195000003
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application is an international application and claims the benefit of priority under 35 U.S.C.§119(e) to U.S. Patent Application No. 17 / 940,749, filed on September 8, 2022, entitled ARTIFICIAL LATENCY FOR MODERATING VOICE COMMUNICATION, the entire content of which is incorporated herein by reference.
Background Art
[0002] Online platforms need a way to provide a safe and polite environment in communication between user devices. Text communication is more forgiving due to the latency when a user sends a text message, so it is easier to moderate than audio communication. In addition, moderating text communication can be done by comparing the text to a list of prohibited or problematic words, which is easier than moderating an audio stream. Conversely, moderating an audio stream is more difficult to analyze due to variations in accent, tone, volume, use of irony, etc.
[0003] The background description provided herein is for the purpose of presenting the context of the present disclosure. The research of the inventors whose current names are listed up to the extent described in this background section, as well as aspects of the text of the specification that may not be eligible as prior art at the time of filing, are not admitted as prior art to the present disclosure, either explicitly or implicitly.
Summary of the Invention
Means for Solving the Problems
[0004] Embodiments generally relate to systems and methods for introducing artificial latency into an audio stream for moderation. According to one embodiment, a computer-implemented method includes the step of receiving an audio stream from a transmitting device. The method further includes providing the audio stream and a speech analysis score, information about one or more vocal emotion parameters, and one or more vocal emotion scores relating to a first user associated with the transmitting device as input to a trained machine learning model, the trained machine learning model being iteratively applied to the audio stream, with each iteration corresponding to a respective portion of the audio stream. The method further includes using the trained machine learning model to generate a level of toxicity in the audio stream as an output. The method further includes transmitting the audio stream to a receiving device, the transmitting step being performed to introduce a time delay in the audio stream based on the toxicity level.
[0005] In some embodiments, the method further includes the steps of identifying instances of toxicity in an audio stream and replacing instances of toxicity in the audio stream with noise or silence before transmitting the audio stream to a receiving device. In some embodiments, the method further includes the step of identifying silence or pauses between words in an audio stream, wherein the silence or pause corresponds to a specific timestamp in the audio stream, and the time delay is introduced as a gap in the audio stream at the specific timestamp of the silence or pause between words. In some embodiments, the method further includes updating a speech analysis score based on the identification of instances of toxicity in an audio stream. In some embodiments, the method further includes the steps of receiving text from a text channel associated with a transmitting device, wherein the text channel is separate from the audio stream, and generating a text score indicating toxicity assessment for the text, the input to a trained machine learning model further includes the text score. In some embodiments, the input to a trained machine learning model further includes toxicity history of a first user, speaker history and metadata associated with the first user, and listener history and metadata associated with a second user associated with a receiving device. In some embodiments, one or more vocal emotion parameters include intonation, pitch, and vocal effort level, which are determined based on one or more previous audio streams from a transmitting device. In some embodiments, the audio stream is provided with a visual signal, and the method further includes the step of synchronizing the visual signal with the audio stream by introducing the same time delay in the audio stream as in the visual signal.In some embodiments, the audio stream is part of the video stream, and the method further includes the steps of: analyzing the audio stream to identify instances of toxicity; detecting a portion of the video stream depicting an offensive gesture in response to the identification of an instance of toxicity, wherein the offensive gesture occurs within a predetermined time period of the instance of toxicity; and modifying at least that portion of the video stream in response to the detection of the offensive gesture by blurring that portion or replacing that portion with pixels that match a background area, one or more of which. In some embodiments, the audio stream is part of the video stream, and the method further includes the steps of: performing motion detection on the video stream to detect an offensive gesture; and modifying at least that portion of the video stream in response to the detection of an offensive action by blurring that portion or replacing that portion with pixels that match a background area, one or more of which. In some embodiments, the time delay is zero seconds if the level of toxicity is below a minimum threshold.
[0006] In one embodiment, the device includes a processor and a memory coupled to the processor, which, when executed by the processor, causes the processor to perform an operation that includes: providing the audio stream and a voice analysis score, information about one or more vocal emotion parameters and one or more vocal emotion scores relating to a first user associated with the transmitting device as input to a trained machine learning model, wherein the trained machine learning model is iteratively applied to the audio stream, with each iteration corresponding to a respective portion of the audio stream; using the trained machine learning model, generating a level of toxicity in the audio stream as output; and transmitting the audio stream to the receiving device, which is performed to introduce a time delay in the audio stream based on the level of toxicity.
[0007] In some embodiments, the operation further includes identifying instances of harmfulness in an audio stream and replacing instances of harmfulness in the audio stream with noise or silence before transmitting the audio stream to a receiving device. In some embodiments, the operation further includes identifying silence or pauses between words in an audio stream, wherein the silence or pause corresponds to a specific timestamp in the audio stream, and the time delay is introduced as a gap in the audio stream at the specific timestamp of the silence or pause between words. In some embodiments, the operation further includes updating a speech analysis score based on the identification of instances of harmfulness in the audio stream.
[0008] According to one embodiment, a non-temporary computer-readable medium, when executed by one or more computers, stores instructions causing one or more computers to perform an action, the action comprising: receiving an audio stream from a transmitting device; providing the audio stream and a speech analysis score, information about one or more vocal emotion parameters, and one or more vocal emotion scores relating to a first user associated with the transmitting device as input to a trained machine learning model, wherein the trained machine learning model is iteratively applied to the audio stream, with each iteration corresponding to a respective portion of the audio stream; using the trained machine learning model, generating a level of toxicity in the audio stream as output; and transmitting the audio stream to a receiving device, which is performed to introduce a time delay in the audio stream based on the level of toxicity.
[0009] In some embodiments, the operation further includes identifying instances of toxicity in an audio stream and replacing instances of toxicity in the audio stream with noise or silence before transmitting the audio stream to a receiving device. In some embodiments, the operation further includes identifying silence or pauses between words in an audio stream, wherein the silence or pause corresponds to a specific timestamp in the audio stream, and the time delay is introduced as a gap in the audio stream at the specific timestamp of the silence or pause between words. In some embodiments, the operation further includes updating a speech analysis score based on the identification of instances of toxicity in the audio stream. In some embodiments, the operation further includes receiving text from a text channel associated with a transmitting device, wherein the text channel is separate from the audio stream, and generating a text score indicating toxicity assessment for the text, the input to a trained machine learning model further includes the text score.
[0010] One way to prevent harmful content in an audio stream is to buffer it, identify instances of harmful content before it is sent from the transmitting device to the receiving device, and remove those instances from the audio stream. However, moderating an audio stream introduces a delay of several seconds. Audio delays exceeding 50 milliseconds result in unnatural delays that disrupt conversations, and audio delays exceeding 250 milliseconds break conversations.
[0011] This application favorably describes a metaverse engine and / or metaverse application that provides a method for identifying instances of harmfulness while selectively inserting gaps or pauses into an audio stream in order to perform moderation without perceptible delay. [Brief explanation of the drawing]
[0012] [Figure 1] This is a block diagram of an example network environment for identifying instances of malfunctions in communications, according to some embodiments described herein. [Figure 2] This is a block diagram of an exemplary computing device for identifying instances of malfunctions in communications, according to some embodiments described herein. [Figure 3] This is a diagram illustrating an exemplary user interface for a video stream in which instances of malware are identified, according to some embodiments described herein. [Figure 4] This is a diagram illustrating an exemplary user interface for a video stream in which an unpleasant action is identified, according to some embodiments described herein. [Figure 5] This is an exemplary flowchart for identifying instances of hazards in communications, according to some embodiments described herein. [Figure 6] This is another exemplary flowchart for identifying instances of malfunction in communications, according to some embodiments described herein. [Modes for carrying out the invention]
[0013] Network environment 100 Figure 1 shows a block diagram of an exemplary network environment 100 for identifying instances of malfunction in communications. In some embodiments, the environment 100 includes a server 101, user devices 115a...n, and a network 105. Users 125a...n may be associated with each user device 115a...n. In Figure 1 and the remaining figures, the letters following a reference number, e.g., "115a", represent a reference to an element having that particular reference number. A reference number in text without following letters, e.g., "115", represents a general reference to an element having that reference number. In some embodiments, the environment 100 may include other servers or devices not shown in Figure 1. For example, server 101 could be multiple servers 101.
[0014] Server 101 includes one or more servers, each including a processor, memory, and network communication hardware. In some embodiments, Server 101 is a hardware server. Server 101 is communicably coupled to Network 105. In some embodiments, Server 101 transmits data to and receives data from User Device 115. Server 101 may include a metaverse engine 103 and a database 199.
[0015] In some embodiments, the metaverse engine 103 includes code and routines capable of receiving communications between two or more users in a virtual metaverse, such as communications in the same location within the metaverse, communications within the same metaverse experience, or communications between friends within a metaverse application. Users interact within the metaverse across various attributes (e.g., different ages, regions, languages, etc.).
[0016] In some embodiments, the metaverse engine 103 receives an audio stream from user device 115a to user device 115n. The metaverse engine 103 provides the audio stream and speech analysis scores, information about one or more vocal emotion parameters, and one or more vocal emotion scores for user 125a associated with device 115a as input to a trained machine learning model. The trained machine learning model is iteratively applied to each portion of the received audio stream, such as a few seconds of the audio stream.
[0017] The metaverse engine 103 uses a trained machine learning model to generate a level of toxicity in the audio stream as an output. The metaverse engine 103 transmits the audio stream to one or more other user devices 115n, and the transmission is performed in such a way that a time delay is introduced into the audio stream based on the toxicity level. In some embodiments, the metaverse engine 103 uses the time delay to identify instances of toxicity and replaces instances of toxicity with noise or silence before transmitting the audio stream to one or more other user devices 115n.
[0018] In some embodiments, the metaverse engine 103 is implemented using hardware including a central processing unit (CPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), any other type of processor, or a combination thereof. In some embodiments, the metaverse engine 103 is implemented using a combination of hardware and software.
[0019] Database 199 can be a non - transient computer - readable memory (e.g., random access memory), cache, drive (e.g., hard drive), flash drive, database system, or another type of component or device capable of storing data. Database 199 can also include multiple storage components (e.g., multiple drives or multiple databases) that can span multiple computing devices (e.g., multiple server computers). Database 199 can store data associated with the metaverse engine 103, such as a training dataset for a trained machine learning model, history and metadata associated with each user 125.
[0020] User device 115 can be a computing device that includes a memory, a hardware processor, and a camera. For example, user device 115 can include a mobile device, tablet computer, mobile phone, wearable device, head - mounted display, mobile email device, portable game player, portable music player, reader device, or another electronic device that can access network 105 and capture images using a camera.
[0021] User device 115a includes metaverse application 104a, and user device 115n includes metaverse application 104b. In some embodiments, user device 115a is a transmitting device and user device 115n is a receiving device. In some embodiments, user 125a uses the metaverse application 104a on the transmitting device to generate a communication, such as an audio stream or video stream, and the communication is transmitted to the metaverse engine 103. When the communication is approved for transmission, the metaverse engine 103 transmits the communication to the metaverse application 104b on the receiving device for user 125n to access.
[0022] In the illustrated embodiment, entities of environment 100 are communicatively coupled via network 105. Network 105 can include a public network (e.g., the Internet), a private network (e.g., a local area network (LAN) or wide area network (WAN)), a wired network (e.g., an Ethernet network), a wireless network (e.g., an 802.11 network, Wi-Fi (registered trademark), or wireless LAN (WLAN)), a cellular network (e.g., a long term evolution (LTE) network), routers, hubs, switches, server computers, or combinations thereof. FIG. 1 shows one network 105 coupled to server 101 and user device 115, but in practice, one or more networks 105 can be coupled to these entities.
[0023] Example 200 of a computing device FIG. 2 is a block diagram of an exemplary computing device 200 that can be used to implement one or more features described herein. Computing device 200 can be any suitable computer system, server, or other electronic or hardware device. In some embodiments, computing device 200 is server 101. In some embodiments, computing device 200 is user device 115.
[0024] In some embodiments, the computing device 200 includes a processor 235, memory 237, an input / output (I / O) interface 239, a microphone 241, a speaker 243, a display 245, and a storage device 247. Depending on whether the computing device 200 is a server 101 or a user device 115, some components of the computing device 200 may be absent. For example, if the computing device 200 is a server 101, the computing device may not include the microphone 241 and the speaker 243. In some embodiments, the computing device 200 includes additional components not shown in Figure 2.
[0025] The processor 235 may be connected to the bus 218 via signal line 222, the memory 237 may be connected to the bus 218 via signal line 224, the I / O interface 239 may be connected to the bus 218 via signal line 226, the microphone 241 may be connected to the bus 218 via signal line 228, the speaker 243 may be connected to the bus 218 via signal line 230, the display 245 may be connected to the bus 218 via signal line 232, and the storage device 247 may be connected to the bus 218 via signal line 234.
[0026] The processor 235 includes an arithmetic logic unit, a microprocessor, a general-purpose controller, or some other processor array to perform calculations and provide instructions to the display device. The processor 235 may include various computing architectures, including composite instruction set computer (CISC) architectures, reduced instruction set computer (RISC) architectures, or architectures that implement combinations of instruction sets for processing data. Figure 2 shows a single processor 235, but multiple processors 235 may be included. In different embodiments, the processor 235 may be a single-core processor or a multi-core processor. Other processors (e.g., graphics processing units), operating systems, sensors, displays, and / or physical configurations may be part of the computing device 200.
[0027] Memory 237 stores instructions and / or data that can be executed by processor 235. Instructions may include code and / or routines for performing the techniques described herein. Memory 237 may be a dynamic random access memory (DRAM) device, static RAM, or any other memory device. In some embodiments, memory 237 also includes non-volatile memory such as static random access memory (SRAM) or flash memory, or similar persistent storage devices and media, including hard disk drives, compact disc read-only memory (CD-ROM) devices, DVD-ROM devices, DVD-RAM devices, DVD-RW devices, flash memory devices, or any other mass storage devices for more persistently storing information. Memory 237 includes code and routines that can operate to run the metaverse engine 103, which are described in more detail below.
[0028] The I / O interface 239 can provide the functionality to enable the computing device 200 to interface with other systems and devices. Interfaced devices can be included as part of the computing device 200 or communicate with the computing device 200 separately. For example, network communication devices, storage devices (e.g., memory 237 and / or storage device 247), and input / output devices can communicate via the I / O interface 239. In another example, the I / O interface 239 can receive data from server 101 and supply the data to components of the metaverse engine 103, such as the metaverse engine 103 or machine learning module 210. In some embodiments, the I / O interface 239 can connect to interface devices such as input devices (keyboard, pointing device, touchscreen, microphone 241, sensor, etc.) and / or output devices (display device, speaker 243, monitor, etc.).
[0029] Some examples of interfaced devices that can be connected to the I / O interface 239 include a display 245 that can be used to display content, such as images, videos, and / or the user interface of the output application described herein, and to receive touch (or gesture) input from the user. The display 245 may include any suitable display device such as a liquid crystal display (LCD), light-emitting diode (LED) or plasma display screen, cathode ray tube (CRT), television, monitor, touchscreen, 3D display screen, or other visual display device.
[0030] Microphone 241 includes hardware for detecting spoken voice. Microphone 241 can transmit voice to the metaverse engine 103 via the I / O interface 239.
[0031] Speaker 243 includes hardware for generating audio for playback. For example, after it is determined that the audio stream does not contain any instances of harmful content, speaker 243 receives a command from the metaverse engine 103 to generate audio from another user. Speaker 243 converts the command into audio and generates audio for the user.
[0032] The storage device 247 stores data related to the metaverse engine 103. For example, the storage device 247 may store training datasets for trained machine learning models, history and metadata associated with each user 125, etc. In embodiments where the computing device 200 is the server 101, the storage device 247 is the same as the database 199 in Figure 1.
[0033] Example: Metaverse engine 103 or metaverse application 104 Figure 2 shows a computing device 200 running an exemplary metaverse engine 103 or metaverse application 104, which includes a history module 202, a voice analyzer 204, a voice emotion analyzer 206, a text module 208, a machine learning module 210, a toxicology module 212, and a user interface module 214. Although the modules are shown as being part of the same metaverse engine 103 or metaverse application 104, those skilled in the art will recognize that the modules may be implemented by the computing device 200. For example, the text module 208 may be part of a user device 115 that provides analysis of text communications before they are sent to the metaverse engine 103, which is part of the server 101, in order to reduce the computing requirements of the server 101.
[0034] The history module 202 generates history information and metadata about users participating in the communication. In some embodiments, the history module 202 includes a set of instructions that can be executed by the processor 235 to generate the history information and metadata. In some embodiments, the history module 202 can be stored in the memory 237 of the computing device 200 and made accessible and executable by the processor 235.
[0035] In some embodiments, after obtaining user permission, the history module 202 stores information about each communication session in the metaverse associated with the user, as well as metadata associated with the user. Communication sessions may include audio streams, video streams, text communications, etc. After obtaining user permission, the history module 202 may store information about instances of toxicity associated with the user. For example, the history module 202 may identify when the user participated in an instance of toxicity, what toxic behavior the user performed (e.g., using foul language, performing offensive actions, bullying another user, threatening another user, etc.), and specific users targeted by the instance of toxicity. In all cases where information about a user is stored, the history module 202 has obtained permission from the user, the user is aware that they can delete the information, and the information is stored securely and in accordance with applicable regulations. Further details are discussed below with reference to the user interface module 214.
[0036] In some embodiments, after obtaining user permission, the history module 202 may store contextual information about instances of harmfulness, such as specific experiences. For example, a user might use a lot of swear words when playing a violent shooting game, but not exhibit harmful behavior in a non-violent role-playing game. The history module 202 may also receive information about instances of harmfulness from other modules, such as the voice analyzer 204 and the harmfulness module 212.
[0037] In some embodiments, after obtaining user permission, the history module 202 stores the listening history and metadata about the user's responses to instances of malice directed at the user by other users. For example, the history module 202 may update the listening history and metadata to indicate whether the user is insensitive to slander, whether the user responds to abusive language with abusive language in return, or whether the user reports a user in response to witnessing offensive behavior. In another example, the history module 202 may track how many times a particular user has been blocked and whether that particular user has participated in an event in which another user who blocked that particular user is also present. In some embodiments, the history module 202 generates a sensitivity score that reflects the user's sensitivity to instances of malice on a scale (e.g., 3 out of 10, 0.9 out of 1, etc.).
[0038] In some embodiments, after obtaining user permission, the history module 202 stores metadata associated with the user. For example, the metadata may include the region where the user lives, other demographic information (such as sex, gender, age, race, preferred pronouns, and orientation), one or more Internet Protocol (IP) addresses associated with the user, and one or more languages spoken by the user. In some embodiments, the history module 202 may characterize the user's responses in combination with the metadata. For example, the history module 202 may identify that the user is generally not sensitive to instances of malice in the metaverse unless they are called by abusive language that matches the user's religious affiliation, gender, race, etc.
[0039] The history module 202 may provide history information and metadata to the machine learning module 210 as input to a trained machine learning model. The history module 202 may also provide history information and metadata to other modules to provide context that influences how instances of harm are calculated. For example, the voice analyzer 204 receives history information and metadata because it uses different rules to identify instances of harm among users aged 13-16, 16-18, or 19 and over.
[0040] In some embodiments, the voice analyzer 204 analyzes speech during a communication session. In some embodiments, the voice analyzer 204 includes a set of instructions that can be executed by the processor 235 to analyze speech during a communication session. In some embodiments, the voice analyzer 204 can be stored in the memory 237 of the computing device 200 and made accessible and executable by the processor 235.
[0041] The audio analyzer 204 receives the audio stream from the transmitting device. Since the audio analysis may take several seconds, the audio analyzer 204 performs a continuous analysis of the audio in the audio stream. The analysis may be retrospective in that the audio analyzer 204 performs the analysis after the audio stream has already been sent to the receiving device, or the analysis may occur whenever an instance of malice is identified, regardless of whether the audio stream has been sent to the receiving device.
[0042] In some embodiments, the speech analyzer 204 includes a machine learning model trained to predict various attributes of the audio stream, such as vocal effort, speech pattern, language, and speech activity.
[0043] The speech analyzer 204 may perform automatic speech recognition (ASR), such as translating speech to text, and compare the translated text to a list of harmful words to identify instances of harmfulness in the audio stream. The speech analyzer 204 may generate a speech analysis score for the user associated with the transmitting device. For example, the speech analyzer 204 may generate a speech analysis score based on identifying instances of harmfulness associated with the audio stream.
[0044] In some embodiments, the voice analyzer 204 generates a voice analysis score based on demographics of a particular user. For example, the voice analyzer 204 applies different rubrics to what constitutes an instance of harm based on whether the user is 13-16 years old, 16-18 years old, 19 years old or older (or 12-15 years old, 15-18 years old, 19 years old or older, etc.), based on the user's location, based on whether the audio stream is being sent to a user with different demographic information (for example, an audio stream may be identified as containing an instance of harm if it is sent to a 13-year-old user), or based on the type of game (e.g., shooting game vs. puzzle game).
[0045] In some embodiments, the speech analyzer 204 performs speech-to-text translation based on one or more languages spoken by the user. For example, the aquatic mammal "seal" in English is called "phoque" in French, which should not be confused with the harmful instance, i.e., the English obscene word "fuck." In some embodiments, the identification of one or more languages spoken by the user is received as part of metadata determined by the history module 202.
[0046] In some embodiments, the speech analyzer 204 periodically provides speech analysis, such as a speech analysis score, to the machine learning module 210 as input to a trained machine learning model. The speech analysis score may be associated with a timestamp so that the speech analysis score is aligned with its position in the audio stream. In some embodiments, whenever the speech analyzer 204 identifies an instance of harm in the audio stream, it sends the speech analysis score to the machine learning module 210 and updates the speech analysis score to reflect the identified instance of harm.
[0047] The vocal emotion analyzer 206 analyzes the emotions in the audio stream. In some embodiments, the vocal emotion analyzer 206 includes a set of instructions that can be executed by the processor 235 to analyze the emotions in the audio stream. In some embodiments, the vocal emotion analyzer 206 can be stored in the memory 237 of the computing device 200 and made accessible and executable by the processor 235.
[0048] In some embodiments, the vocal emotion analyzer 206 identifies different speakers in an audio stream and associates each speaker with a corresponding audio identifier. The vocal emotion analyzer 206 analyzes several vocal parameters, such as one or more of each user's tone, pitch, and vocal effort level in the audio stream. In some embodiments, the vocal emotion analyzer 206 analyzes tone by determining the positivity and energy of the user's voice. For example, the vocal emotion analyzer 206 detects whether the user appears excited, irritated, neutral, sad, etc. In some embodiments, the vocal emotion analyzer 206 determines the speaker's emotional state based on an emotion quadrant, which includes four states: nervous, happy, angry, and sad. The vocal emotion analyzer 206 may use converter-based techniques for emotion detection, such as using wav2vec2.0 as part of its front-end.
[0049] In some embodiments, the vocal emotion analyzer 206 analyzes pitches in the range from 60 Hz to 2 kHz. In some embodiments, the vocal emotion analyzer 206 establishes the fundamental frequency of the sound and determines the range of pitches occurring in the audio stream.
[0050] In some embodiments, the vocal emotion analyzer 206 analyzes the vocal effort level by determining the level of noise and comparing that level to a predetermined description of vocal effort. For example, the vocal emotion analyzer 206 may determine that a whisperer produces a vocal effort of 20–30 decibels (dB), a soft speaker produces a vocal effort of 30–55 dB, a speaker at an average level produces a vocal effort of 55–65 dB, a loud speaker or shouter produces a vocal effort of 65–80 dB, and a screamer produces a vocal effort of 80–120 dB. In some embodiments, the vocal emotion analyzer 206 may also identify whether the vocal effort is increasing as a function of time, as this may indicate a conversation escalating into an argument where harassment may occur.
[0051] In some embodiments, the vocal emotion analyzer 206 generates a vocal emotion score for the user associated with the audio stream. In some embodiments, the vocal emotion analyzer 206 generates separate scores for each of the tone, pitch, and vocal effort level, which are determined based on one or more previous audio streams from the transmitting device. In some embodiments, the vocal emotion analyzer 206 generates a vocal emotion score that is a combination of tone, pitch, and vocal effort. For example, a user may not yell when angry, so it may not be clear that a user is angry unless there is a combination of angry tone, large fluctuations in pitch, and a low vocal effort level.
[0052] The voice emotion analyzer 206 can analyze the emotions in an audio stream regardless of whether the user has used voice modulation software. For example, if the user selects voice modulation software that makes them sound like a popular cartoon character, the voice emotion analyzer 206 will detect that modulation is occurring and perform emotion analysis regardless of the modulation.
[0053] The vocal emotion analyzer 206 may periodically generate one or more vocal emotion scores and transmit information about the vocal emotion parameters and one or more vocal emotion scores to the machine learning module 210. For example, whenever the voice analyzer 204 identifies an instance of harmfulness in the audio stream, the vocal emotion analyzer 206 transmits information about the vocal emotion parameters and one or more vocal emotion scores to the machine learning module 210 and updates the voice analysis score to reflect the identified instance of harmfulness. In another example, the vocal emotion analyzer 206 transmits information about the vocal emotion parameters and one or more vocal emotion scores whenever there is a change in the vocal emotion parameters, such as when there is a change in tone, pitch, or vocal effort level, or whenever a change in tone, pitch, or vocal effort level exceeds one or more predetermined thresholds, such as when the user moves from a vocal effort level associated with normal speech to a vocal effort level associated with shouting.
[0054] In some embodiments, the vocal emotion analyzer 206 can be stored on the transmitting device and perform a rapid analysis of vocal emotion. For example, the transmitting device may include a vocal emotion analyzer 206 with an accuracy of 60% in detecting tone. The vocal emotion analyzer 206 on the transmitting device may transmit information to a vocal emotion analyzer 206 on the server 101 that performs a more detailed analysis of vocal emotion.
[0055] The text module 208 analyzes text from a text channel. In some embodiments, the text module 208 includes a set of instructions that can be executed by the processor 235 to analyze the text. In some embodiments, the text module 208 can be stored in the memory 237 of the computing device 200 and made accessible and executable by the processor 235.
[0056] In some embodiments, a user may participate in an audio stream and use a separate text channel to send text messages. For example, during a video call, a user may give a presentation and also add additional information to the chat box associated with the video call. In another example, a user may participate in a game using an audio stream while simultaneously sending text messages directly to a specific user through the game software. Analyzing both audio streams and text messages can be helpful, as some users may be polite on the audio stream but use direct messages during the game to directly insult a specific user.
[0057] In some embodiments, the text module 208 compares the text to a list of harmful words and identifies instances of harmfulness in the text. Based on the text message, the text module 208 may generate a text score indicating a harmfulness rating for the text. The text module 208 may periodically generate the text score and send it to the machine learning module 210. In some embodiments, the text module 208 sends the text score whenever an instance of harmfulness in the text message is identified, and as a result, the text module 208 updates the text score.
[0058] In some embodiments, the text module 208 removes instances of malice from the text before sending it to the receiving device. The text module 208 may simply remove instances of malice, such as a first user threatening another user. Alternatively, the text module 208 may include a warning and an explanation of why instances of malice were removed from the text.
[0059] In some embodiments, the text module 208 is stored on the transmitting device as part of the metaverse application 104a, and the text is analyzed on the transmitting device to conserve computational resources on the server 101. The text module 208 on the transmitting device may send the text score to the metaverse engine 103 as input to the machine learning module 210.
[0060] The machine learning module 210 trains a machine learning model (or multiple models) to output the level of toxicity in an audio stream. In some embodiments, the machine learning module 210 includes a set of instructions that can be executed by the processor 235 to train the machine learning model to output the level of toxicity in an audio stream. In some embodiments, the machine learning module 210 can be stored in the memory 237 of the computing device 200 and made accessible and executable by the processor 235.
[0061] In some embodiments, the machine learning module 210 acquires a training dataset having an audio stream that is manually labeled and paired with outputs from one or more of the following: the history module 202, the speech analyzer 204, the vocal emotion analyzer 206, and the text module 208. In some embodiments, the manual labeling includes instances of toxicity in the audio stream and metadata such as language, emotional state, and vocal effort. For example, the training dataset may include an audio stream paired with outputs from the history module 202 that include toxicity history for a first user (i.e., speaker), speaker history and metadata associated with the first user, and listener history and metadata associated with a second user, outputs from the speech analyzer 204 that include periodically transmitted speech analysis scores for the first user, outputs from the vocal emotion analyzer 206 that include individual speech analysis scores for tone, pitch, and vocal effort level, and outputs from the text module 208 that include text scores.
[0062] In some embodiments, the training dataset further includes automatically labeled audio streams that have been preprocessed for offline toxicity detection. The labels may indicate timestamps within the audio stream where instances of toxicity occurred.
[0063] In some embodiments, the training dataset further includes a synthetic audio stream translated from speech to text, the synthetic audio stream containing a large corpus of harmful and non-harmful speech. The synthetic audio stream may be labeled as containing harmful and non-harmful speech and may include timestamps that detail where instances of harmfulness occurred.
[0064] In some embodiments, the training dataset further includes audio streams augmented with different parameters to assist in training a machine learning model to output levels of harmfulness in audio streams under different circumstances. For example, the training dataset may be augmented with audio streams that include varying speaker pitch, noise, codecs, echoes, distracting background elements (e.g., traffic, nature sounds, people speaking illegibly), music, and playback speed.
[0065] The machine learning module 210 trains a machine learning model using a training dataset in a controlled learning manner. The training dataset includes examples of audio streams with non-harmful content and audio streams with one or more instances of harm, which allows the machine learning module 210 to train its machine learning model to classify input audio streams as containing harmful and non-harmful activities, using the distinction between harmful and non-harmful activities as labels during the controlled learning process.
[0066] In some embodiments, the training data used for a machine learning model includes audio streams collected with the user's permission for training purposes and labeled by a human reviewer. For example, a human reviewer listens to the audio streams in the training data, identifies whether each audio stream contains instances of harmful content, and, if so, times-stamps the location in the audio stream where the instances of harmful content occur. The human-generated data is called ground truth labels. Such training data is then used to train a model, for example, during training, the model generates a label for each audio stream in the training data, this label is compared to the ground truth labels, and a feedback function based on the comparison is used to update one or more model parameters.
[0067] In some embodiments, the training data used for the machine learning model also includes video streams. The training dataset may be labeled to include examples of video streams with non-offensive actions and examples of video streams with one or more offensive actions, which allows the machine learning module 210 to train the machine learning model to classify input video streams as containing offensive and non-offensive actions, using the distinction between offensive and non-offensive actions as labels during a controlled learning process. For example, a training dataset with offensive actions may include a set of actions corresponding to lip movements forming swear words, arm movements forming movements associated with swear words during a threshold time period, and precursory movements of offensive actions, such as the arm beginning to move in a particular way that could result in an offensive action. The video stream may include a video of a user or a video of a user's avatar.
[0068] In some embodiments, the machine learning module 210 is a deep neural network. Types of deep neural networks include convolutional neural networks, deep belief networks, stacked autoencoders, generative adversarial networks, variational autoencoders, flow models, recurrent neural networks, and attention-based models. A deep neural network uses multiple layers to progressively extract higher-level features from raw inputs, where the inputs to the layers are various types of features extracted from other modules, and the output is a decision on whether or not to perform moderation.
[0069] The machine learning module 210 can generate layers that identify increasingly detailed features and patterns within the audio stream, with the output of one layer serving as input to subsequent, more detailed layers until the final output reaches a level of toxicity within the audio stream. Examples of various layers within the deep neural network may include token embeddings, segment embeddings, and positional embeddings.
[0070] In some embodiments, the machine learning module 210 trains a machine learning model using a backpropagation algorithm. The backpropagation algorithm modifies the internal weights of the input signal (e.g., at each node / layer of a multilayer neural network) based on feedback that may be a function of the output label generated by the model during training (e.g., "This part of the audio stream has level 1 harmfulness") using ground truth labels contained in the training data (e.g., "This part of the audio stream has level 2 harmfulness"). Such weight adjustments can improve the accuracy of the model during training.
[0071] After the machine learning module 210 has trained the machine learning model, the trained machine learning model receives a speech analysis score from the speech analyzer 204 regarding a first user associated with the transmitting device, information about one or more vocal emotion parameters and one or more vocal emotion scores regarding the first user from the vocal emotion analyzer 206, and an audio stream from the transmitting device. The information about one or more vocal emotion parameters may include information about the user's tone, pitch, and vocal effort level in the audio stream.
[0072] In some embodiments, the machine learning module 210 also receives from the history module 202 the torrent history of a first user, speaker history and metadata associated with the first user, and listener history and metadata associated with a second user associated with the receiving device. The metadata can be used to identify whether a speaker is likely to engage in behavior that violates community guidelines. For example, an expelled user may create a new user profile, but the metadata includes indicators of it being the same user, such as IP address and demographic information. In some embodiments, the machine learning module 210 also receives from the text module 208 a text score indicating a torrent rating for the text. In some embodiments, the trained machine learning model periodically receives a speech analysis score, information about one or more speech emotion parameters, and one or more speech emotion scores. In some embodiments, the trained machine learning model continuously receives an audio stream, and the trained machine learning model is iteratively applied to the audio stream, with each iteration corresponding to a different portion of the audio stream. For example, the trained machine learning model may be applied to the audio stream every second, every two seconds, every 0.5 seconds, etc.
[0073] The trained machine learning model generates an output representing the level of toxicity within the audio stream. This level of toxicity reflects how toxic the audio stream is currently and predicts how toxic the audio stream may become.
[0074] In some embodiments, the machine learning module 210 transmits the level of toxicity in the audio stream to the toxicity module 212. The machine learning module 210 may provide the toxicity level whenever the toxicity level is generated by the trained machine learning model.
[0075] The toxicity module 212 introduces a time delay in the audio stream based on the toxicity level and analyzes the audio stream for instances of toxicity. In some embodiments, the toxicity module 212 includes a set of instructions that can be executed by the processor 235 to introduce the time delay and analyze the audio stream for instances of toxicity. In some embodiments, the user interface module 214 is stored in the memory 237 of the computing device 200 and can be accessed and executed by the processor 235.
[0076] In some embodiments, the toxicity module 212 receives toxicity levels for an audio stream from the machine learning module 210. The toxicity level may correspond to the time delay of transmitting the audio stream to the receiving device. For example, a level of 0 may indicate that there is no likelihood of toxicity in the audio stream and that the audio stream can be transmitted without time delay.
[0077] In some embodiments, the toxicity module 212 decides whether to introduce a time delay when transmitting the audio stream to have sufficient time to analyze the audio stream for instances of toxicity. For example, the time delay may be between 0 and 5 seconds. In some embodiments, if the toxicity level is below a minimum threshold, the time delay is zero seconds, and the toxicity module 212 transmits the audio stream to the receiving device. In some embodiments, if the toxicity level exceeds a minimum threshold, the toxicity module 212 determines the amount of delay to apply based on the increase in toxicity level. In some embodiments, the toxicity module 212 also applies a stricter level of scrutiny if the toxicity level indicates a higher likelihood that the audio stream is toxic. A larger delay in transmission is a negative feedback mechanism that can suppress toxic behavior without additional moderation by delaying the interaction or hindering effective communication. In some embodiments, the toxicity level may be high enough for the toxicity module 212 to mute the user. For example, if a user uses abusive language word for word, it may be easier to simply mute the audio stream until the user's abusive language ends.
[0078] In some embodiments, the toxicity module 212 performs speech-to-text translation of the audio stream or receives translations from one of the other modules. In some embodiments, the toxicity module 212 performs speech-to-text translation after the speaker has completed a sentence. In some embodiments, the toxicity module 212 identifies instances of toxicity within the audio stream without first converting the audio stream to text.
[0079] The toxicity module 212 identifies instances of toxicity within the audio stream. In some embodiments, the toxicity module 212 progressively adjusts the time delay in transmitting the audio stream to avoid creating perceptible audio artifacts, targeting changes in gaps between words or sentences. For example, the toxicity module 212 introduces gaps where silence or pauses are present to make the gaps less noticeable. Since pauses and silences are perceptible for less than 100 milliseconds, the time it takes for the toxicity module 212 to identify pauses and silences is shorter than waiting for the sentence to end. In some embodiments, the toxicity module 212 tracks timestamps of all audio signals / packets to help add gaps where appropriate and make transitions between utterances and silences more seamless.
[0080] In some embodiments, the toxicity module 212 censors instances of toxicity in the audio stream by replacing them with noise or by replacing them with silence.
[0081] In some embodiments, the audio stream is part of the visual signal. The visual signal may be an animation, such as an avatar animation or a physical animation, and the visual signal may be a video stream, in which the audio stream is part of the video stream. The visual signal is presented to one or more other users, all of whom participate in the metaverse, where the users' representations (e.g., the users' avatars) are in the same area of the metaverse so that each avatar can see the other avatars while the users interact with each other. If the toxicity module 212 introduces a time delay in the transmission of audio, the toxicity module 212 synchronizes the time delay with the video signal so that the visual signal has the same time delay as the audio stream.
[0082] In some embodiments, the toxicity module 212 analyzes the video stream for instances of toxicity. In response to the toxicity module 212 identifying instances of toxicity in the audio stream, the toxicity module 212 may analyze the video stream for offensive actions occurring within a predetermined time period of the toxicity instance. For example, looking at Figure 3, an exemplary user interface 300 of a video stream is shown, where instances of toxicity in the audio stream are identified by the toxicity module 212. The toxicity module 212 performs image recognition on the video stream of the user's avatar and identifies locations in the video stream where the speaker's mouth moves to form a word corresponding to the instance of toxicity in the audio stream. The toxicity module 212 instructs the user interface module 214 to overlay a graphic 305 on the speaker's mouth while the speaker is performing the offensive action. Since the graphic 305 draws attention to the offensive action, other mitigation actions are possible, such as adding a blur to the mouth or replacing the mouth with pixels that match the background.
[0083] In some embodiments, the toxicity module 212 performs motion detection and / or object detection on the video to identify offensive actions. In response to the detection of offensive actions, the toxicity module 212 blurs the offensive actions in the video stream or replaces the offensive actions with pixels that match the background. In some embodiments, the toxicity module 212 may analyze the user's avatar while performing motion detection and / or object detection to determine whether the avatar appears agitated and use the avatar as a signal that the user may be performing an offensive action.
[0084] Turning to Figure 4, another exemplary user interface 400 of a video stream where offensive actions are identified is shown. In this example, the toxicity module 212 determines that the user is about to perform an offensive action and that the user appears angry. The toxicity module 212 generates a mask 405 that replaces pixels associated with the hand with pixels from the background so that the offensive action is not visible.
[0085] The user interface module 214 generates the user interface. In some embodiments, the user interface module 214 includes a set of instructions that can be executed by the processor 235 to generate the user interface. In some embodiments, the user interface module 214 can be stored in the memory 237 of the computing device 200 and made accessible and executable by the processor 235.
[0086] The user interface module 214 generates a user interface for user 125 associated with user device 115. The user interface can be used to initiate audio communication with other users, participate in games in the metaverse, send texts to other users, initiate video communication with other users, and so on. In some embodiments, the user interface includes options for adding user preferences, such as the ability to block other users 125.
[0087] In some embodiments, before a user participates in the metaverse, the user interface module 214 generates a user interface that includes information about how the user's information is collected, stored, and analyzed. For example, the user interface requests permission from the user to use any information associated with the user. The user is informed that user information may be deleted by the user and that the user may have the option to choose which types of information are provided for different purposes. The use of information is subject to applicable regulations, and the data is stored securely. Data collection is not performed in specific locations or against specific user categories (e.g., based on age or other demographics), data collection is temporary (i.e., the data is discarded after a certain period), and the data is not shared with third parties. Some of the data may be anonymized, aggregated across users, or otherwise modified in a way that makes it impossible to identify a particular user.
[0088] In some embodiments, the user interface module 214 provides a user interface that explains to the user that the metaverse engine 103 can automatically detect instances of harmful content and store the audio or video of those instances of harmful content in association with the user's account.
[0089] Exemplary Method Figure 5 is an exemplary flowchart 500 for identifying instances of malfunction in communications. In flowchart 500, thick lines represent data flows containing audio streams, and thin lines represent data flows containing information.
[0090] Flowchart 500 includes audio and text communication from the transmitting device 505 to the receiving device 510. The transmitting device 505 receives audio input via a microphone, performs analog-to-digital conversion of the audio stream, compresses the audio stream, and sends the audio stream to the real-time server 515. The real-time server 515 sends the audio stream to a fixed-length multi-second buffer 520, which then sends the audio stream to a module 525 that performs continuous retrospective speech analysis. The continuous retrospective speech analysis 525 is not performed in real time in order to allow sufficient time to improve the accuracy of the analysis. The module 525 that performs continuous retrospective speech analysis sends the speech analysis to a machine learning module 530. If the machine learning module 530 determines that the audio stream does not need to be delayed, the real-time server 515 sends the audio stream to a stream selector / mute / noise module 555 so that it can be forwarded to the receiving device 510. If the real-time server 515 determines that the machine learning module 530 needs to delay the audio stream, it sends the audio stream to an adjustable multi-second buffer 545.
[0091] The audio stream is also sent to module 535, which performs vocal sentiment analysis. The vocal sentiment analysis is then sent as input to machine learning module 530.
[0092] The transmitting device 505 also receives text input via the keyboard. The transmitting device 505 performs text encoding and sends the text to a module that performs text moderation 540. The module that performs text moderation 540 sends the text to the receiving device 510 as input to the machine learning module 530.
[0093] The machine learning module 530 also receives per-game toxicity history, speaker history and metadata, and listener history and metadata as input.
[0094] The machine learning module 530 predicts the likelihood of an immediate harm occurring and determines how long to buffer the audio. If there is no likelihood of an harm occurring, there is no delay in the audio stream, and the machine learning module 530 sends the audio stream directly to the receiving device 510. If there is a likelihood of an harm occurring, the machine learning module 530 sends the audio stream to an adjustable multi-second buffer 545. The adjustable multi-second buffer 545 sends the audio stream to the module 550 that detects actual harm. The module 550 that detects actual harm sends an instance of the harm to a module 555 that decides whether to select the stream, mute the stream, or add noise to the stream. The stream selector / mute / noise module 555 sends the audio stream to the receiving device 510.
[0095] Figure 6 is another exemplary flowchart 600 for identifying instances of malfunction in communications according to some embodiments described herein. In some embodiments, the metaverse engine 103 is stored on the server 101. In some embodiments, the metaverse engine 103 is stored on the user device 115. In some embodiments, part of the metaverse engine 103 is stored on the server 101 and part of it is stored on the user device 115.
[0096] Method 600 may begin in block 602, where an audio stream from the transmitting device is received. Block 602 may be followed by block 604.
[0097] In block 604, input to a trained machine learning model is provided, including an audio stream, a speech analysis score, information about one or more vocal emotion parameters, and one or more vocal emotion scores for a first user associated with the transmitting device. The trained machine learning model is iteratively applied to portions of the audio stream, with each iteration corresponding to a different portion of the audio stream. Block 604 may be followed by block 606.
[0098] In block 606, the trained machine learning model generates the level of toxicity in the audio stream as its output. Block 608 may follow block 606.
[0099] In block 608, the audio stream is sent to the receiving device. The transmission is performed in such a way that a time delay is introduced into the audio stream based on the level of toxicity.
[0100] The methods, blocks, and / or operations described herein may be executed in an order different from the order illustrated or described, and / or may be executed concurrently (partially or completely) with other blocks or operations as necessary. Some blocks or operations may be executed on a portion of the data and then executed again later on another portion of the data, for example. Not all of the described blocks and operations must be executed in various implementations. In some implementations, blocks and operations may be executed multiple times in a method, in different orders, and / or at different times.
[0101] The various embodiments described herein include acquiring data from various sensors in a physical environment, analyzing such data, generating recommendations, and providing a user interface. Data acquisition is performed only with the permission of a specific user and in accordance with applicable regulations. The data is stored in accordance with applicable regulations, including anonymizing the data or otherwise modifying the data to protect the user's privacy. Users are provided with clear information regarding data collection, storage, and use, and are given the option to select the types of data that may be collected, stored, and used. Furthermore, users control the devices on which data may be stored (e.g., user devices only, client devices + server devices) and the devices on which data analysis is performed (e.g., user devices only, client devices + server devices). The data is used for the specific purposes described herein. The data is not shared with third parties without the explicit permission of the user.
[0102] In the above description, numerous specific details are provided for explanatory purposes and to provide a complete understanding of this specification. However, it will be apparent to those skilled in the art that this disclosure can be implemented without these specific details. In some cases, structures and devices are shown in block diagram form to avoid obscuring the description. For example, embodiments may be described above with reference primarily to user interfaces and specific hardware. However, embodiments can be applied to any computing device capable of receiving data and commands and any peripheral device providing services.
[0103] Any reference in this specification to “some embodiments” or “some examples” means that certain features, structures, or characteristics described in relation to the embodiments or examples may be included in at least one implementation of the description. The phrase “in some embodiments” appearing in various places within this specification does not necessarily refer to the same embodiment.
[0104] Some parts of the detailed explanation above are presented in terms of algorithms and symbolic representations of operations on data bits in computer memory. These algorithmic explanations and representations are the means used by those skilled in data processing techniques to most effectively communicate the content of their work to others skilled in the art. Here, an algorithm is generally considered to be a self-consistent set of steps leading to a desired result. The steps require the physical manipulation of physical quantities. Usually, though not always, these quantities take the form of electrical or magnetic data that can be stored, transferred, combined, compared, and other manipulated. For reasons of general use, it has sometimes proven convenient to refer to these data as bits, values, elements, symbols, characters, terms, numbers, etc.
[0105] However, it should be kept in mind that all of these terms and similar terms should be associated with appropriate physical quantities and are merely convenient labels applied to those quantities. As will be evident from the following discussion, unless otherwise noted, discussions throughout this explanation using terms such as “processing,” “computing,” “calculating,” “decision,” or “display” are understood to refer to the actions and processes of a computer system or similar electronic computing device that manipulate and transform data expressed as physical (electronic) quantities in the registers and memory of a computer system into other data similarly expressed as physical quantities in the computer system memory or registers, or other such information storage, transmission, or display devices.
[0106] Embodiments of this specification may also relate to a processor for performing one or more steps of the methods described above. The processor may be a dedicated processor that is selectively activated or reconfigured by a computer program stored in the computer. Such computer programs may be stored in non-temporary computer-readable storage media, including, but not limited to, any type of disk including optical disks, ROM, CD-ROM, magnetic disk, RAM, EPROM, EEPROM, magnetic card or optical card, flash memory including a USB key with non-volatile memory, or any type of medium suitable for storing electronic instructions, each coupled to a computer system bus.
[0107] This specification may take the form of several entirely hardware embodiments, several entirely software embodiments, or several embodiments that include both hardware and software elements. In some embodiments, this specification is implemented in software, including, but not limited to, firmware, resident software, microcode, etc.
[0108] Furthermore, the description may take the form of a computer program product accessible from a computer-enabled or computer-readable medium that provides program code for use by or in connection with a computer or any instruction execution system. For the purposes of this description, the computer-enabled or computer-readable medium may be any device that can store, communicate, propagate or transfer a program for use by or in connection with an instruction execution system, device, or drive.
[0109] A data processing system suitable for storing or executing program code includes at least one processor directly or indirectly coupled to a memory element via a system bus. The memory element may include local memory used during the actual execution of the program code, bulk storage, and cache memory that provides temporary storage for at least some of the program code to reduce the number of times the code must be retrieved from bulk storage during execution. [Explanation of Symbols]
[0110] 100 Network environment, environment 101 Server 103 Metaverse Engine 104, 104a, 104b Metaverse Applications 105 Network 115, 115a...n User devices 125, 125a...n users 199 Databases 200 computing devices 202 History Module 204 Voice analyzer 206 Voice Emotion Analyzer 208 Text Modules 210 Machine Learning Modules 212 Hazardous Module 214 User Interface Module 218 Bus 222 signal line 224 signal line 226 signal line 228 signal line 230 signal line 232 signal line 234 signal line 235 processors 237 memory 239 Input / Output (I / O) Interface, I / O Interface 241 Microphone 243 speakers 245 displays 247 Storage Devices 300 User Interfaces 305 Graphics 400 User Interfaces 405 Masks 500 Flowcharts 505 Transmitter device 510 Receiving device 515 Real-time Server 520 Multiple-second buffer 525 Module for performing continuous retrospective speech analysis, continuous retrospective speech analysis 530 Machine Learning Modules 535 Module for performing vocal emotion analysis 540 Text Moderation 545 Adjustable multiple-second buffer 550 Module for detecting actual hazards 555 Stream Selector / Mute / Noise Module, a module that determines whether to select a stream, mute a stream, or add noise to a stream, Stream Selector / Mute / Noise Module
Claims
1. A computer-based method for determining whether to introduce latency into an audio stream from a particular speaker, The steps include receiving an audio stream from the transmitting device, A step of providing the audio stream and speech analysis score, information about one or more vocal emotion parameters, and one or more vocal emotion scores relating to a first user associated with the transmitting device as input to a trained machine learning model, wherein the trained machine learning model is iteratively applied to the audio stream, with each iteration corresponding to a respective portion of the audio stream. The steps include: using the trained machine learning model to generate an output of the level of toxicity in the audio stream; A step of identifying silence or pauses between words in the audio stream, A step in which the silence or pause corresponds to a specific timestamp in the audio stream, A step of transmitting the audio stream to a receiving device, wherein a time delay is introduced into the audio stream based on the level of harm, the time delay being introduced as a gap in the audio stream at the specific timestamp of the silence or pause between words, and Methods that include...
2. The steps include identifying instances of harmful content within the aforementioned audio stream, Before transmitting the audio stream to the receiving device, the steps include replacing instances of the harmful elements in the audio stream with noise or silence. The method according to claim 1, further comprising:
3. The step of generating the level of toxicity in the audio stream further includes the step of generating the likelihood that the toxicity occurs within a predetermined time period, the method is The step further includes determining the buffering time of the audio stream based on the likelihood that the harmful effect occurs within a predetermined time period, The method according to claim 1, wherein the transmission step is further based on the buffering time.
4. The method according to claim 2, further comprising the step of updating the audio analysis score based on the identification of instances of the harmfulness in the audio stream.
5. A step of receiving text from a text channel associated with the transmitting device, wherein the text channel is separate from the audio stream; A step of generating a text score that indicates the toxicity assessment for the text. It further includes, The method according to claim 1, wherein the input to the trained machine learning model further includes the text score.
6. The method according to claim 1, wherein the input to the trained machine learning model further includes the first user's harmfulness history, speaker history and metadata associated with the first user, and listener history and metadata associated with a second user associated with the receiving device.
7. The method according to claim 1, wherein the one or more vocal emotion parameters include tone, pitch, and vocal effort level, which are determined based on one or more previous audio streams from the transmitting device.
8. The method according to claim 1, wherein the audio stream is provided with a visual signal, and the method further includes the step of synchronizing the visual signal with the audio stream by introducing the same time delay in the visual signal as the time delay in the audio stream.
9. The audio stream is part of the video stream, and the method is The steps include analyzing the audio stream to identify instances of harmfulness, Steps of detecting a portion of the video stream that depicts an unpleasant action in response to the identification of an instance of the harmfulness, wherein the unpleasant action occurs within a time period of the instance of the harmfulness, In response to the detection of the aforementioned unpleasant action, the steps include modifying at least the portion of the video stream by blurring the portion or replacing the portion with pixels that match the background area, or The method according to claim 1, further comprising:
10. The audio stream is part of the video stream, and the method is The steps include performing motion detection on the video stream to detect an unpleasant gesture, In response to the detection of the aforementioned unpleasant gesture, the steps include modifying at least a portion of the video stream by blurring the portion or replacing the portion with pixels that match the background area, and The method according to claim 1, further comprising:
11. The method according to claim 1, wherein the time delay is zero seconds when the level of toxicity is below a minimum threshold.
12. Processor and A memory coupled to the aforementioned processor, which, when executed by the aforementioned processor, the processor, Receiving an audio stream from the transmitting device, The present invention provides, as input to a trained machine learning model, the audio stream and speech analysis scores, information about one or more vocal emotion parameters, and one or more vocal emotion scores relating to a first user associated with the transmitting device, wherein the trained machine learning model is iteratively applied to the audio stream, with each iteration corresponding to a respective portion of the audio stream. Using the aforementioned trained machine learning model, the level of toxicity in the audio stream is generated as an output. Identifying silence or pauses between words in the audio stream, wherein the silence or pause corresponds to a specific timestamp in the audio stream. Transmitting the audio stream to a receiving device, wherein the transmission is performed such that a time delay is introduced within the audio stream based on the level of harm, and the time delay is introduced as a gap in the audio stream at the specific timestamp of the silence or pause between words. A memory containing instructions that cause an operation including A device equipped with the following features.
13. The aforementioned operation, Identifying instances of harmful content within the aforementioned audio stream, Before transmitting the audio stream to the receiving device, replace instances of the harmful elements in the audio stream with noise or silence. The device according to claim 12, further comprising:
14. The device according to claim 12, wherein the audio stream is part of the metaverse.
15. The aforementioned operation, Updating the audio analysis score based on the identification of instances of the harmfulness within the audio stream. The device according to claim 12, further comprising:
16. A non-temporary computer-readable medium that, when executed by one or more computers, stores instructions that cause the one or more computers to perform an action, wherein the action is Receiving an audio stream from the transmitting device, The present invention provides, as input to a trained machine learning model, the audio stream and speech analysis scores, information about one or more vocal emotion parameters, and one or more vocal emotion scores relating to a first user associated with the transmitting device, wherein the trained machine learning model is iteratively applied to the audio stream, with each iteration corresponding to a respective portion of the audio stream. Using the aforementioned trained machine learning model, the level of toxicity in the audio stream is generated as an output. Identifying silence or pauses between words in the audio stream, wherein the silence or pause corresponds to a specific timestamp in the audio stream. Transmitting the audio stream to a receiving device, wherein the transmission is performed such that a time delay is introduced within the audio stream based on the level of harm, and the time delay is introduced as a gap in the audio stream at the specific timestamp of the silence or pause between words. Computer-readable media, including [specific examples of computer-readable media].
17. The aforementioned operation, Identifying instances of harmful content within the aforementioned audio stream, Before transmitting the audio stream to the receiving device, replace instances of harmful content in the audio stream with noise or silence. The computer-readable medium according to claim 16, further comprising:
18. The computer-readable medium according to claim 16, wherein the audio stream is part of the metaverse.
19. The aforementioned operation, Updating the audio analysis score based on the identification of instances of the harmfulness within the audio stream. The computer-readable medium according to claim 16, further comprising:
20. The aforementioned operation, Receiving text from a text channel associated with the transmitting device, wherein the text channel is separate from the audio stream, To generate a text score indicating a toxicity assessment for the aforementioned text. It further includes, The computer-readable medium according to claim 16, wherein the input to the trained machine learning model further comprises the text score.