Abuse of three-dimensional avatars

CN122804256APending Publication Date: 2026-09-22ROBLOX CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480076063.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-29
Filing Date
2024-10-30
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

检测诸如化身的身体或服饰上的滥用性文本等滥用是困难的,因为虚拟世界中的用户生成内容体量极其庞大,使得难以对所有内容进行实时人工查阅与审核

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122804256A_ABST
    Figure CN122804256A_ABST
Patent Text Reader

Abstract

A method includes receiving a three-dimensional (3D) avatar. The method also includes generating two-dimensional (2D) images of the 3D avatar from different angles around the 3D avatar. The method also includes providing the 2D images as input to a trained machine learning model. The method also includes generating, with the machine learning model, a patch embedding for the 2D images. The method also includes analyzing, with the machine learning model, at least one attribute associated with the 3D avatar based on the patch embedding, wherein the at least one attribute is selected from a group of shapes of the 3D avatar, attire on the 3D avatar, and combinations thereof. The method also includes outputting, with the machine learning model, a determination that the at least one attribute of the 3D avatar is abusive.
Need to check novelty before this filing date? Find Prior Art

Description

Cross-references to related applications

[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 615,977, filed December 29, 2023, entitled "Detecting Abuse Text on Three-Dimensional Avatars," the contents of which are incorporated herein by reference in their entirety, pursuant to 35 USC § 119(e). Background Technology

[0002] Abuse in virtual environments occurs in a variety of ways. For example, avatars may wear offensive clothing, players may perform offensive actions, players may utter offensive language, and players may type offensive words in group chats. Detecting abuse such as abusive text on an avatar's body or clothing is difficult because the volume of user-generated content in virtual worlds is so vast that real-time manual review and moderation of all content is impossible. Furthermore, avatars can quickly change their clothing or appearance, requiring continuous monitoring. The longer the delay between a violation and its consequences, the greater the likelihood that players will continue to abuse in the virtual environment.

[0003] The background description provided herein is intended to present the context of this disclosure. Within the scope described in this background section, the work of the inventors listed herein, and aspects of the description that may not conform to the prior art at the time of filing, are neither explicitly nor implicitly acknowledged as prior art to this disclosure. Summary of the Invention

[0004] A method includes receiving a three-dimensional (3D) avatar. The method also includes generating two-dimensional (2D) images of the 3D avatar from different angles surrounding it. The method further includes feeding the 2D images as input to a trained machine learning model. The method also includes generating a stitched embedding of the 2D images using the machine learning model. The method further includes using the machine learning model to analyze at least one attribute associated with the 3D avatar based on the stitched embedding, wherein the at least one attribute is selected from groups of the 3D avatar's shape, clothing worn by the 3D avatar, and combinations thereof. The method also includes using the machine learning model to output a determination that at least one attribute of the 3D avatar is abusive.

[0005] In some embodiments, the method further includes: detecting text on one or more 2D images in a 2D image of the 3D avatar; extracting text from the one or more 2D images; and classifying the extracted text to determine whether the extracted text is abusive, wherein the determination of whether an attribute is abusive is based on classifying the extracted text as abusive. In some embodiments, the 3D avatar is generated by a user, and the method further includes providing a user with notification, in response to an output determination, of one or more of the following: attributes of the 3D avatar that are abusive, an identifier of an abuse category, the severity of the abuse of the attribute, the reason why the attribute is abusive, and combinations thereof. In some embodiments, the method further includes: receiving an updated 3D avatar from a user; generating an updated 2D image of the updated 3D avatar from different angles surrounding the updated 3D avatar; providing the updated 2D image as input to a machine learning model; utilizing the machine learning model to output a determination that an attribute is acceptable; and providing the user with the option to use the updated 3D avatar in a virtual environment. In some embodiments, the machine learning model is trained using training data, and the method further includes: in response to the updated 3D avatar being used in a virtual environment, receiving an abuse report from a player in the virtual environment describing the updated 3D avatar as having abusive attributes; determining that the updated 3D avatar has abusive attributes; and updating the training data associated with the machine learning model to include the updated 2D image and a label with an abuse category.

[0006] In some embodiments, the machine learning model includes a convolutional neural network (CNN) that generates a stitched embedding of 2D images by: extracting a corresponding view embedding for each 2D image using pooling layers; and stitching the view embeddings to form a stitched embedding. In some embodiments, the machine learning model is trained using synthetic training data for text, and the synthetic training data is generated by: identifying abusive categories of text, wherein the determination that an attribute is abusive is associated with a confidence score below a threshold confidence value; generating abusive text associated with the identified categories; generating a training 3D avatar; projecting the abusive text onto the training 3D avatar; generating a training 2D image from the training 3D avatar having the projected abusive text; applying a label including the abusive category to the training 2D image; and training the machine learning model to minimize the difference between the predicted abusive category for the training 2D image and the label for the training 2D image. In some embodiments, the synthetic training data is also generated by shuffling the training 2D image and enhancing the training 2D image using random color jitter. In some embodiments, the machine learning model is trained using training data, and the method further includes updating the training data associated with the machine learning model to include training images of attributes recently identified as abusive. In some embodiments, generating 2D images of 3D avatars includes using a virtual camera system to capture high-resolution thumbnails.

[0007] A non-transitory computer-readable medium having instructions that, when executed by one or more processors at a client device, cause the one or more processors to perform operations. These operations include: receiving a 3D avatar; generating 2D images of the 3D avatar from different angles surrounding the 3D avatar; providing the 2D images as input to a trained machine learning model; generating a stitched embedding of the 2D images using the machine learning model; analyzing, based on the stitched embedding, at least one attribute associated with the 3D avatar using the machine learning model, wherein the at least one attribute is selected from a group consisting of the shape of the 3D avatar, clothing worn by the 3D avatar, and combinations thereof; and outputting a determination that at least one attribute of the 3D avatar is abusive using the machine learning model.

[0008] In some embodiments, these operations further include: detecting text on one or more 2D images in a 2D image of the 3D avatar; extracting text from the one or more 2D images; and classifying the extracted text to determine whether the extracted text is abusive, wherein the determination of whether an attribute is abusive is based on classifying the extracted text as abusive. In some embodiments, the 3D avatar is generated by a user, and these operations further include providing a user with notification, in response to an output determination, of one or more of the following: attributes of the 3D avatar that are abusive, an identifier of an abuse category, the severity of the abuse of the attribute, the reason why the attribute is abusive, and combinations thereof. In some embodiments, these operations further include: receiving an updated 3D avatar from a user; generating an updated 2D image of the updated 3D avatar from different angles surrounding the updated 3D avatar; providing the updated 2D image as input to a machine learning model; utilizing the machine learning model to output a determination that an attribute is acceptable; and providing the user with the option to use the updated 3D avatar in a virtual environment. In some embodiments, the machine learning model is trained using training data, and these operations further include: receiving an abuse report from a player in the virtual environment describing the updated 3D avatar as having abusive attributes in response to the updated 3D avatar being used in the virtual environment; determining that the updated 3D avatar has abusive attributes; and updating the training data associated with the machine learning model to include the updated 2D image and a label with an abuse category.

[0009] A system comprising: a processor; and a memory coupled to the processor, the memory storing instructions that, when executed by the processor, cause the processor to perform operations. These operations include: receiving a 3D avatar; generating 2D images of the 3D avatar from different angles surrounding the 3D avatar; providing the 2D images as input to a trained machine learning model; generating a stitched embedding of the 2D images using the machine learning model; analyzing, based on the stitched embedding, at least one attribute associated with the 3D avatar using the machine learning model, wherein the at least one attribute is selected from a group consisting of the shape of the 3D avatar, clothing worn by the 3D avatar, and combinations thereof; and outputting a determination regarding the at least one attribute of the 3D avatar as abusive using the machine learning model.

[0010] In some embodiments, these operations further include: detecting text on one or more 2D images in a 2D image of the 3D avatar; extracting text from the one or more 2D images; and classifying the extracted text to determine whether the extracted text is abusive, wherein the determination of whether an attribute is abusive is based on classifying the extracted text as abusive. In some embodiments, the 3D avatar is generated by a user, and these operations further include providing a user with notification, in response to an output determination, of one or more of the following: attributes of the 3D avatar that are abusive, an identifier of an abuse category, the severity of the abuse of the attribute, the reason why the attribute is abusive, and combinations thereof. In some embodiments, these operations further include: receiving an updated 3D avatar from a user; generating an updated 2D image of the updated 3D avatar from different angles surrounding the updated 3D avatar; providing the updated 2D image as input to a machine learning model; utilizing the machine learning model to output a determination that an attribute is acceptable; and providing the user with the option to use the updated 3D avatar in a virtual environment. In some embodiments, the machine learning model is trained using training data, and these operations further include: receiving an abuse report from a player in the virtual environment describing the updated 3D avatar as having abusive attributes in response to the updated 3D avatar being used in the virtual environment; determining that the updated 3D avatar has abusive attributes; and updating the training data associated with the machine learning model to include the updated 2D image and a label with an abuse category. Attached Figure Description

[0011] Figure 1 This is a block diagram of an example network environment based on some embodiments described herein.

[0012] Figure 2A This is a block diagram of an example computing device according to some embodiments described herein.

[0013] Figure 2B This is a block diagram of an example machine learning module based on some embodiments described herein.

[0014] Figure 3 These are example illustrations of a process for creating synthetic training data according to some embodiments described herein.

[0015] Figure 4 These are examples of processes for creating two-dimensional images of three-dimensional avatars according to some embodiments described herein.

[0016] Figure 5 This is an example of a process for classifying text on an avatar according to some embodiments described herein.

[0017] Figure 6These are examples of example outputs from a machine learning model that classifies text, based on some embodiments described herein.

[0018] Figure 7 These are example multi-view fusion models based on some embodiments described herein.

[0019] Figure 8 Examples of user interfaces illustrating warnings to users regarding potentially disciplining actions that may result from abusive avatars generated during avatar creation, according to some embodiments described herein.

[0020] Figure 9 This is a flowchart of an example method for generating synthetic training data for a machine learning model, according to some embodiments described herein.

[0021] Figure 10 This is a flowchart of an example method for training a machine learning model to identify abusive incarnations, based on some embodiments described herein.

[0022] Figure 11 This is a flowchart of an example method for auditing avatars, based on some embodiments described herein. Detailed Implementation

[0023] Overview Detecting abusive avatars in virtual environments in real time is challenging due to the inherent delays in waiting for auditors to review 3D avatars for violations. Designing machine learning models to identify abusive 3D avatars for virtual experiences has been difficult because too many false positives (i.e., identifying too many instances of abuse that aren't actually abusive) can alienate users and reduce their willingness to create avatars for virtual environments. Conversely, too many false negatives (i.e., failing to identify instances of abuse) risk exposing too many users to abusive and insecure virtual experiences.

[0024] One way machine learning models fail to identify abusive 3D avatars is by examining a two-dimensional (2D) image of the avatar alone. This leads to too many false negatives because a 2D image might appear acceptable when the violation can be identified on the 3D avatar. For example, an avatar with an inappropriate body shape might not be identifiable when viewed from the front, but it might be identifiable from the side.

[0025] The technique described below advantageously describes the use of an auditing application that incorporates a machine learning model to determine whether a 3D avatar is abusive, based on text, shape, and clothing. The auditing application generates multiple 2D images of the 3D avatar from different angles surrounding it.

[0026] The review application determines whether the 3D avatar contains abusive text by performing text detection on each 2D image. In response to the detection of text on the 2D image, the text is extracted. The extracted text is then categorized to determine whether it is abusive. This advantageously ensures that abusive text can be detected even when it is not readily apparent, such as text located on one side of the avatar's torso or distorted based on clothing fit.

[0027] The review application feeds 2D images to a machine learning model, such as a convolutional neural network (CNN), which extracts a corresponding view embedding for each of the multiple 2D images and stitches the view embeddings together to form a stitched embedding. This stitched embedding is then fed into a multilayer perceptron (MLP) that outputs a determination of whether a 3D avatar is abusive. The MLP can be trained to recognize different abusive shapes or abusive clothing. In some embodiments, the stitched embedding is fed to two different MLPs that are respectively trained to identify abusive shapes or abusive clothing.

[0028] In some embodiments, an auditing application is used whenever a user creates a 3D avatar, creates clothing for the 3D avatar, or responds to an abuse report submitted by a player. This improves the accuracy of the auditing application, making the virtual experience safe and enjoyable for players.

[0029] Example network environment Figure 1 An example network environment 100 according to some embodiments of the present disclosure is illustrated. Figure 1 The same reference numerals are used to identify the same elements as in other figures. Letters following the reference numerals, such as "110a", indicate the element specifically referred to in the text that has that particular reference numeral. Reference numerals without a letter in the text, such as "110", refer to any or all elements in the figure that bear that reference numeral (e.g., "110" in the text refers to reference numerals "110a", "110b", and / or "110n" in the figure).

[0030] The network environment 100 (also referred to herein as the “platform”) includes an online virtual experience server 102, a data storage device 108, and client devices 110 (or multiple client devices), all of which are connected via a network 122.

[0031] The online virtual experience server 102 may include a virtual experience engine 104, one or more virtual experiences 105, and an auditing application 130, etc. In some implementations, the online virtual experience server 102 may be configured to provide virtual experiences 105 to one or more client devices 110 and audit 3D avatars via the auditing application 130.

[0032] Data storage device 108 is shown coupled to online virtual experience server 102, but in some embodiments, it may also be provided as part of online virtual experience server 102. In some embodiments, the data storage device may be configured to store advertising data, user data, engagement data, and / or other contextual data associated with the review application 130.

[0033] Client devices 110 (e.g., 110a, 110b, 110n) may include virtual experience applications 112 (e.g., 112a, 112b, 112n) and I / O interfaces 114 (e.g., 114a, 114b, 114n) to interact with online virtual experience servers 102 and view, for example, a graphical user interface (GUI) through a computer monitor or display (not illustrated). In some implementations, client devices 110 may be configured to execute and display virtual experiences, which may include virtual user engagement portals as described herein.

[0034] Network environment 100 is provided for illustration only. In some implementations, network environment 100 may include interfaces with... Figure 1 The same, fewer, more, or different elements configured in the same or different ways as shown.

[0035] In some implementations, network 122 may include a public network (e.g., the Internet), a private network (e.g., a local area network (LAN) or a wide area network (WAN)), a wired network (e.g., Ethernet), a wireless network (e.g., an 802.11 network, a Wi-Fi® network, or a wireless LAN (WLAN)), a cellular network (e.g., a Long Term Evolution (LTE) network), a router, a hub, a switch, a server computer, or a combination thereof.

[0036] In some implementations, the data storage device 108 may be a non-transitory computer-readable storage device (e.g., random access memory), a cache, a drive (e.g., a hard disk drive), a flash drive, a database system, or another type of component or device capable of storing data. The data storage device 108 may also include multiple storage components (e.g., multiple drives or multiple databases) that may span multiple computing devices (e.g., multiple server computers).

[0037] In some implementations, the online virtual experience server 102 may include a server with one or more computing devices (e.g., a cloud computing system, rack server, server computer, physical server cluster, virtual server, etc.). In some implementations, the server may be included in the online virtual experience server 102, and may be a standalone system or part of another system or platform. In some implementations, the online virtual experience server 102 may be a single server, or any combination of multiple servers, load balancers, network devices, and other components. The online virtual experience server 102 may also be implemented on a physical server, but in some implementations, virtualization technology may be utilized. Other variations of the online virtual experience server 102 are also applicable.

[0038] In some implementations, the online virtual experience server 102 may include one or more computing devices (such as rack servers, router computers, server computers, personal computers, mainframe computers, laptop computers, tablet computers, desktop computers, etc.), data storage devices (e.g., hard disks, memory, databases), networks, software components, and / or hardware components that can be used to perform operations on the online virtual experience server 102 and provide access to the online virtual experience server 102 to users (e.g., users 114 via client device 110).

[0039] The online virtual experience server 102 may also include a website (e.g., one or more web pages) or application backend software that can be used to provide users with access to content provided by the online virtual experience server 102. For example, a user (or developer) may access the online virtual experience server 102 using the virtual experience application 112 on the client device 110.

[0040] In some implementations, the online virtual experience server 102 may include digital asset and digital virtual experience generation capabilities. For example, the platform may provide an administrator interface that allows for design, modification, individualized customization, and other modifications. In some implementations, virtual experiences may include, for example, two-dimensional (2D) games, three-dimensional (3D) games, virtual reality (VR) games, or augmented reality (AR) games. In some implementations, virtual experience creators and / or developers may be able to search for virtual experiences, combine portions of virtual experiences, customize virtual experiences for specific activities (e.g., group virtual experiences), and other features provided by the virtual experience server 102.

[0041] In some implementations, the online virtual experience server 102 or client device 110 may include a virtual experience engine 104 or a virtual experience application 112. In some implementations, the virtual experience engine 104 may be used for the development or execution of the virtual experience 105. For example, the virtual experience engine 104 may include a rendering engine (“renderer”) for 2D, 3D, VR, or AR graphics, a physics engine, a collision detection engine (and collision response), a sound engine, scripting capabilities, a haptic engine, an artificial intelligence engine, networking capabilities, streaming capabilities, memory management capabilities, threading capabilities, scene graph capabilities, or video support for story animations, and other features. Components of the virtual experience engine 104 may generate commands (e.g., rendering commands, collision commands, physics commands, etc.) to aid in the calculation and rendering of the virtual experience.

[0042] The online virtual experience server 102 using virtual experience engine 104 can execute some or all of the virtual experience engine functions (e.g., generating physics commands, rendering commands, etc.), or offload some or all of the virtual experience engine functions to the virtual experience engine 104 on the client device 110 (not illustrated). In some implementations, each virtual experience 105 may have a different ratio between the virtual experience engine functions executed on the online virtual experience server 102 and the virtual experience engine functions executed on the client device 110.

[0043] In some implementations, virtual experience instructions may refer to instructions that allow client device 110 to render gameplay, graphics, and other features of the virtual experience. Instructions may include one or more of user input (e.g., physical object location), character location and velocity information, or commands (e.g., physics commands, rendering commands, collision commands, etc.).

[0044] In some embodiments, client devices 110 may each include computing devices such as personal computers (PCs), mobile devices (e.g., laptops, mobile phones, smartphones, tablets, or netbooks), network-connected televisions, game consoles, etc. In some embodiments, client device 110 may also be referred to as "client device 110". In some embodiments, one or more client devices 110 may connect to the online virtual experience server 102 at any given time. It should be noted that the number of client devices 110 is provided as illustrative rather than limiting. In some embodiments, any number of client devices 110 may be used.

[0045] In some implementations, each client device 110 may include an instance of a virtual experience application 112. The virtual experience application 112 may be rendered for interaction at the client device 110. During user interaction within a virtual experience on the online platform 100 or within another GUI, a user can create avatars including different body parts from different libraries. An auditing application 130 may receive the 3D avatar, generate two-dimensional (2D) images of the 3D avatar from different angles surrounding it, and determine that at least one attribute of the 3D avatar is abusive.

[0046] Example computing device Figure 2A This is a block diagram of an example computing device 200 that can be used to implement one or more features described herein. The computing device 200 can be any suitable computer system, server, or other electronic or hardware device. In some embodiments, the computing device 200 is a client device 110. In some embodiments, the computing device 200 is an online virtual experience server 102.

[0047] In some embodiments, the computing device 200 includes a processor 235, a memory 237, an input / output (I / O) interface 239, a microphone 241, a speaker 243, a display 245, and a storage device 247, all of which are coupled via a bus 218. In some embodiments, the computing device 200 includes Figure 2A Additional components not illustrated herein. In some embodiments, computing device 200 includes components not shown. Figure 2A The example shows fewer components. For example, in audit application 130, the components are stored... Figure 1 In the case of an online virtual experience server 102, the computing device may not include a microphone 241, a speaker 243, or a monitor 245.

[0048] Processor 235 can be coupled to bus 218 via signal line 222, memory 237 can be coupled to bus 218 via signal line 224, I / O interface 239 can be coupled to bus 218 via signal line 226, microphone 241 can be coupled to bus 218 via signal line 228, speaker 243 can be coupled to bus 218 via signal line 230, display 245 can be coupled to bus 218 via signal line 232, and storage device 247 can be coupled to bus 218 via signal line 234.

[0049] Processor 235 includes an arithmetic logic unit, a microprocessor, a general-purpose controller, or some other processor array to perform computations and provide instructions to a display device. Processor 235 processes data and may include various computing architectures, including Complex Instruction Set Computer (CISC) architecture, Reduced Instruction Set Computer (RISC) architecture, or architectures implementing instruction set combinations. In some embodiments, processor 235 may include dedicated units, such as machine learning processors, audio / video encoding and decoding processors, etc. Although Figure 2A A single processor 235 is illustrated, but multiple processors 235 may be included. In different embodiments, processor 235 may be a single-core processor or a multi-core processor. Other processors (e.g., a graphics processing unit), operating system, sensors, display, and / or physical configuration may be part of computing device 200, such as a keyboard, mouse, etc.

[0050] Memory 237 stores instructions and / or data that can be executed by processor 235. Instructions may include code and / or routines for performing the techniques described herein. Memory 237 may be a dynamic random access memory (DRAM) device, static RAM, or some other memory device. In some embodiments, memory 237 also includes non-volatile memory, such as a static random access memory (SRAM) device or flash memory, or similar permanent storage devices and media, including hard disk drives, optical disc read-only memory (CD-ROM) devices, DVD-ROM devices, DVD-RAM devices, DVD-RW devices, flash memory devices, or some other high-capacity storage devices for more permanent storage of information. Memory 237 includes code and routines operable to execute audit application 130, which is described in more detail below.

[0051] I / O interface 239 can provide functionality that enables computing device 200 to interface with other systems and devices. Interface devices can be included as part of computing device 200, or they can be independent and communicate with computing device 200. For example, network communication devices, storage devices (e.g., memory 237 and / or storage device 247), and input / output devices can communicate via I / O interface 239. In another example, I / O interface 239 can receive data from online virtual experience server 102 and deliver the data to auditing application 130 and components of auditing application 130, such as user interface module 202. In some embodiments, I / O interface 239 can connect to interface devices such as input devices (keyboard, pointing device, touchscreen, microphone 241, sensors, etc.) and / or output devices (display 245, speaker 243, etc.).

[0052] Some examples of interface devices that can be connected to I / O interface 239 may include display 245, which can be used to display content such as images, videos, and / or a user interface of the metaverse as described herein, and to receive touch (or gesture) input from a user. Display 245 may include any suitable display device, such as a liquid crystal display (LCD), a light-emitting diode (LED) or plasma display, a cathode ray tube (CRT), a television, a monitor, a touch screen, a 3D display, a projector (e.g., a 3D projector), or other visual display device.

[0053] Microphone 241 includes hardware, such as one or more microphones for detecting audio spoken by a person. Microphone 241 can send audio to auditing application 130 via I / O interface 239.

[0054] The speaker 243 includes hardware for generating audio for playback. In some embodiments, the speaker 243 may include audio hardware that supports playback via an external, stand-alone speaker (e.g., wired or wireless headphones, external speakers, or other audio playback devices) coupled to the computing device 200.

[0055] Storage device 247 stores data related to auditing application 130. For example, storage device 247 may store user profiles associated with user 125, lists of blocked avatars, training data, embeddings from machine learning models, etc.

[0056] Example audit application Figure 2A A computing device 200 is illustrated that executes an example auditing application 130, which includes a user interface module 202, a text recognition engine 204, a machine learning module 206, and an abuse module 208. In some embodiments, a single computing device 200 includes Figure 2A All components illustrated herein. In some embodiments, one or more of these components may reside on different computing devices 200. For example, client device 110 may include user interface module 202, while text recognition engine 204, machine learning module 206, and abuse module 208 are implemented on online virtual experience server 102. In some embodiments, different portions of one or more of modules 202, 204, 206, and 208 may be implemented on client device 110 and online virtual experience server 102. For example, machine learning module 206 may reside on online virtual experience server 102 to mitigate security risks arising from machine learning module 206 residing on client device 110.

[0057] User interface module 202 generates graphical data for displaying a user interface to a user associated with client device 110 to participate in a virtual experience. In some embodiments, before a user participates in the virtual experience, user interface module 202 generates a user interface that includes information on how user information may be collected, stored, and / or analyzed. For example, the user interface requires the user to grant permission to use any information associated with the user. The user is informed that user information can be deleted by the user, and the user can have the option to choose what types of information are provided for different purposes. The use of information complies with applicable regulations, and the data is stored securely. Data collection is not performed at specific locations and for specific user categories (e.g., based on age or other demographics), data collection is temporary (i.e., data is discarded after a period of time), and data is not shared with third parties. Some data may be anonymized, aggregated across users, or otherwise modified to make it impossible to determine the specific user's identity.

[0058] User interface module 202 generates graphical data for displaying the user interface, which is used to design the 3D avatar to be used in the virtual experience. The user interface includes options for selecting the type of character (e.g., human, humanoid, non-human, etc.) and selecting character attributes (including gender, height, body shape, clothing, etc.). The user interface includes options for adding text to the 3D avatar by adding text to clothing, skin, designing tattoos containing text for the 3D avatar, etc. The user interface includes options for rotating the 3D avatar and placing text or other attributes on different parts of the avatar's body.

[0059] In some embodiments, once a user has completed the design of their 3D avatar, the user can finalize the 3D avatar and request that it be saved in their associated user account. In some embodiments, the user interface module 202 generates 2D images of the 3D avatar from different angles surrounding it. In the example described below, the user interface module 202 generates eight 2D images of the 3D avatar, but any number of 2D images can be generated. The text recognition engine 204 determines whether the 3D avatar includes text, and if so, extracts the text. The machine learning module 206 determines whether the extracted text and / or other attributes (e.g., shape and clothing) are abusive. In some embodiments, the abusiveness of the 3D avatar's attributes is analyzed in a two-stage process, where the 3D avatar's body is analyzed first, followed by any clothing created for the 3D avatar.

[0060] In response to a determination that any of these attributes is abusive, the user interface module 202 provides a notification to the user. The notification includes that the attribute is abusive and may include an identifier of the abuse category (e.g., attributes associated with hate groups, nudity, innuendo; text that is profane, racist, sexist, bullying, etc.), the severity of the abuse (e.g., mild, moderate, severe), and / or the reason for the abuse (e.g., the 3D avatar has inappropriate exposure). In some embodiments, the notification may include a warning. In some embodiments, the abuse module 208 may apply disposition based on the severity of the abuse, such as if the user is associated with an abuse report or a history of abusive behavior.

[0061] In some embodiments, the notification includes an option to appeal the determination of abuse. The user interface module 202 can provide a 3D avatar of a 2D image to one or more reviewers who can confirm or reject the determination of abuse.

[0062] Users can generate an updated 3D avatar and resubmit it for approval. If the updated 3D avatar has acceptable attributes, the user interface module 202 can notify the user that the updated 3D avatar can be saved to the user's associated user account and used in the virtual experience 105.

[0063] The user interface module 202 receives user input from the user during gameplay in the virtual experience 105. For example, user input can cause the avatar to move around, perform actions, change postures, and talk to other users in the virtual experience. The user interface module 202 generates graphical data to display the avatar's position, actions, postures, etc., within the virtual experience 105.

[0064] Users can interact with other users in a virtual experience. Some of these interactions may be negative, and in some embodiments, the user interface module 202 generates graphical data for the user interface that allows the user to restrict exposure to other users they want to avoid. For example, the user interface module 202 may include options to mute other users, block other users, and report abuse occurring in the virtual experience. For example, another avatar may wear offensive clothing (e.g., a T-shirt or hat), an avatar may hold inappropriate objects (e.g., a flag associated with a hate group, an object with an offensive shape, etc.), an avatar may perform offensive actions (e.g., an avatar may use spray paint to create inappropriately revealing images), or an avatar may utter inappropriate phrases (e.g., in a chat box or directly via voice chat with the user). An avatar may also be associated with multiple types of abuse, such as wearing inappropriate clothing while performing offensive actions.

[0065] The text recognition engine 204 detects text on one or more 2D images generated from the 3D avatar by the user interface module 202. The text can be digital text, handwritten text, containing different fonts, styles, strokes, colors, etc., and can be in any language. In response to the text detection, the text recognition engine 204 extracts the text. In some embodiments, the text recognition engine 204 uses optical character recognition (OCR) to classify the text.

[0066] Machine learning module 206 trains a machine learning model to output a determination of whether at least one attribute of the 3D avatar is abusive. In some embodiments, machine learning module 206 includes different machine learning models trained for different functions. For example, Figure 2B This is a block diagram 250 of an example machine learning module 206, which includes three different machine learning models: a text machine learning model 255 for determining whether text is abusive, a shape machine learning model 260 for determining whether a shape is abusive, and a clothing machine learning model 265 for determining whether clothing is abusive. In some embodiments, the machine learning module 206 uses a cross-entropy loss function to train the text machine learning model 255, the shape machine learning model 260, and / or the clothing machine learning model 265. In some embodiments, 3D avatars are filtered to remove duplicates (or similar types of avatars) before training the machine learning models for 3D avatars, to avoid overtraining the machine learning models for specific types of 3D avatars.

[0067] In some embodiments, machine learning module 206 trains text machine learning model 255 to detect and classify offensive keywords and / or offensive phrases. Machine learning module 206 uses a training dataset to train the text machine learning model. In some embodiments, the training dataset includes both abusive and non-abusive words. In some embodiments, the training dataset also includes baseline truth label data. Baseline truth label data can be generated when an auditor reviews the labels assigned by text machine learning model 255 and confirms or modifies the labels. Labels can include an identifier of whether the text is abusive or non-abusive, and the category of abuse, such as profanity, racism, sexism, bullying, etc. In some embodiments, baseline truth labels can also be derived from abuse reports, where a user complains about another user's behavior, and an auditor verifies that the text is abusive and applies labels to the text submitted with the abuse report.

[0068] In some embodiments, the machine learning module 206 generates synthetic datasets to supplement the training dataset. For example, the machine learning module 206 may identify abuse categories of text in cases of insufficient examples and / or unreliable results provided by the text machine learning model. In some embodiments, the machine learning module 206 generates synthetic data for determining whether an attribute is abusive and associating it with confidence scores below a threshold confidence value.

[0069] Go to Figure 3 Example 300 illustrates a process for creating synthetic training data according to some embodiments described herein. Machine learning module 206 selects a 3D avatar 304. In some embodiments, the 3D avatar 304 is identified in the virtual experience. The 3D avatar 304 can be a human-like, realistic cartoon character, abstract figure, etc. The 3D avatar 304 can be unclothed, clothed (e.g., shirt, trousers, etc.), wearing accessories (e.g., hat), or have tattoos. Machine learning module 206 identifies the 3D mesh of the 3D avatar 304 in the virtual experience and separates the 3D avatar 304 from the surrounding scene.

[0070] Machine learning module 206 receives a set of images containing representative text samples. The text can be handwritten; numerical; contain different numbers of fonts, styles, strokes, colors; include different languages, etc. Text machine learning model 255 can be trained on abusive text (e.g., inappropriate word use) and non-abuseful text (e.g., chill, hello, world).

[0071] Machine learning module 206 utilizes depth information available in the 3D environment to isolate the text from the background. The isolated text is then projected onto the top of the 3D avatar 304 at different angles and orientations using a virtual camera system. For example, the 3D avatar 304 is combined with a projected image 310. The 3D avatar with the projected image 310 is then captured at different virtual camera positions 315 to generate multiple 2D images 305a-305h of the 3D avatar 304 including the projected image. In some embodiments, machine learning module 206 identifies specific distances and camera angles such that the text is consistently projected onto the 3D avatar 304. In some embodiments, instead of capturing discrete images at different angles, the camera system captures video of the 3D avatar 304 with the projected text by moving around the 3D avatar 304. 2D images 305 are created from various types of avatars to synthesize a diverse set of examples of text that differ depending on whether the avatar is tall or short, smooth or wavy, straight or spherical, etc.

[0072] Go to Figure 4This document illustrates an example of a process 400 for creating a 2D image 402 of a 3D avatar 403 according to some embodiments described herein. The input 3D avatar 403 can be a human, a humanoid, or a non-human, and can take the form of a realistic cartoon character, an abstract figure, etc. The 3D avatar 403 can be unclothed or clothed (e.g., a shirt, pants, a hat, etc.).

[0073] In this example, multiple high-resolution images of the 3D avatar 403 are captured from different angles, for example, using virtual camera systems 401a-401h. In some embodiments, the virtual camera system 401 is positioned in a virtual 3D space to capture the 3D avatar 403 from a specific viewpoint. The virtual camera system 401 adjusts camera parameters such as field of view, focal length, and depth of field to control the final composition. The virtual camera system 401 generates multiple 2D images 402a-402h of the 3D avatar. Each 2D image 402 is captured by the virtual camera system 401 from a different angle or orientation. The 2D images 402 are preprocessed to enhance quality and make them suitable for text detection, such as by adjusting the size of the 2D images 402, cropping them, and adjusting their brightness and contrast. In some embodiments, the 2D images 402 are saved as high-resolution thumbnails.

[0074] In some embodiments, the machine learning module 206 employs Figure 4 The process 400 generates synthetic training data for different types of attributes. For example, the machine learning module 206 can generate synthetic training data for different categories of abusive clothing and different categories of abusive shapes for 3D avatars. In some embodiments, the user manually creates different 3D avatars and uses... Figure 4 The process 400 is used to generate synthetic training data. For example, a user can manually create 3D avatars with exaggerated shapes for different examples of inappropriate body shapes. In another example, a user can create 3D avatars with different amounts of clothing (also known as modesty layers) for different body shapes (e.g., modesty layers that are insufficient to cover a particular body part are flagged as abusive for violating nudity standards).

[0075] Go to Figure 5 An example process 500 for classifying text on an avatar is illustrated. In some embodiments, the user interface 202 generates 2D thumbnail images 510 of the 3D avatar 505. For example, the user interface module 202 may include a camera system for capturing multiple 2D thumbnail images 510 of the 3D avatar 505. The 2D thumbnail images 510 of the 3D avatar 505 are captured from multiple angles and different orientations surrounding the 3D avatar using a virtual camera system.

[0076] The text recognition engine 204 receives a 2D thumbnail image 510 and performs text detection 515. The text recognition engine 204 identifies regions in the image that may contain text. In some embodiments, the text recognition engine 204 uses OCR to detect the presence of text.

[0077] The text recognition engine 204 is trained on a dataset including images with annotated text regions to help the model learn to recognize text in various orientations and angles. The text recognition engine 204 is language-independent and applicable to text in any language.

[0078] Text recognition engine 204 performs text extraction 520. For example, if the phrase “I love ZZZ” is located on the skin of 3D avatar 505, where “ZZZ” is replaced with a specific inappropriate word, then 2D thumbnail image 510 may include a first 2D thumbnail image 510 of the avatar at a frontal angle where the text appears slightly distorted, a second 2D thumbnail image 510 of the avatar directly facing the camera, and a third 2D thumbnail image 510 of the avatar at a relative frontal angle where the text appears slightly distorted (i.e., rotated 45 degrees from the first 2D thumbnail image 510). Text recognition engine 204 can extract text by cropping a portion of the image from the three different 2D thumbnail images 510, including the phrase “I love ZZZ” (where “ZZZ” refers to a specific word with sexual connotations).

[0079] The text recognition engine 204 provides the extracted text to the text machine learning model 255. The text machine learning model 255 receives the extracted text as input and classifies the extracted text. The text machine learning model 255 can output a determination of whether the extracted text is abusive, and if so, output the abuse category and / or the severity of the abuse.

[0080] The text machine learning model 255 can use OCR to perform classification. In some embodiments, the text machine learning model 255 is a deep learning machine learning model trained using a machine learning model to detect abusive text. Types of deep neural networks include CNNs, deep belief networks, transformer models, generative adversarial networks, bidirectional encoder representations from transformers (BERT), stream models, recurrent neural networks, and Universal Language Model Fine-tuning (ULMFiT). Deep neural networks use multiple layers to progressively extract higher-level features from the raw input, where the input to a layer is different types of features extracted from other modules, and the output is a determination of whether the text contains abusive content.

[0081] Text machine learning model 255 converts the extracted text into a machine-readable format. Text machine learning model 255 performs text classification 525 by classifying the extracted text into predefined categories or labels that predict violation categories (if any). Text machine learning model 255 performs text-based analysis to identify instances of abuse or violation. In some embodiments, text classification module 375 utilizes OCR to train or fine-tune a dataset including examples of abusive and non-abuse text.

[0082] In some embodiments, the text machine learning model 255 uses a text classification system to generate labels for the identified text. For example, “Hello” is associated with a non-abuse label, and “XXX” (where “XXX” refers to a specific word with a profanity connotation) is associated with an abuse label. In some embodiments, abuse labels can be further refined into categories such as bullying, sex, and profanity labels. Other labels are also possible. For example, abuse labels can be expanded to include bullying and harassment, real-world dangerous activities, discrimination and hate, extortion and blackmail, sexual content, violent and gory content, threats of violence, illegal and regulated content, dating and romance, profanity, spam, political content, misleading impersonation or false statements, deception and exploitation, etc. For example, “Dog,” “World,” and “Chill” are associated with non-abuse labels; “YYY” (where “YYY” refers to a specific word with a sexual connotation) is associated with a sexual abuse label. Labels are added to the training dataset. Labels predict the violation category (if any). For example, “XXX” (where “XXX” refers to a specific word with a profane connotation) is associated with the abusive label of blasphemy and falls under the category of violation.

[0083] Figure 6 Example 600 illustrates a 3D avatar 604 with detected text 605. Although illustrated as “XXX”, the text in this example refers to a specific word with profanity. Text recognition engine 204 identifies the 3D avatar 604 as containing text 605 and extracts text 605 from the avatar. The extracted text is fed to text machine learning model 255, which outputs a determination of abuse.

[0084] In this example, the text machine learning model 255 determines that the extracted text “XXX” (where “XXX” refers to a specific word with profanity) printed on the avatar's t-shirt is abusive 601 and associated with “blasphemy” as abusive category 602. The text machine learning model 255 also determines that the severity of the abusive text “XXX” (where “XXX” refers to a specific word with profanity) is high 603.

[0085] In some embodiments, the machine learning module 206 receives 2D images of the 3D avatar and provides them as input to the shape machine learning model 260 and / or the clothing machine learning model 265. In some embodiments, the clothing machine learning model 265 receives the 2D images if the avatar is wearing clothing, otherwise it does not. In some embodiments, a user can create a 3D avatar approved by the shape machine learning model 260, and later the user creates clothing for the 3D avatar, which is then provided to the clothing machine learning model 265 for approval.

[0086] Machine learning module 206 trains shape machine learning model 260 to output a determination of whether an avatar's shape is abusive. In some embodiments, the criteria used to determine abusive shapes are policy-based. For example, a policy may include specific examples of prohibited shapes, such as including inappropriate body shapes, shapes of non-humanoid creatures designed to harass other players (e.g., characters known to make other players uncomfortable or associated with hate groups), etc.

[0087] In some embodiments, the shape machine learning model 260 is trained using examples of incarnations of shapes that are considered abusive and are labeled as abusive (including abusive categories). In some embodiments, the shape machine learning model 260 is periodically updated with new examples of abusive shapes to reflect new trends and / or more subtle differences in shapes. For example, the shape machine learning model 260 may receive training data for fine-tuning, which is updated using 2D images identified by human reviewers as incarnations of abusive shapes, or where a determination made by the shape machine learning model 260 is rejected and determined not to be an abusive shape.

[0088] Machine learning module 206 trains clothing machine learning model 265 to output a determination that an avatar's clothing is abusive. In some embodiments, the criteria used to determine abusive clothing are policy-based. For example, the policy may include specific examples of clothing associated with known hate groups, clothing that includes inappropriate images, etc. In some embodiments, clothing machine learning model 265 is trained using examples of avatars that are considered abusive and are labeled as abusive (including abusive categories).

[0089] In some embodiments, the clothing machine learning model 265 is trained using examples of avatars of clothing that are considered abusive and are labeled as abusive (including abusive categories). For example, clothing may be associated with hate groups, clothing may be overly revealing, clothing may include problematic symbols, etc. In some embodiments, the clothing machine learning model 265 is periodically updated with new examples of abusive clothing to reflect new trends and / or more subtle distinctions (e.g., the difference between a regular t-shirt and a racist t-shirt). For example, the shape machine learning model 260 may receive training data updated with images of avatars identified by human reviewers as wearing abusive clothing, or where a determination made by the clothing machine learning model 265 is rejected and determined not to be abusive clothing.

[0090] In some embodiments, shape machine learning model 260 and / or clothing machine learning model 265 use a multi-view fusion model. Go to Figure 7 An example of a multi-view fusion model 700 is shown. The multi-view fusion model 700 includes multiple machine learning models. Specifically, the multi-view fusion model 700 includes a CNN 701 and a multilayer perceptron (MLP) 704.

[0091] CNN 701 receives multiple 2D images 402 of the 3D avatar from different angles surrounding the 3D avatar (e.g., reference). Figure 4 The 2D image 402 is described as a high-resolution thumbnail of a 3D avatar captured using a camera system. In some embodiments, the 2D image 402 is preprocessed to enhance image quality. In some embodiments, during the training of CNN 701, the 2D image 402 may be shuffled and enhanced with various random color jitters to train CNN 701 against greater variations in the 2D image 402.

[0092] CNN 701 performs image classification. In some embodiments, CNN 701 is a residual network with convolutional layers and batch normalization layers. In some embodiments, CNN 701 includes a series of convolutional layers, batch normalization, activation functions, and residual blocks. CNN 701 may include one or more convolutional layers that reduce the number of channels in a 2D image 402 for faster processing, one or more convolutional layers that extract spatial details from the reduced 2D image 402, and one or more convolutional layers that restore the reduced 2D image 402 to its original number of channels. In some embodiments, CNN 701 includes skip connections that add the unchanged input directly to the output of the convolutional layers, which advantageously preserves information from earlier convolutional layers. Convolutional layer output feature maps.

[0093] In some embodiments, the CNN 701 includes a pooling layer that slides a two-dimensional filter across each channel of the feature map and aggregates features located within the regions covered by the filter. The pooling layer generates multiple view embeddings 702 for the 2D image 402, with the CNN 701 outputting a view embedding 702 for each 2D image 402. The view embeddings 702 provide a compact representation of the 2D image, which can be used for comparison, classification, or text detection across different views or orientations. Multiple view embeddings 702 are concatenated to form a larger multi-view embedding 703. The concatenated embedding 703 is provided as input to the MLP 704.

[0094] MLP 704 comprises multiple connected neural network layers (e.g., three neural network layers). MLP 704 processes the splicing embedding 703 through a nonlinear transformation to extract patterns and features. MLP 704 analyzes attributes associated with the 3D avatar, as described by the splicing embedding. For example, MLP 704 can be trained to identify abusive shapes or abusive clothing.

[0095] In some embodiments, MLP 704 outputs classification results for different categories. For example, MLP 704 may output classifications of attributes that match offensive clothing category 705, nudity category 706, suggestive category 707, or no violation determination 708. In some embodiments, MLP 704 also outputs the severity of abuse. For example, the shape of a persona with a more exaggerated inappropriate body shape may be classified as more severe than a persona with a lesser degree of inappropriate body shape. In some embodiments, each determination of an abusive attribute may be associated with a confidence score, where the confidence score reflects the level of confidence in the determination of the abusive attribute.

[0096] The abuse module 208 performs a disposition action based on the determination of abuse. For example, the disposition action may be a warning when abuse is first determined, and a prohibition, such as prohibiting the customization of an avatar associated with the user who committed the abuse, when abuse is determined a second time. In some embodiments, the abuse module 208 determines the disposition action based on the severity of the abuse and / or based on the user's past violation history. In some embodiments, the determination of abuse is associated with a confidence score, and the customization of the user's avatar is not prohibited unless the confidence score meets a threshold confidence value.

[0097] In some embodiments, the abuse module 208 applies different groups of rules based on the age of the first player. In some embodiments, the rules are different for the following groups: 13-16 years old, 16-18 years old, or 18 years and older. For example, if the first user is 18 years or older, the consequences of abuse committed by the first user may escalate faster than if the first user is a minor (i.e., under 18 years old).

[0098] In some embodiments, the abuse module 208 may provide a warning to a first user when it first determines that the user is committing abuse before applying the disposition action.

[0099] Example User Interface Figure 8 This includes a sample user interface 800 that warns users that the 3D avatar violates community standards. User interface 800 includes reasons 802 for the avatar being abusive (i.e., because a specific body part contains blasphemy and nudity) and the potential consequences of continued violations. The community standards text consists of links to the community standards. Clicking the links may take users to different pages that contain all the community standards.

[0100] The user interface 800 also includes a confirmation button 805 that instructs the first user to confirm the warning. If the user disagrees with the warning, the user can click an objection button 810 with the text "Was the judgment incorrect? Please let us know." In some embodiments, the abuse module 208 tracks the number of times the user selects the confirmation button 805 and the objection button 810.

[0101] Actions can take several forms and are based on whether the user is associated with a previous action. In some embodiments, the abuse module 208 imposes a temporary ban on initial violations and a permanent ban on more serious and / or repeated violations. Bans prevent users from participating in virtual experiences for a period of time. For example, a player ban could include disabling login credentials.

[0102] Example user interface 800 can be displayed in response to user creation of avatars including abusive attributes, in response to player submission of an abuse report that triggers confirmation of user creation of abusive avatars, in response to avatars being randomly identified by reviewers, etc.

[0103] Example Method Figure 9 This is a flowchart of an example method 900 for generating synthetic training data for a machine learning model, according to some embodiments described herein. Method 900 can be... Figure 2A The computing device 200 in the middle executes.

[0104] Method 900 can begin with box 902. At box 902, the abuse category is identified for the defined text relating the attribute to a confidence score below a threshold confidence value. Categories may include profanity, racism, sexism, or bullying. Box 902 can be followed by box 904.

[0105] At box 904, generate abusive text associated with the identified category. Box 904 may be followed by box 906.

[0106] At box 906, a training 3D avatar is generated. In some embodiments, the training 3D avatar is generated by identifying the 3D avatar in the virtual experience, a 3D mesh identifying the 3D avatar in the virtual experience, and separating the 3D avatar from the surrounding scene. Box 906 may be followed by box 908.

[0107] At box 908, the abusive text is projected onto the training 3D model. Box 908 can be followed by box 910.

[0108] At box 910, a training 2D image is generated from the training 3D avatar having the projected abusive text. In some embodiments, method 900 further includes shuffling the training 2D image and using random color jitter to enhance the training 2D image. Box 910 may be followed by box 912.

[0109] At box 912, labels including the abuse category will be applied to the training 2D image. Box 912 may be followed by box 914.

[0110] At box 914, a machine learning model is trained to minimize the difference between the misuse category predicted for the training 2D image and the label for the training image.

[0111] Figure 10 This is a flowchart of an example method 1000 for training a machine learning model to identify abusive embodiments, according to some embodiments described herein. Method 1000 can be... Figure 2A The computing device 200 in the middle executes.

[0112] Method 1000 can begin at box 1002. At box 1002, the 3D avatar is received. Box 1004 can follow box 1002.

[0113] At box 1004, 2D images of the 3D avatar are generated from different angles surrounding the 3D avatar. Box 1004 can be followed by box 1006.

[0114] At box 1006, determine whether one or more of the 2D images contain text. If one or more of the 2D images contain text, then box 1008 may follow box 1006. If none of the 2D images contain text, then box 1010 may follow box 1006.

[0115] At box 1008, extract text from one or more 2D images. Box 1008 may be followed by box 1014.

[0116] At box 1010, a 2D image is fed as input to the trained machine learning model. Box 1010 may be followed by box 1012.

[0117] At box 1012, the machine learning model generates a stitched embedding of the 2D image. Box 1012 can be followed by box 1014.

[0118] At box 1014, the machine learning model analyzes at least one attribute associated with the 3D avatar, wherein the attribute is selected from text, the shape of the 3D avatar, the clothing of the 3D avatar, and groups of combinations thereof. Box 1014 may be followed by box 1016.

[0119] At box 1016, the machine learning model outputs a determination that at least one property of the 3D avatar is abusive.

[0120] Figure 11 This is a flowchart of an example method 1100 for auditing avatars, based on some embodiments described herein. Method 1100 can be... Figure 2A The computing device 200 in the middle executes.

[0121] At box 1102, the 3D avatar is received. Box 1104 can follow box 1102.

[0122] At box 1104, 2D images of the 3D avatar are generated from different angles surrounding the 3D avatar. In some embodiments, generating 2D images of the 3D avatar includes using a virtual camera system to capture high-resolution thumbnails. Box 1104 may be followed by box 1106.

[0123] At box 1106, a 2D image is fed as input to the trained machine learning model. Box 1106 may be followed by box 1108.

[0124] At box 1108, the machine learning model generates a stitched embedding of the 2D images. In some embodiments, the machine learning model includes a CNN that generates the stitched embedding by: extracting the corresponding view embedding of each 2D image from the 2D images using pooling layers; and stitching the view embeddings to form the stitched embedding. Box 1108 may be followed by box 1110.

[0125] At box 1110, the machine learning model analyzes at least one attribute associated with the 3D avatar based on splicing embeddings, wherein the at least one attribute is selected from the shape of the 3D avatar, the clothing on the 3D avatar, and combinations thereof. Box 1110 may be followed by box 1112.

[0126] At box 1112, the machine learning model outputs a determination that at least one attribute of the 3D avatar is abusive. In some embodiments, the 3D avatar is generated by the user, and method 1100 further includes providing the user with notification, in response to the output determination, of one or more of the following: attributes of the 3D avatar being abusive, an identifier of the abuse category, the severity of the abuse of the attribute, the reason why the attribute is abusive, and combinations thereof. In some embodiments, method 1100 further includes: receiving an updated 3D avatar from the user; generating an updated 2D image of the updated 3D avatar from different angles surrounding the updated 3D avatar; providing the updated 2D image as input to the machine learning model; utilizing the machine learning model's output determination that the attribute is acceptable; and providing the user with the option to use the updated 3D avatar in a virtual environment. In some embodiments, the machine learning model is trained using training data, and method 1100 further includes: in response to the updated 3D avatar being used in a virtual environment, receiving an abuse report from a player in the virtual environment describing the updated 3D avatar as having abusive attributes; determining that the updated 3D avatar has abusive attributes; and updating the training data associated with the machine learning model to include the updated 2D image and a label with an abuse category.

[0127] In some embodiments, method 1100 further includes: detecting text on one or more 2D images in a 3D avatar; extracting text from the one or more 2D images; and classifying the extracted text to determine whether the extracted text is abusive, wherein the determination of whether an attribute is abusive is based on classifying the extracted text as abusive. In some embodiments, the machine learning model is trained using training data, and method 1100 further includes updating the training data associated with the machine learning model to include training images that have recently been identified as abusive.

[0128] Where appropriate, the methods, blocks, and / or operations described herein may be performed in an order different from that shown or described, and / or simultaneously (partially or completely simultaneously) with other blocks or operations. Some blocks or operations may be performed on a portion of the data, and some blocks or operations may be performed again later, for example, on another portion of the data. Not all described blocks and operations need to be performed in the various embodiments. In some embodiments, these blocks and operations may be performed multiple times in a method, in different orders, and / or at different times.

[0129] The various embodiments described herein include acquiring data from various sensors in the physical environment, analyzing such data, generating recommendations, and providing a user interface. Data collection is performed only with specific user permission and in accordance with applicable regulations. Data storage complies with applicable regulations, including anonymizing or otherwise modifying data to protect user privacy. Explicit information is provided to users regarding data collection, storage, and use, and users are given the option to select the types of data that can be collected, stored, and utilized. Furthermore, users control the devices that can store data (e.g., client-only devices; client + server devices, etc.) and the devices that perform data analysis (e.g., client-only devices; client + server devices, etc.). Data is used for the specific purposes described herein. Data is not shared with third parties without explicit user permission.

[0130] In the foregoing description, numerous specific details have been set forth for purposes of explanation in order to provide a thorough understanding of this specification. However, it will be apparent to those skilled in the art that this disclosure may be practiced without these specific details. In some instances, structures and devices have been shown in block diagram form to avoid obscuring the description. For example, these embodiments have been described above primarily with reference to user interfaces and specific hardware. However, these embodiments are applicable to any type of computing device capable of receiving data and commands, as well as any peripheral devices providing services.

[0131] References to "some embodiments" or "some examples" in the specification mean that a particular feature, structure, or characteristic described in connection with an embodiment or example may be included in at least one embodiment of the specification. The phrase "some embodiments" appearing throughout the specification does not necessarily refer to the same embodiment.

[0132] Some parts of the detailed description above are presented based on algorithms and symbolic representations of operations on data bits within computer memory. These algorithmic descriptions and representations are the most effective means for those skilled in the art of data processing to communicate the essence of their work to others skilled in the art. An algorithm here, and generally in the sense of the general term, is considered a self-consistent sequence of steps leading to a desired result. These steps are those that require physical manipulation of physical quantities. Typically, but not necessarily, these quantities take the form of electrical or magnetic data that can be stored, transferred, combined, compared, and otherwise manipulated. Sometimes, primarily for reasons of common use, it has proven convenient to refer to these data as bits, values, elements, symbols, characters, terms, numbers, etc.

[0133] However, it should be remembered that all these and similar terms are associated with appropriate physical quantities and are merely convenient labels applied to those quantities. As will be apparent from the following discussion, unless otherwise explicitly stated, it should be understood that throughout the description, the use of terms such as “processing” or “calculating” or “estimating” or “determining” or “displaying” refers to the actions and processes of a computer system or similar electronic computing device that manipulate and convert data represented as physical (electronic) quantities within the registers and memories of the computer system into other data similarly represented as physical quantities within the computer system's memory or registers or other such information storage, transmission, or display devices.

[0134] Embodiments of this specification may also relate to a processor for performing one or more steps of the methods described above. The processor may be a dedicated processor that is selectively activated or reconfigured by a computer program stored in a computer. Such a computer program may be stored in a non-transitory computer-readable storage medium, including but not limited to any type of disk, including optical discs, ROMs, CD-ROMs, magnetic disks, RAM, EPROMs, EEPROMs, magnetic cards or optical cards, flash memory including a USB key with non-volatile memory, or any type of medium suitable for storing electronic instructions, each coupled to a computer system bus.

[0135] This specification may take the form of entirely hardware embodiments, entirely software embodiments, or embodiments that include both hardware and software elements. In some embodiments, this specification is implemented in software, including but not limited to firmware, resident software, microcode, etc.

[0136] Furthermore, the description may take the form of a computer program product accessible from a computer-usable or computer-readable medium that provides program code for use by or in connection with a computer or any instruction execution system. For the purposes of this description, a computer-usable or computer-readable medium can be any means that can contain, store, communicate, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.

[0137] A data processing system suitable for storing or executing program code will include at least one processor directly or indirectly coupled to memory elements via a system bus. Memory elements may include local memory used during the actual execution of the program code, mass storage devices, and cache memory that provides temporary storage for at least some of the program code to reduce the number of times code must be retrieved from mass storage devices during execution.

Claims

1. A computer-implemented method, the computer-implemented method comprising: Receive three-dimensional (3D) avatars; Generate two-dimensional (2D) images of the 3D avatar from different angles surrounding the 3D avatar; The 2D image is used as input to the trained machine learning model; The machine learning model is used to generate the stitched embedding of the 2D image; The machine learning model is used to analyze at least one attribute associated with the 3D avatar based on the stitched embedding, wherein the at least one attribute is selected from the shape of the 3D avatar, the clothing on the 3D avatar, and combinations thereof; and The machine learning model outputs a determination that at least one attribute of the 3D avatar is abusive.

2. The method according to claim 1, further comprising: Detect text on one or more 2D images in the 2D images of the 3D avatar; Extract the text from one or more 2D images in the 2D images; as well as The extracted text is categorized to determine whether it is abusive, wherein the determination of whether the attribute is abusive is based on classifying the extracted text as abusive.

3. The method of claim 1, wherein the 3D avatar is generated by a user, and the method further comprises, in response to outputting the determination, providing the user with notification regarding the 3D avatar including the attribute being abusive and an identifier of the abuse category, the severity of the abuse of the attribute, the reason why the attribute is abusive, and one or more combinations thereof.

4. The method according to claim 3, further comprising: Receive the updated 3D avatar from the user; An updated 2D image of the updated 3D avatar is generated from different angles surrounding the updated 3D avatar; The updated 2D image is used as input to the machine learning model; The machine learning model outputs a determination that the attribute is acceptable; and The user is provided with the option to use the updated 3D avatar in a virtual environment.

5. The method of claim 4, wherein the machine learning model is trained using training data, and the method further comprises: In response to the updated 3D avatar being used in the virtual environment, an abuse report is received from a player in the virtual environment describing the updated 3D avatar as having abusive attributes; It has been determined that the updated 3D avatar possesses the aforementioned abusive properties; as well as The training data associated with the machine learning model is updated to include the updated 2D image and a label with the abuse category.

6. The method of claim 1, wherein the machine learning model comprises a convolutional neural network (CNN), the CNN generating the stitched embedding of the 2D image in the following manner: The pooling layer is used to extract the corresponding view embedding of each 2D image in the 2D image; and The view embeddings are stitched together to form the stitched embedding.

7. The method of claim 1, wherein the machine learning model is trained using synthetic training data for text, and the synthetic training data is generated in the following manner: The text is categorized as abused, wherein the determination that the attribute is abused is associated with a confidence score below a threshold confidence value; Generate abusive text associated with the identified category; Generate training 3D avatars; Project the abusive text onto the training 3D model; Generate training 2D images from the training 3D avatar with the projected abusive text; The labels, including the abuse category, are applied to the training 2D images; as well as The machine learning model is trained to minimize the difference between the predicted abuse category for the training 2D image and the label for the training 2D image.

8. The method of claim 7, wherein the synthetic training data is further generated by shuffling the training 2D image and enhancing the training 2D image using random color jitter.

9. The method of claim 1, wherein the machine learning model is trained using training data, and the method further comprises updating the training data associated with the machine learning model to include training images of attributes recently identified as abusive.

10. The method of claim 1, wherein generating the 2D image of the 3D avatar comprises capturing a high-resolution thumbnail using a virtual camera system.

11. A non-transitory computer-readable medium having instructions that, when executed by one or more processors at a client device, cause the one or more processors to perform operations, the operations including: Receive three-dimensional (3D) avatars; Generate two-dimensional (2D) images of the 3D avatar from different angles surrounding the 3D avatar; The 2D image is used as input to the trained machine learning model; The machine learning model is used to generate the stitched embedding of the 2D image; The machine learning model is used to analyze at least one attribute associated with the 3D avatar based on the stitched embedding, wherein the at least one attribute is selected from the shape of the 3D avatar, the clothing on the 3D avatar, and combinations thereof; and The machine learning model outputs a determination that at least one attribute of the 3D avatar is abusive.

12. The non-transitory computer-readable medium of claim 11, wherein the operation further comprises: Detect text on one or more 2D images in the 2D images of the 3D avatar; Extract the text from one or more 2D images in the 2D images; as well as The extracted text is categorized to determine whether it is abusive, wherein the determination of whether the attribute is abusive is based on classifying the extracted text as abusive.

13. The non-transitory computer-readable medium of claim 11, wherein the 3D avatar is generated by a user, and the operation further includes providing the user with notification regarding the 3D avatar in response to outputting the determination, including the attribute being abusive and an identifier of the abuse category, the severity of the abuse of the attribute, the reason why the attribute is abusive, and one or more combinations thereof.

14. The non-transitory computer-readable medium of claim 13, wherein the operation further comprises: Receive the updated 3D avatar from the user; An updated 2D image of the updated 3D avatar is generated from different angles surrounding the updated 3D avatar; The updated 2D image is used as input to the machine learning model; The machine learning model outputs a determination that the attribute is acceptable; and The user is provided with the option to use the updated 3D avatar in a virtual environment.

15. The non-transitory computer-readable medium of claim 14, wherein the machine learning model is trained using training data, and the operation further comprises: In response to the updated 3D avatar being used in the virtual environment, an abuse report is received from a player in the virtual environment describing the updated 3D avatar as having abusive attributes; It has been determined that the updated 3D avatar possesses the aforementioned abusive properties; as well as The training data associated with the machine learning model is updated to include the updated 2D image and a label with the abuse category.

16. A system comprising: processor; and A memory coupled to the processor, the memory storing instructions that, when executed by the processor, cause the processor to perform operations, the operations including: Receive three-dimensional (3D) avatars; Generate two-dimensional (2D) images of the 3D avatar from different angles surrounding the 3D avatar; The 2D image is used as input to the trained machine learning model; The machine learning model is used to generate the stitched embedding of the 2D image; The machine learning model is used to analyze at least one attribute associated with the 3D avatar based on the stitched embedding, wherein the at least one attribute is selected from the shape of the 3D avatar, the clothing on the 3D avatar, and combinations thereof; and The machine learning model outputs a determination that at least one attribute of the 3D avatar is abusive.

17. The system of claim 16, wherein the operation further comprises: Detect text on one or more 2D images in the 2D images of the 3D avatar; Extract the text from one or more 2D images in the 2D images; as well as The extracted text is categorized to determine whether it is abusive, wherein the determination of whether the attribute is abusive is based on classifying the extracted text as abusive.

18. The system of claim 16, wherein the 3D avatar is generated by a user, and the operation further includes providing the user with a notification in response to outputting the determination regarding the 3D avatar including the attribute being abusive and an identifier of the abuse category, the severity of the abuse of the attribute, the reason why the attribute is abusive, and one or more combinations thereof.

19. The system of claim 18, wherein the operation further comprises: Receive the updated 3D avatar from the user; An updated 2D image of the updated 3D avatar is generated from different angles surrounding the updated 3D avatar; The updated 2D image is used as input to the machine learning model; The machine learning model outputs a determination that the attribute is acceptable; and The user is provided with the option to use the updated 3D avatar in a virtual environment.

20. The system of claim 19, wherein the machine learning model is trained using training data, and the operation further includes: In response to the updated 3D avatar being used in the virtual environment, an abuse report is received from a player in the virtual environment describing the updated 3D avatar as having abusive attributes; It has been determined that the updated 3D avatar possesses the aforementioned abusive properties; as well as The training data associated with the machine learning model is updated to include the updated 2D image and a label with the abuse category.