System and method for enhancing the security of digital avatars
The system addresses the limitations of existing deepfake detection by integrating biometric verification and watermarking to secure avatar generation, ensuring authenticity and ethical use, thereby enhancing trust and safety in digital avatars.
Patent Information
- Application Number
- PCT/IL2025/050485
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-06
- Filing Date
- 2025-06-05
- Publication Date
- 2025-12-11
AI Technical Summary
Current solutions for detecting deepfake avatars lack comprehensive integration of security measures and rely heavily on post-facto detection methods, making them less effective against sophisticated fake content generation techniques.
A computer-based system and method for secure avatar generation that includes user enrollment through biometric verification using selfie videos, dynamic challenges, and identity logs, along with real-time moderation and watermarking to ensure authenticity and ethical use.
Enhances the security of digital avatars by preventing unauthorized generation and ensuring ethical use through robust real-time moderation and multi-layered security features, maintaining high standards of trust and safety in digital content.
Smart Images

Figure IL2025050485_11122025_PF_FP_ABST
Abstract
Description
SYSTEM AND METHOD FOR ENHANCING THE SECURITY OF DIGITALAVATARSFIELD OF THE INVENTION
[0001] Embodiments of the present invention relate generally to the field of computerized avatar generation. Some embodiments of the present invention relate to enhancing the security of digital avatars of persons.BACKGROUND
[0001] In computing, an avatar may relate to a digital or computerized graphical representation of a person or a character. Avatars may be or may include two dimensional (2D) or three dimensional (3D) models of the face and / or body of a person (or character) they represent. Avatars may be used in virtual worlds (e.g., computer-simulated environments), video games, in generating video clips (e.g., animations), etc.
[0002] An avatar of a person may be created from multiple images and / or videos depicting the person, using an artificial intelligence (Al) process that employs one or more machine learning (ML) models. Once the avatar is generated, a new video or animation can be produced by applying ML models that animate movements and speech for the avatar. For example, in order to generate a video, the ML model may utilize the avatar, a voice recording (which could be of the represented person, another individual, or a synthetic voice), and text that the avatar is intended to speak. Generative ML models can then produce the audio for the video, often using text-to-speech technology, and generate lip movements synchronized with the speech. As a result, the final video may depict the avatar speaking the provided text, with the audio track including the spoken text. In some cases, the voice may be synthesized to match a specific person's voice, either the person modeled by the avatar or someone else, based on the voice sample provided to the ML model. Depending on the sophistication of the ML models used, additional gestures and movements of the avatar can also be generated, as known in the art.
[0003] Avatars can be either photorealistic or stylized. As technology advances, ALgenerated video clips featuring photorealistic avatars are likely to appear more natural, realistic, and authentic. Consequently, it may become increasingly difficult for both humans and machines todistinguish between genuine video footage of a person and deepfake videos created using that person's avatar.
[0004] While avatars are often generated for legitimate purposes such as video games, social media, entertainment, for generating training material, etc., the same technology may be used for illegitimate purposes such as imposing, performing privacy invasions, providing misinformation, performing digital impersonation, and for performing other types of fraud. Current solutions for detecting deepfake exist; however, those solutions lack comprehensive integration of security measures or rely heavily on post-facto detection methods which may be less effective against new and sophisticated fake content generation techniques.SUMMARY
[0005] According to some embodiments, a computer-based system and method for secure generation of avatars may include any combination of and in any order: requesting a user to record a selfie video with a dynamic challenge; verifying the selfie video against the dynamic challenge; enrolling the user on the computerized avatar generation platform and generating an identity log of the user if the verification of the selfie video is successful, and refusing to enroll the user on the computerized avatar generation platform otherwise, wherein the identity log comprises biometric identification data e.g. image(s) and audio templates / samples, extracted from the selfie video of the enrolled user; receiving a request from a requester to generate an avatar of the requester, the request may include biometric identification data of the requester; verifying the biometric identification data of the requester against the biometric identification data in the identity log of the enrolled user; and generating an avatar for the requester if the verification of the identity of the requester against the identity log of the enrolled user is successful, and denying the request otherwise.
[0006] According to some embodiments, enrolling the user on the computerized avatar generation platform may include any combination of and in any order: authenticating the user by performing at least one of: a shared authentication scheme, email verification, two-factor authentication (2FA), multiple factor authentication (MFA), automated know your customer (KYC) verification and credit card verification; and receiving user consent.
[0007] According to some embodiments, the dynamic challenge may include at least one of: text that the user is requested to say and an action that the user is requested to perform.
[0008] According to some embodiments, the biometric identification data in the identity log may include at least one of: data indicative of the facial image of the enrolled user extracted from the selfie video and data indicative of the voice of the enrolled user extracted from the selfie video.
[0009] According to some embodiments, verifying the biometric identification data of the requester against the biometric identification data in the identity log of the enrolled user may include verifying a selfie of the requester against the data indicative of the facial image of the enrolled user.
[0010] According to some embodiments, verifying the selfie of the requester against the data indicative of the facial image of the enrolled user may be performed using an ML model.[Oi l] According to some embodiments, verifying the biometric identification data of the requester against the biometric identification data in the identity log of the enrolled user may include verifying a voice recording of the requester against data indicative of the voice of the enrolled user.
[0012] According to some embodiments, verifying a voice recording of the requester against data indicative of the voice of the enrolled user may be performed using an ML model.
[0013] Some embodiments may include generating a video comprising the avatar; and inserting a watermark into the generated video, the watermark comprising data identifying the requester.
[0014] Some embodiments may include receiving a complaint associated with an examined video presenting an examined avatar; analyzing content of the examined video to determine whether the content of examined video is legitimate or illegitimate; in case of illegitimate content, extracting from the examined video data identifying a user that generated the examined video; and banning the user that generated the examined video.
[0015] Some embodiments may include generating a video comprising the avatar; and inserting a watermark into the generated video, the watermark comprises data identifying the computerized avatar generation platform.
[0016] Some embodiments may include receiving a complaint associated with an examined video presenting an examined avatar; determining whether the examined video was generated by the computerized avatar generation platform based on the watermark; and handling the complaint only if the examined video was generated by the computerized avatar generation platform.
[0017] Some embodiments may include generating a video comprising the avatar; generating video identifying data from the generated video; and storing the video identifying data in the computerized avatar generation platform.
[0018] According to some embodiments, the video identifying data may include at least one of: a video identification number embedded in a watermark inserted into the generated video, a hash of the generated video and an embedding of the video.
[0019] Some embodiments may include receiving a complaint associated with an examined video presenting an examined avatar; determining whether the examined video was generated by the computerized avatar generation platform based on the video identifying data; and handling the complaint only if the examined video was generated by the computerized avatar generation platform.
[0020] According to some embodiments, a computer-based system and method for secure generation of avatars, may include: enrolling a user by: providing a dynamic challenge to the user;
[0021] requesting the user to record a selfie video with the dynamic challenge; verifying the selfie video against the dynamic challenge; and enrolling the user and generating an identity log of the user if the verification of the selfie video is successful, and refusing to enroll the user if the verification of the selfie video fails; and receiving a request from a requester to generate an avatar of the requester; verifying the identity of the requester against an identity log of the enrolled user; generating an avatar for the requester if the verification of the identity of the requester against the identity log of the enrolled user is successful, and denying the request if the verification of the identity of the requester against the identity log of the enrolled user fails.BRIEF DESCRIPTION OF THE DRAWINGS
[0022] For a better understanding of embodiments of the disclosure and to show how the same can be carried into effect, reference will now be made, purely by way of example, to the accompanying drawings in which like numerals designate corresponding elements or sections throughout.
[0023] In the accompanying drawings:
[0024] Fig. 1 is a block diagram of an exemplary computing device which may be used with some embodiments;
[0025] Fig. 2 is a block diagram of a system for secure generation of avatars and videos, according to some embodiments;
[0026] Fig. 3 is a block diagram of a system for complaint resolution, according to some embodiments;
[0027] Fig. 4 is a flowchart of a method for enrolling a new user on an avatar generation platform, according to some embodiments;
[0028] Fig. 5 is a flowchart of a method for generating an avatar and a video by an avatar generation platform, according to some embodiments;
[0029] Fig. 6 is a flowchart of a method for conducting a complaint resolution process by an avatar generation platform, according to some embodiments; and
[0030] Fig. 7 is a block diagram of an avatar generation platform, according to embodiments of the invention.
[0031] It will be appreciated that, for simplicity and clarity of illustration, elements shown in the figures have not necessarily been drawn to scale. For example, the dimensions of some of the elements may be exaggerated relative to other elements for clarity. Further, where considered appropriate, reference numerals may be repeated among the figures to indicate corresponding or analogous elements.DETAILED DESCRIPTION
[0032] In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the disclosure. However, it will be understood by those skilled in the art that the present disclosure can be practiced without these specific details. In other instances, well-known methods, procedures, and components, modules, units and / or circuits have not been described in detail so as not to obscure the disclosure.
[0033] According to embodiments of the invention, ML models disclosed herein, also referred herein as networks, may include one or more artificial neural networks (NN). NNs are mathematical models of systems made up of computing units typically called neurons (which are artificial neurons or nodes, as opposed to biological neurons) communicating with each other via connections, links or edges. In common NN implementations, the signal at the link between artificial neurons or nodes can be for example a real number, and the output of each neuron or node can be computed by function of the (typically weighted) sum of its inputs, such as a rectified linear unit (ReLU) function. NN links or edges typically have a weight that adjusts as learning or training proceeds typically using a loss or cost function, which may for example be a function describingthe difference between a NN output and the ground truth (e.g., correct answer). The weight may increase or decrease the strength of the signal at a connection. Typically, NN neurons or nodes are divided or arranged into layers, where different layers can perform different kinds of transformations on their inputs and can have different patterns of connections with other layers. NN systems can learn to perform tasks by considering example input data, generally without being programmed with any task-specific rules, being presented with the correct output for the data, and self-correcting, or learning using the loss function. A NN may be configured or trained for a specific task, e.g., image processing, pattern recognition or classification. Training a NN for the specific task may involve adjusting these weights based on examples (e.g., labeled data included in the training dataset). Each neuron of an intermediate or last layer may receive an input signal, e.g., a weighted sum of output signals from other neurons, and may process the input signal using a linear and / or nonlinear function (e.g., an activation function). The results of the input and intermediate layers may be transferred to other neurons and the results of the output layer may be provided as the output of the NN. For example, in a NN algorithm known as the gradient descent algorithm, the results of the output layer may be compared to the labels of the samples in the training dataset, and a loss or cost function (such as the root-mean-square error) may be used to calculate a difference between the results of the output layer and the labels. The weights of some of the neurons may be adjusted using the calculated differences, in a process that iteratively minimizes the loss or cost until satisfactory metrics are achieved or satisfied.
[0034] A processor, e.g., central processing units (CPU), graphical processing units or fractional graphical processing units (GPU), tensor processing units (TPU) or a dedicated hardware device may perform the relevant calculations on the mathematical constructs representing the NN. As used herein a NN may include deep neural networks (DNN), convolutional neural networks (CNN), recurrent neural networks (RNN), long short-term memory networks (LSTM) probabilistic neural networks (PNN), time delay neural network (TDNN), deep stacking network (DSN), generative adversarial networks (GAN), recurrent neural network (RNN), long short-term memory (LSTM), Siamese networks, etc.
[0035] Some embodiments of the invention may include other deep architectures such as transformers, that may include series of layers of self-attention mechanisms and feedforward neural networks, used for processing input data. Transformers may be used in light of their capacity of parallelism and their multi-headed self-attention which facilitate features extraction.
[0036] Some algorithms for training a NN model such as gradient descent may enable training the NN model using samples taken from a training dataset. Each sample may be fed into, e.g., provided as input to, the NN model and a prediction may be made. At the end of a training session, the resulting predictions may be compared to the expected output variables, and loss or cost function may be calculated. The loss or cost function is then used to train the NN model, e.g., to adjust the model weights, for example using backpropagation and / or other training methods. Embodiments of the invention may use a loss function. The loss function may be used in the training process to adjust weights and other parameters in the various networks in a back propagation process.
[0037] A digital image, also referred to herein simply as an image, may include a visual (e.g., optical) representation of physical objects, specifically, a face of a human, provided in any applicable digital and computer format. Images may include a simple 2D array or matrix of computer pixels, e.g., values representing one or more light wavelengths or one or more ranges of light wavelength, within the visible light, in specified locations, or any other digital representation, provided in any applicable digital format such as jpg, bmp, tiff, etc. A digital image may be provided in a digital image file containing image data.
[0038] A voice or speech sample may be provided in any applicable computerized audio format such as the MP3, MP4, M4A, WMA (windows media audio), FLAC (free lossless audio codec), ALAC (apple lossless audio codec), WAV (waveform audio file), etc. The generated animation or video clip, e.g., a moving visual media with or without audio, may be provided in any applicable computerized video format such as the MP4 (MPEG-4 part 14), MOV (QUICKTIME movie), WMV (Windows media viewer), AVI (audio video interleave), AVCHD (advanced video coding high definition), Flash video formats FLV, F4V, and SWF (Shockwave Flash) etc.
[0039] Embodiments of the invention may provide a system and method for secure generation of avatars using a computerized avatar generation application or platform. Embodiments of the invention may include an enrolment or registration phase for registering new users to the systems, e.g., users that are not already enrolled into the system and request to use the system for the first time. In the enrollment or registration phase, the new user is identified, e.g., using two-factor authentication (2FA) and / or any other applicable identification process, and an identity log is generated for the user. The identity log may include data related to or indicative of a facial image of the new user (e.g., one or more facial images of the new user, an embedding generated for the one or more facial images, etc.), and data related to or indicative of a voice of the new user (e.g.,one or more voice recordings of the new user, an embedding generated for the one or more voice recordings, etc.).
[0040] The enrolment or registration phase may further include liveness verification of the new user, e.g., a process by which embodiments of the invention may determine that the new user is a human user and not a machine. For example, embodiments of the invention may present a dynamic challenge to the new user. A dynamic challenge may refer to a task for the user, where the task is not constant, e.g., the task may be different for different users. The task may include specific words or phrases that the user has to say and / or specific movements that the user has to do (e.g., move your head to the right’). In order to register on the platform, the new user may record a selfie video showing the user performing the dynamic challenge. For example, the user may have to say the specific words and perform the specific movements included in the dynamic challenge. The user may record the selfie video using the existing features of the user’ s mobile phone.
[0041] The recorded selfie video may be verified against the dynamic challenge, e.g., the recorded selfie video may be analyzed to determine which words or phrases were said in the recorded selfie video, and the video may be analyzed to determine which movements were made in the recorded selfie video. If the text and the movements in the recorded selfie video are the same as in the dynamic challenge, then it may be determined that the verification is successful, and the new user may be enrolled or registered on the system. If, however, the text and / or the movements in the recorded selfie video are not the same as in the dynamic challenge, then it may be determined that the verification is not successful, and the new user may be refused and the registration denied. In addition, the recorded selfie video may be used to prepare the identity log, e.g., the data indicative of a facial image of the new user and the data indicative of a voice of the new user may be derived from the recorded selfie video.
[0042] Requesting the new user to record a selfie video with a dynamic challenge may verify that the user as a live human, not a video or deepfake. In addition, using a dynamic challenge instead of a constant and predictable challenge may prevent fraudsters from uploading a fake or prepared videos. Thus, embodiments of the invention may improve the technology of avatar generation, by preventing fraudsters from generating avatars.
[0043] Once a user is enrolled or registered on the system, the user may generate its own avatar, provided that the identity of the user is verified against the identity log of the user, if stored,otherwise he may have to verify his identity again. For example, when a user that is already enrolled or registered on the system requires to generate its own avatar, the identity of the requester may be verified against the identity log of the enrolled user. If the verification of the identity of the requester against the identity log of the enrolled user is successful an avatar for the requester may be generated and stored. If, however, the verification of the identity of the requester against the identity log of the enrolled user is not successful, the request may be denied, and no avatar may be generated. Verification of the requester against the identity log of the user may be performed by requesting the user to upload a selfie image, a voice recording and / or a selfie video of the requester, and comparing the facial image from the uploaded selfie image or selfie video with the data indicative of the facial image included in the identity log of the user and comparing the voice recording or the voice sample from the selfie video with and the data indicative of a voice of the user in the identity log.
[0044] Once an avatar is generated, the avatar may be used for generating videos or animations. Generating a video may include receiving from the user a text or an audio script or a voice-over script, e.g., the words or phrases that the avatar should say or pronounce in the video. Embodiments of the invention may include real-time moderation tools for the uploaded audio, to make sure only appropriate content is generated by the computerized avatar generation platform. For example, the text uploaded by the user may be scanned by dedicated ML model to detect and ban inappropriate language. The video may be generated using known in the art Al tools, where in the final video, the avatar may be saying the provided text or voice-over script. According to the sophistication of the used Al tools, lips movements that are coordinated with the text may be added, with or without other movements and gestures.
[0045] According to embodiment of the invention one or more visible and / or invisible watermarks may be added to the video and / or audio. The watermarks may include data identifying the user, e.g., a user identification number (ID), data identifying the video, e.g., video ID and / or a representation of the video, and / or data identifying the computerized avatar generation platform. The representation of the video may be generated or calculated from the video, by applying mathematical formula or an ML model. The representation of the video may be or may include a hash of the video, an embedding of the video, e.g., generated by a NN trained for that purpose, or any other data that may identify the video. The final video may be provided to the user, and the video identifying data may be stored for later use.
[0046] Embodiments of the invention may further provide a complaint resolution service. For example, a user may file a complaint about a video that includes an avatar. The complaint may refer to inappropriate content in the video and / or to unauthorized use of an avatar of the user in the video. The complaint resolution service may include automatically performing a series of tests and checks, by a computer, to resolve the complaint. For example, a first test may include determining if the examined video was generated using the avatar generation platform. This may be performed in various ways, for example, using the data identifying the computerized avatar generation platform embedded in the watermark, using the video ID embedded in the watermark, by generating a representation for the examined video and searching for the same representation in the system's database, etc. If it is determined that the video was not generated by the computerized avatar generation platform, then a notice that the video was not generated by the computerized avatar generation platform may be provided to the user, and the complaint may be resolved. If, however, it is determined that the video was generated by the computerized avatar generation platform, examination may continue. For example, a next step may include identifying the video and the user that generated the video. The video and the user that generated the video may be identified based on data embedded in the watermarks and / or based on the representation of the video. In case the complaint is about inappropriate content, the content of the video may be analyzed, e.g., by a machine or by a human observer. If the analysis results suggest that the content of the video is indeed inappropriate, the user that has prepared the video may be banned, similarly, if the complaint refers to an unauthorized use of an avatar of the complaining user, the user that has prepared the video may be banned.
[0047] Thus, embodiments of the invention may improve the technology of avatar generation, by enhancing the security and ethical use of digital human avatars. Embodiments of the invention may employ an enrolment phase, in which the identity of the enrolled user is identified, and liveliness of the user is examined. This may prevent non-human, e.g., deepfake, users from generating avatars. In addition, embodiments of the invention may add watermarks to the generated videos, to enable identifying which platform has generated the avatar and video, which user has generated the avatar and video, etc. Thus, embodiments of the invention may enable resolving complaints and detecting fraud.
[0048] Embodiments of the invention may include technical products, processes, and policies to prevent unauthorized or unethical applications of artificial intelligence (Al) generated avatars.Embodiments of the invention may addresses issues related to deepfake technology which can undermine digital trust and security, such as privacy invasion, misinformation, and digital impersonation. Embodiments of the invention may implement a series of checks and balances, including user identity verification, moderation tools, and watermarking to ensure each digital avatar created or manipulated through the technology is ethical and authorized.
[0049] Current Al and deepfake detection technologies exist, but none is linked to a specific person’s identity before the creation of the avatar. Current solutions in the field lack comprehensive integration of security measures or rely heavily on post-facto detection methods which may be less effective against new and sophisticated fake content generation techniques.
[0050] Embodiments of the invention may integrate preventive measures directly into the avatar creation process, provide a robust real-time moderation, and employ multi-layered security features which collectively may enhance overall system integrity. Embodiments of the invention may improve the technology of digital human avatars by providing enhanced trust and safety in digital content, comprehensive approach combining prevention, detection, and response strategies, and continuous innovation cycle with red and blue teams.
[0051] Embodiments of the invention may be implemented with varying levels of security checks and third-party integrations, allowing for customization based on user needs and risk levels. Some embodiments of the invention may include using a full suite of proprietary tools and third-party integrations for maximum security and compliance. This may include rigorous user and avatar verification processes, real-time moderation, and robust watermarking technologies to ensure the authenticity and ethical use of digital avatars. The comprehensive approach of embodiments of the invention may maintain high standards of security and trust, crucial for user acceptance and regulatory compliance in sensitive applications involving digital identities.
[0052] Reference is now made to Fig. 1, which is a block diagram of an exemplary computing device which may be used with some embodiments. Computing device 100 may include a controller or processor 105 that may be or include, for example, logic, one or more central processing unit processor(s) (CPU), one or more graphics processing unit(s) (GPU), one or more data processing unit(s) (DPU), one or more application- specific integrated circuits (ASICs), one or more field programmable arrays (FPGAs), one or more integrated circuits (ICs), a system-on-chip (SoC), a chip or any suitable computing or computational device, an operating system 115, a memory 120, a storage 130, input devices 135 and output devices 140, all connected to each other.Each of modules and equipment such as enrolment module 210, avatar generation module 230, video generation module 250 and complaint resolution module 270, shown in Fig. 2, complaint resolution module 380, shown in Fig. 3, and other modules or equipment mentioned herein may be or include, or may be executed by, a computing device such as included in Fig. 1 or specific components of Fig. 1, although various units among these entities may be combined into one computing device. In one embodiment, each of the above -listed modules comprise software that is used to program one or more processors (e.g., processor 105) to perform the functions described herein.
[0053] Operating system 115 may be or may include any code segment designed and / or configured to perform tasks involving coordination, scheduling, arbitration, supervising, controlling or otherwise managing operation of computing device 100, for example, scheduling execution of programs. Memory 120 may be or may include, for example, a random access memory (RAM), a read only memory (ROM), a dynamic RAM (DRAM), a synchronous DRAM (SD-RAM), a double data rate (DDR) memory chip, a Flash memory, a volatile memory, a non-volatile memory, a cache memory, a buffer, a short term memory unit, a long term memory unit, or other suitable memory units or storage units. Memory 120 may be or may include a plurality of possibly different memory units. Memory 120 may store for example, instructions to carry out a method (e.g., executable code 125), and / or data such as a state file, contexts, various queues as disclosed herein, etc.
[0054] Executable code 125 may be any executable code, e.g., software or firmware, an application, a program, a process, task or script (e.g., any of the above-listed modules). Executable code 125 may be executed by processor 105 possibly under control of operating system 115. In some embodiments, more than one computing device 100 or components of device 100 may be used for multiple functions described herein. For the various modules and functions described herein, one or more computing devices 100 or components of computing device 100 may be used. Devices that include components similar or different to those included in computing device 100 may be used, and may be connected to a network and used as a system. One or more processor(s) 105 may be configured to carry out embodiments by, for example, executing software or code.
[0055] Storage 130 (which is non-transitory) may be or may include, for example, a hard disk drive, a floppy disk drive, a compact Disk (CD) drive, a CD-recordable (CD-R) drive, a universal serial bus (USB) device or other suitable removable and / or fixed storage unit. Data, simulation data, may be stored in a storage 130 and may be loaded from storage 130 into a memory 120 whereit may be processed by processor 105. In some embodiments, some of the components shown in Fig. 1 may be omitted. Storage 130 may store data required for performing embodiments of the invention, such as parameters and weights of ML models. In addition, storage 130 may store databases and / or data repositories such as identity log database 132, video log database 134 and avatar database 136.
[0056] Input devices 135 may be or may include a mouse, a keyboard, a touch screen or pad or any suitable input device. It will be recognized that any suitable number of input devices may be operatively connected to computing device 100 as shown by block 135. Output devices 140 may include one or more displays, speakers and / or any other suitable output devices. It will be recognized that any suitable number of output devices may be operatively connected to computing device 100 as shown by block 140. Any applicable input / output (I / O) devices may be connected to computing device 100, for example, a wired or wireless network interface card, a modem, printer or facsimile machine, a universal serial bus (USB) device or external hard drive may be included in input devices 135 and / or output devices 140.
[0057] Embodiments may include one or more article(s) (e.g., memory 120 or storage 130) such as a computer or processor non-transitory readable medium, or a computer or processor non- transitory storage medium, such as for example a memory, a disk drive, or a USB flash memory, encoding, including or storing instructions, e.g., computer-executable instructions, which, when executed by a processor or controller, carry out methods disclosed herein.
[0058] Reference is made to Fig. 2, which is a block diagram a system 200 for secure generation of avatars and videos, according to some embodiments of the invention. It should be understood in advance that the components and functions shown in Fig. 2 are intended to be illustrative only and embodiments of the invention are not limited thereto. While in some embodiments the system of Fig. 2 is implemented using systems as shown in Fig. 1, in other embodiments other systems and equipment can be used. System 200 may be or may be a part of an avatar generation platform 700 shown in Fig. 7.
[0059] Enrolment module 210 may enroll or register users on system 200. Enrolling a user on system 200 may include a series of operations intended to verify the identity of the user, verify that the user is a human and not a deepfake generated by a machine, and to generate an identity log 220 for the user. Identity log 220 may include a user ID, biometric identification data, the avatar of the user, data identifying the avatar of the user, and other data related to the user, that may enable theavatar generation platform to identify the user in future times when the user logs into system 200 or to identify the avatar in case a complaint is filed of misuse of the avatar. In some embodiments, identity log 220 may be stored in identity log database 132. In some embodiments, the user may download and store (e.g., in a storage separate from system 200) an encrypted and digitally signed version of identity log 220 for future use. In such case, identity log 220 may not be stored in identity log database 132 or only parts of identity log 220 may be stored in identity log database 132. Thus, when a new avatar is to be made and verification is required again, the user may upload the stored version of identity log 220 and system 200 may perform signature verification and decryption to restore identity log 220, and verify the user.
[0060] In some embodiments, enrolment or registration of a user on system 200 by enrolment module 210 may include a sign-up stage. The sign-up stage may include using a collection of methods for identifying and authenticating the user. For example, enrolment module 210 may use one or more of the following techniques to identify and authenticate the user: shared authentication schemes, email verification, two-factor authentication (2FA) or multiple factor authentication (MFA), automated know your customer (KYC) verification, credit card verification (with or without charge) and / or a visual verification of the user against a national document. An example for a shared authentication scheme may include single sign-on (SSO) that may be implemented using the security assertion markup language (SMAL) protocol. Other methods may be used.
[0061] A shared authentication scheme, and specifically SSO, is an authentication process that may allow a user to access multiple applications or services with one set of login credentials such as a username and password. Once the user logs in through the SSO system, and as long as the SSO session is open, the user can seamlessly access all connected applications without needing to reenter credentials for each one. If SSO is used for user authentication, then when the user attempts to register or enroll on system 200, enrolment module 210 may redirect the user to an identity provider (IdP) for authentication. The user may authenticate with the IdP, e.g., via login credentials, biometric, etc., or may be already authenticated (if an SSO session is already open for the user). After successful authentication, the IdP may issue an authentication token and send the authentication token to enrolment module 210, e.g., using the SMAL protocol.
[0062] 2FA is a security process that requires users to provide two different types of identification before gaining access to an account or system. This additional layer of security may make it more difficult for unauthorized individuals to access system 200, even if they have the password of theuser. MFA is similar to 2FA, but requires users to provide more than two different types of identification before gaining access to an account or system. The factors may include something you know, e.g., a password or personal identification number (PIN), something you have, e.g., a smartphone, hardware token, or security card and / or something you are, e.g., a fingerprint, facial recognition, or other biometric data. If 2FA or MFA is used for user authentication, then when the user attempts to register or enroll into system 200, enrolment module 210 may obtain user credentials (e.g., password), and prompts for a second factor (e.g., a code sent to a phone of the user, fingerprint scan, etc.). The user may provide this second factor, and the process may be repeated for the required number of factors. To authenticate the user, enrolment module 210 may verify all factors.
[0063] Email verification is a process of confirming that an email address provided by the user is valid, exists, and belongs to the user. If email verification is used for user authentication, then when a user attempts to register or enroll on system 200, enrolment module 210 may request the user to provide an email address, and enrolment module 210 may send a verification email with a unique link or code to the email address provided by the user. Once the link is clicked or the code is obtained from the user, the email address is verified.
[0064] Automated KYC is the use of a collection of techniques to digitally verify the identity of a user. If automated KYC is used for user authentication, then when a user attempts to register or enroll into system 200, enrolment module 210 may request the user to upload an ID document and take a selfie image or video, and may use Al to verify document authenticity (e.g., by chasing watermarks, fonts, expiration, tampering of the document), face match (biometric comparison of the face in the uploaded selfie to facial image of the user in a national document) and liveness (if a selfie video is uploaded it may be checked for blinks, head turns, etc.).
[0065] Once sign-up stage is completed, e.g., enrolment module 210 has successfully authenticated the user using one or more of the above described methods (or other suitable methods), enrolment module 210 may request user consent for generating or preparing an avatar of the user and videos with the avatar of the user. Enrolment module 210 may request the user consent using a consent checkbox and / or by presenting the terms of use (TOU) to the user and using a TOU checkbox.
[0066] Enrolment module 210 may further verify a credit card of the user. Credit card verification may include verifying the authenticity of a credit or debit card of the user. If credit card verification is used, either for user authentication (e.g., without payment), or for payment, then enrolmentmodule 210 may request the user to provide the credit card details (e.g., credit card number, expiration date, etc.) of the user, and verify the credit card details against the credit card provider databases. Enrolment module 210 may further check that there are no fraud risk indicators associated with the uploaded credit card.
[0067] In order to register a new avatar on system 200, enrolment module 210 may perform a dynamic challenge test to the user. The dynamic challenge test may include presenting a dynamic challenge to the user, requesting the user to record a selfie video with a dynamic challenge and verifying the selfie video against the dynamic challenge. The dynamic challenge may include a script including, for example text (e.g., specific words or phrases) that the user is requested to say and / or one or more specific movements or actions that the user is requested to perform (e.g., ‘move your head to the right’). Enrolment module 210 may generate the script dynamically, e.g., generate a script that is unexpected and different for different users, so that the user can’t know the challenge in advance and fake or upload a preprepared video. For example, enrolment module 210 may randomly select the text and the action from a repository of texts and actions. Other methods for generating the dynamic challenge may be used. Enrolment module 210 may verify the selfie video against the dynamic challenge by analyzing the selfie video uploaded by the user, e.g., using speech to text conversion and gesture recognition or ML models that perform speech to text conversion and gesture recognition, and comparing the text and movements extracted from uploaded video with the text and movements in the dynamic challenge. If the verification of the selfie video is successful, enrolment module 210 may enroll the user on system 200 (the computerized avatar generation platform). If, however, the verification of the selfie video fails, enrolment module 210 may refuse to enroll the user on the system 200. Using a dynamic challenge may verify the user as a live human, not a video or deepfake, using the existing features of the user’s mobile phone. It is noted that the uploaded video may be used to verify the identity of the individual, e.g., the same selfie video may be used for avatar enrolment and for the automated KYC.
[0068] Once the selfie video is verified against the dynamic challenge, enrolment module 210 may generate an identity log 220 of the user, and store the identity log 220, e.g., in identity log database 132. Identity log 220 may include biometric identification data extracted from the selfie video of the enrolled user, e.g., data indicative of a facial image of the user, data indicative of a voice of the user, gate patten (e.g., limbs movements extracted from the video), heartbeats, veins (e.g., if the video camera has some Infra-Red sensitivity then veins can be seen as well as heartbeats) and / orfingerprints, etc. The data indicative of a facial image of the user may include one or more facial images of the user, extracted from the selfie video and / or a representation of the one or more facial images of the user, e.g., an embedding generated for example by an ML model or by another tool. The data indicative of the voice of the user may include one or more voice recordings of the user, extracted from the selfie video and / or a representation of the one or more voice recordings of the user, e.g., an embedding generated for example by an ML model or by another tool.
[0069] Avatar generation module 230 may receive a request from a user (referred to herein as the requester) to generate an avatar of the requester. Before generating the avatar, avatar generation module 230 may verify the identity of the requester against an identity log 220 of an enrolled user, e.g., verify that the requester and the enrolled user are the same person. As noted, in some embodiments, following a successful enrollment, the user may download an encrypted and digitally signed version of identity log 220. In this case, the user may upload the stored version of identity log 220 and avatar generation module 230 may perform signature verification and decryption to restore identity log 220, and verify the user. The avatar generation module 230 may generate avatar 240 for the requester only if the verification of the identity of the requester against the identity log 220 of an enrolled user (or the uploaded identity log 220) is successful, e.g., only if it is verified that the requester and the enrolled user are the same person. Avatar generation module 230 may deny the request if the verification of the identity of the requester against the identity log of the enrolled user is not successful, suggesting that the requester and the enrolled user are different people.
[0070] Avatar generation module 230 may perform the verification by requesting the requester to upload a biometric identification data of the requester, e.g., a selfie image, a voice recording and / or a selfie video of the requester. Avatar generation module 230 may verify the identity of the requester against the identity log of the enrolled user by verifying the biometric identification data uploaded by the requester against the biometric identification data in the identity log 220 of the enrolled user. For example, avatar generation module 230 may verify a selfie of the requester (e.g., the selfie image uploaded by the requester or a frame form the selfie video uploaded by the requester that includes the face of the requester) against the data indicative of the facial image of the enrolled user stored in identity log 220 of the enrolled user, and / or by verifying a voice recording of the requester (e.g., the voice sample uploaded by the requester or a voice sample extracted from the selfie video uploaded by the requester) against data indicative of the voice ofthe enrolled user stored in identity log 220. Verification of the selfie of the requester against the data indicative of the facial image of the enrolled user, may include verifying that both belong or depict the same person, and may be performed by dedicated ML models trained for that purpose. Verification of the voice recording of the requester against the data indicative of the voice of the enrolled user, may include verifying that both belong to the same person, and may be performed by dedicated ML models trained for that purpose. Other types of biometric data may be verified, e.g., veins, heartbeats, fingerprints, etc.
[0071] If the data indicative of the facial image of the enrolled user includes a representation or embedding of the facial image of the enrolled user generated by an ML model, then a representation or embedding of the facial image of the requester may be generated by the same ML model. The level of match between the two images may be estimated by calculating a distance metric (e.g., cosine similarity or Euclidian distance) and comparing the distance metric to a threshold. If the data indicative of the facial image of the enrolled user includes the facial image of the enrolled user, a Siamese network may be used. The Siamese network may use the facial image of the requester as input and the facial image of the enrolled user as a reference image, and generate a similarity score denoting the probability that the two input images are of the same person. Similar principles may be applied for verifying the voice recording of the requester against the data indicative of the voice of the enrolled user. Other methods may be used.
[0072] As noted, avatar generation module 230 may generate avatar 240 for the requester only if the verification of the identity of the requester against the identity log of the enrolled user is successful. Avatar 240 may be two-dimensional or three-dimensional, and may be generated using any applicable method from the selfie image, voice recording and / or selfie video uploaded by the requester. Avatar 240 may be stored in avatar database 136 and be associated to the user ID.
[0073] According to some embodiments, avatar generation module 230 may skip the step of verifying the identity of the requester if a user that has just enrolled on system 200 requests to generate an avatar within the same session of enrollment, e.g., without logging out or disconnecting from system 200. The user may either actively log out of system 200, or system 200 may terminate the session and disconnect the user after a certain time with no activity in the session. If neither has happened, avatar generation module 230 may generate avatar 240 without the verification using the video uploaded for the dynamic challenge. However, whenever an enrolled user wants to upload a new selfie image, selfie video or audio recording for generating avatar 240 (even within the samesession), avatar generation module 230 may verify the identity of the requester against the identity log of the enrolled user and generate avatar 240 only if the verification is successful, e.g., only if the person in the uploaded selfie image, selfie video and audio recording is the same person as the enrolled user. Thus, an enrolled user may only generate avatar 240 of himself / herself. A user may not be able to enroll on system 200, upload a selfie image, selfie video or audio recording of a different user and generate avatar 240 of the different user.
[0074] If an enrolled user logs out or disconnects from system 200 and wants to login again, the enrolled user may have to repeat a part or all of the identification process performed by enrolment module 210, only instead of generating an identity log 220 that already exists, the selfie video uploaded for the dynamic challenge verification may be verified against the already existing identity log 220 to make sure that this is the same person.
[0075] By verifying a selfie of the requester (e.g., the selfie image uploaded by the requester or a frame form the selfie video uploaded by the requester that includes the face of the requester) against the data indicative of the facial image of the enrolled user stored in identity log 220 of the enrolled user, and / or by verifying a voice recording of the requester (e.g., the voice sample uploaded by the requester or a voice sample extracted from the selfie video uploaded by the requester) against data indicative of the voice of the enrolled user stored in identity log 220.
[0076] After avatar 240 is generated for the requester, video generation module 250 may generate video 260 using avatar 240. To generate video 260, also referred to as animation 260, video generation module 250 may receive a text or an audio script or a voice-over script, e.g., the words or phrases that the avatar should say or pronounce in the video, from the requester. Video generation module 250 may include real-time moderation tools for the uploaded audio, to make sure only appropriate content is generated by system 200. For example, the text uploaded by the requester may be scanned by dedicated ML model to detect and ban inappropriate language. Video generation module 250 may generate video 260 using dedicated ML models such as generative GANs, style GANs, etc., where in the final video 260, avatar 240 may be saying the provided text or voice-over script. Video generation module 250 may use ML models to add lips movements that are coordinated with the text, with or without other movements and gestures.
[0077] Video generation module 250 may add or insert one or more visible and / or invisible watermarks 262-264 to video 260. For example, watermark 262 may be visible, e.g., visible watermark 262 may be seen by a viewer when video 260 is presented on a display or a screen.Invisible watermark 264 may be an invisible watermark inserted into the image part of video 260 (the part containing image data), e.g., invisible watermark 264 may not be seen by a viewer when video 260 is played on a display or a screen, and invisible audio watermark 266 may be an invisible watermark inserted into the audio part of video 260, e.g., invisible video watermark 264 may not be seen or heard by a viewer when video 260 is played. Each of watermarks 262-264 may include data identifying the requester (e.g., user ID), data identifying the video (e.g., video ID), and / or data identifying system 200, e.g., the computerized avatar generation platform. Video generation module 250 may add metadata to video 260, including any applicable data related to video 160, e.g., time of creation. In addition, video generation module 250 may generate video log 270, e.g., the video identifying data, from video 260. Video log 270 may include a representation of video 260, e.g., a hash or embedding of video 260, e.g., generated using a mathematical formula or by an ML model or an NN trained for that purpose, a video ID number, or any other data that may identify video 260. Video generation module 250 may provide video 260 (e.g., a file comprising video 260) to the user, and store video log 270 in video log database 134 for later use.
[0078] Reference is made to Fig. 3, which is a block diagram of a system 300 for complaint resolution, according to embodiments of the invention. It should be understood in advance that the components and functions shown in Fig. 3 are intended to be illustrative only and embodiments of the invention are not limited thereto. While in some embodiments the system of Fig. 3 is implemented using systems as shown in Fig. 1 , in other embodiments other systems and equipment can be used. System 300 may be a part of an avatar generation platform 700 and work in conjunction with system 200.
[0079] Complaint resolution module 380 may receive a complaint 320 associated with an examined video 360 presenting an avatar 340. For example, complaint 320 may be about inappropriate content in video 360, or about illegitimate use of avatar 340, e.g., that an avatar 340 of the user that filed the complaint (referred to herein as complainant) was used without permission. Other complaints may also be filed. Complaint resolution module 380 may determine whether examined video 360 was generated by computerized avatar generation platform 200 based on the one or more watermark 362-366. For example, complaint resolution module 380 may analyze examined video 360, e.g. attempt to extract one or more watermarks 362-366 from examined video 360, to identify whether examined video 360 was generated by system 200 of the avatar generation platform or by another system. Other methods may be used to determine whether examined video360 was generated by computerized avatar generation platform 200. Complaint resolution module 380 may handle complaint 350, e.g., determine whether the content of examined video 360 is appropriate and / or whether avatar 340 is misused, only if examined video 360 was generated by system 200 of the avatar generation platform.
[0080] For example, complaint resolution module 380 may extract data identifying examined video 360. The data may be extracted from visible watermark 362, from invisible video watermark 364 and or from invisible audio watermark 366. The data extracted from watermarks 362-366 may include data identifying the user that has generated video 360, a video identification number, and / or data identifying system 200. Thus, if data is identified successfully from watermarks 362-366, complaint resolution module 380 may determine, based on the extracted data, whether examined video 360 was generated by system 200 and which user has generated examined video 360.
[0081] If the data extracted from watermarks 362-366 suggests that examined video 360 was not generated by system 200, then complaint resolution module 380 may notify the complainant that the examined video 360 was not generated by system 200, and the compliant resolution process may terminate. If, however, the data extracted from watermarks 362-366 suggests that examined video 360 was generated by system 200, investigation may continue.
[0082] If the subject of complaint 350 is that the content of examined video 360 is inappropriate, complaint resolution module 380 may continue the examination by analyzing the content of examined video 360. Complaint resolution module 380 may analyze the content of examined video 360 using automatic tools, e.g., large language models (LLMs) and other ML models. For example, complaint resolution module 380 may analyze that audio part of examined video 360 by converting the audio track of examined video 360 into text using speech-to-text engine and prompt a general purpose LLM to identify inappropriate content in the text of examined video 360. Additionally or alternatively, complaint resolution module 380 may use LLMs trained for identifying inappropriate content. Other methods may be used. Complaint resolution module 380 may further use image classifiers and image analysis tools to analyze the visual part of examined video 360 to identify inappropriate visuals in examined video 360. Complaint resolution module 380 may combine the analysis of the audio part with the analysis of the visual part of examined video 360 to generate an appropriateness score of examined video 360. Complaint resolution module 380 may compare the appropriateness score to a threshold to determine wither the content in examined video 360 to obtain a resolution 310 whether the complaint is justified or not, e.g., whether the content of thevideo is appropriate or inappropriate. Additionally or alternatively, complaint resolution module 380 may present examined video 360 to a human reviewer. In some embodiments, complaint resolution module 380 may present examined video 360 to a human reviewer only if the appropriateness score accedes a threshold. If examined video 360 is presented to a human reviewer, then complaint resolution module 380 may receive the appropriateness score or determination from the human reviewer.
[0083] If resolution 310 determined by complaint resolution module 380 is that the content in examined video 360 is appropriate, then complaint resolution module 380 may provide a notice to the complainant that the complaint has been found unjustified, and the complaint resolution process may terminate. If, however, resolution 310 is that the content in examined video 360 is indeed inappropriate, then complaint resolution module 380 may extract data identifying the user that generated examined video 360 from examined video 360. The data identifying the user that generated the examined video may be or may include the user ID of the user that generated examined video 360. For example, complaint resolution module 380 may extract the video ID of examined video 360 from watermarks 362-366, use the video ID to retrieve video log 370 of examined video 360 from video log database 134, and extract the user ID from video log 370. Alternatively, complaint resolution module 380 may extract the user ID directly from watermarks 362-366.
[0084] Once the user ID is found, complaint resolution module 380 may ban the user that has prepared examined video 360 from using system 200 for making additional videos. For example, complaint resolution module 380 may add a note or a remark indicating that the user is banned to identity log 320 of the user that has prepared examined video 360, or the user ID of the user that has prepared examined video 360 may be added to a denylist of banned users. Complaint resolution module 380 may perform other actions if the content of examined video 360 is determined to be inappropriate such as notifying the complainant that complaint 350 has been found justified and that the user that has prepared examined video 360 has been banned, notify the user that has prepared examined video 360 that he / she are banned, send a report to a system administrator indicating the video ID of examined video 360, the user ID of the banned user, and the resons for banning the user, etc.
[0085] In case that examined video 360 has been changed, damaged or tampered in a way that prevents complaint resolution module 380 from being able to extract data from watermarks 362-366, then complaint resolution module 380 may extract the data identifying the user that generated the examined video by generating a representation (e.g., an embedding) from examined video 360, e.g., using the same ML model or an NN as used by video generation module 250. Complaint resolution module 380 may perform similarity checks between the representation of examined video 360 and the representations stored in video log database 134 to identify a matching representation in video log database 134. If a matching representation is found in video log database 134, then the video log 370 that includes the matching representation may be retrieved and the user ID of the user that have prepared examined video 360 may be identified from video log 370. If, however, a matching representation or embedding is not found in video log database 134, then complaint resolution module 380 may notify the complainant that it has failed to identify examined video 360 and the user that has generated examined video 360. A notice may be sent to a system administrator as well. A matching representation may be a representation that is the closest to the representation of examined video 360 within video log database 134, e.g., with a lowest distance metric, and with the distance metric below a distance threshold.
[0086] If the complaint is about illegitimate use of avatar 340, e.g., that avatar 340 of the complainant was used without permission, then complaint resolution module 380 may retrieve complainant avatar 336 from avatar database 136, and compare avatar 340 in video 360 with the complainant avatar 336. If it is determined that the video was generated by the user associated with identity log 320, but the avatar 340 in video 360 is complainant avatar 336, then the complaint may be found justified. In case of a justified compliant, complaint resolution module 380 may ban the user that has generated examined video 360 (e.g., place a notice in the identity log 320 of the user that has generated video 360 or place the user ID of the user that has generated video 360 in a denylist), send a report to the complainant, to the user that has generated examined video 360, and / or to a system administrator. If avatar 340 in video 360 does not match the avatar of the complainant, then complaint resolution module 380 may determine that the complaint is unjustified and notify this resolution 310 to the complainant. Other methods may be used.
[0087] Reference is now made to Fig. 4, which is a flowchart of one embodiment of a method for enrolling a new user on an avatar generation platform. While in some embodiments the operations of Fig. 4 are carried out using systems as shown in Figs. 1 and 2, in other embodiments other systems and equipment can be used. Enrolling a new user on an avatar generation platform may include a series of operations intended to verify the identity of the user, verify that the user is ahuman and not a deepfake generated by a machine, and to generate an identity log for the user. A new user may be a user that is not already enrolled on the avatar generation platform and uses the avatar generation platform for the first time.
[0088] In operation 410, a processor (e.g., processor 105 depicted in Fig. 1 executing code to carry out the method for enrolling a new user to an avatar generation platform according to embodiments of the present invention) may identify and authenticate the user, e.g., using one or more of a shared authentication scheme (e.g., such as SSO), email verification, 2FA or MFA, automated KYC, credit card verification (with or without charge) and / or a visual verification of the user against a national document. In operation 420, the processor may request user consent for generating an avatar of the user and videos with the avatar of the user, using e.g., a consent checkbox and / or a TOU checkbox. In operation 430, the processor may perform credit card verification, with or without payment. In operation 440, the processor may perform a dynamic challenge test to the user. The dynamic challenge test may include providing a dynamic challenge to the user, requesting the user to record a selfie video while performing the dynamic challenge and verifying the selfie video against the dynamic challenge. The dynamic challenge may include a script including, for example text (e.g., specific words or phrases) that the user is requested to say and / or one or more movements or actions that the user is requested to perform (e.g., “move your head to the right”). If operations 410-440 are successful, then in operation 450 the processor may enroll the user on the avatar generation platform and generate an identity log to the user. The identity log may include a user ID, biometric identification data extracted from the selfie video of the enrolled user such as data indicative of a facial image of the user and / or data indicative of a voice of the user, and other data related to the user, that may enable the avatar generation platform to identify the user when the user logs into the avatar generation platform. If any of operations 410-440 is not successful, then the user may be refused and not enrolled on the system.
[0089] Reference is now made to Fig. 5, which is a flowchart of one embodiment of a method for generating an avatar and a video by an avatar generation platform, according to some embodiments. While in some embodiments the operations of Fig. 5 are carried out using systems as shown in Figs. 1 and 2, in other embodiments other systems and equipment can be used.
[0090] In operation 510, a processor (e.g., processor 105 depicted in Fig. 1 executing code to carry out the method for generating an avatar and a video according to embodiments of the present invention) may receive a request from a requester to generate an avatar for the requester. If analready verified selfie of the requester is used for generating the avatar (e.g., the selfie video uploaded by the requester in operation 440) then, as indicated in operation 520, the method may move to generate an avatar for the requester, operation 580. If, however, the requester wants to use a new selfie for the avatar, then in operation 530 the processor may receive biometric identification data of the requester, e.g., a selfie image, a voice recording and / or a selfie video uploaded by the requester.
[0091] The processor may verify the biometric identification data of the requester against the biometric identification data in the identity log of the enrolled user. In case the user selected to download and store the encrypted and digitally signed user identity log, the user would be required to upload the encrypted and digitally signed user identity log back into the system. For example, in operation 540, the processor may verify a selfie image of the requester (e.g., the selfie image uploaded in operation 530 or a frame of the selfie video depicting an image of the face uploaded in operation 530) against the identity log of the requester, e.g., the identity log generated in operation 450. For example, the identity log of the requester may be retrieved from identity log database 132 using the user ID of the requester (or uploaded by the user as disclosed herein), and the selfie image of the requester may be verified against the data indicative of the facial image of the user stored in the identity log. Verification of the selfie image of the requester against the data indicative of the facial image of the user stored in the identity log may include determining whether the selfie image of the requester and the data indicative of the facial image of the user stored in the identity log represent the same person. If so, the verification is successful, if not the verification fails.
[0092] In operation 550, the processor may verify a voice sample of the requester (e.g., the voice sample uploaded in operation 530 or a voice sample from the selfie video uploaded in operation 530) against the voice sample stored in the identity log. The voice sample of the requester may be verified against the data indicative of the voice of the user stored in the identity log. Verification of the voice sample of the requester against the data indicative of the voice of the user stored in the identity log may include determining whether the voice sample of the requester and the data indicative of the voice of the user stored in the identity log represent the same person. If so, the verification is successful, if not the verification fails.
[0093] In some embodiments, successful verification of the biometric identification data of the requester against the biometric identification data in the identity log of the enrolled user inoperation 560, requires that both the verifications in operation 540 and in operation 550 succeed. If one of the verifications in operation 540 and in operation 550 fails, then in operation 570 the request to generate an avatar is denied, no avatar is generated, and the process terminates. If both the verifications in operation 540 and in operation 550 are successful, then in operation 580 the processor may generate an avatar for the requester. In some embodiments, only one of operations 540 and 550 is performed, and successful verification of the biometric identification data of the requester against the biometric identification data in the identity log of the enrolled user in operation 560, requires that the verification in the one of operations 540 and 550 is successful. Other types of biometric data and verification processes may be used.
[0094] In operation 590, the processor may receive a voice script from the requester, in operation 592, the processor may generate a video including the avatar generated in operation 580 performing the voice script, optionally with corresponding lip movements and other movements and gestures. In operation 594, the processor may insert watermarks into the video generated in operation 592. The watermarks may be visible or invisible, and include data identifying the user, e.g., a user ID, data identifying the video, e.g., video ID, and / or data identifying the computerized avatar generation platform. The invisible watermarks may be inserted into the image part of the video and / or into the audio part of the video. In operation 596, the processor may provide the generated video to the user, in any applicable file format.
[0095] Reference is now made to Fig. 6, which is a flowchart of a method for conducting a complaint resolution process by an avatar generation platform, according to some embodiments. While in some embodiments the operations of Fig. 6 are carried out using systems as shown in Figs. 1 and 3, in other embodiments other systems and equipment can be used.
[0096] In operation 610, a processor (e.g., processor 105 depicted in Fig. 1 executing code to carry out the method for conducting complaint resolution according to embodiments of the present invention) may receive a complaint about a video that includes an avatar. The complaint may be filed by a complainant and may include the video that the complainant is complaining about, a description of the complaint, and a user ID of the complainant, in case the complainant is a registered user in the avatar generation platform. In case the user selected to download and store the encrypted and digitally signed user identity log, the user would be required to upload the encrypted and digitally signed user identity log back into the system.
[0097] In operation 612, the processor may determine whether the video was generated by the avatar generation platform. The processor may determine whether the video was generated by the avatar generation platform by extracting one or more watermarks (e.g., the watermarks inserted into the video in operation 594) from the video and extracting the data identifying the computerized avatar generation platform from the watermarks. If the processor is not able to extract the watermarks from the video, the processor may try to generate a representation of the video, and search for the representation in video log database 134. If the representation is found in log database 134, then the video is identified as generated by the avatar generation platform. If the processor is not able to extract the watermarks and the representation is not found in log database 134, then the processor may determine that the video was not generated by the avatar generation platform, and terminate the complaint, as indicated in operation 614. Terminating the complaint may include notifying the user of the reason why the complaint is terminated, e.g., that video was not generated by the avatar generation platform.
[0098] If the processor determines in operation 612 that the video was generated by the avatar generation platform, then in operation 616 the processor may attempt to extract the user ID of the user that generated the video. The user ID may be extracted directly from the watermarks, or indirectly by extracting the video ID from the watermarks, fetching the video log of the video from video log database 134, and extracting the user ID from the video log. If the processor is not able to extract the watermarks from the video, the processor may try to generate a representation of the video, and search for the representation in video log database 134. If the representation is found in log database 134, then the processor may retrieve the video log associated with representation from log database 134 and extract the user ID of the user that generated the video from the video log in operation 622. If the processor fails to extract the user ID, then the processor may terminate the complaint resolution process, as indicated in operation 614. The processor may notify the complainant that the complaint cannot be resolved since the user ID of the user that generated the video cannot be identified.
[0099] In operation 624, the processor may determine how to process the complaint based on the type of the complaint. If the complaint is about inappropriate content, then in operation 626, the processor may analyze the content of the video to determine if indeed the content is inappropriate. The analysis may be performed automatically using ML models, for example by prompting an LLM to identify inappropriate content and / or using dedicated ML models trained for that purpose.Additionally or alternatively, the processor may receive a determination on the content from a human reviewer. If the processor determines that the content is appropriate, then in operation 632 the complaint resolution process terminates. The processor may notify the complainant that the content in the video has been found to be appropriate. If the processor determines that the content is inappropriate, then in operation 682 the processor may take an action. For example, the processor may ban the user that has generated the video, the processor may notify a system administrator of the complaint, including the complaint details such complainant ID, the user ID, type of complaint, complaint resolution, etc.
[0100] If the complaint is about avatar misuse, then in operation 630, the processor may determine if indeed the avatar of the complainant was used by a different user. To determine a misuse, the processor may compare the avatar in the video to the avatar of the complainant. If the avatar in the video matches avatar of the complainant, then the processor may determine that the complaint is justified and take action, as indicated in operation 628. If the avatar in the video does not match the avatar of the complainant, then the processor may determine that the complaint is unjustified and the compliant resolution process may terminate, as indication in operation 632.
[0101] One skilled in the art will realize the embodiments may be embodied in other specific forms without departing from the spirit or essential characteristics thereof. The foregoing embodiments are therefore to be considered in all respects illustrative rather than limiting of the embodiments described herein. Scope of the embodiments is thus indicated by the appended claims, rather than by the foregoing description, and all changes that come within the meaning and range of equivalency of the claims are therefore intended to be embraced therein.
[0102] In the foregoing detailed description, numerous specific details are set forth in order to provide an understanding of the embodiments . However, it will be understood by those skilled in the art that the embodiments can be practiced without these specific details. In other instances, well-known methods, procedures, and components, modules, units and / or circuits have not been described in detail so as not to obscure the embodiments . Some features or elements described with respect to one embodiment can be combined with features or elements described with respect to other embodiments.
[0103] Although embodiments of the embodiments are not limited in this regard, discussions utilizing terms such as, for example, “processing,” “computing,” “calculating,” “determining,” “establishing”, “analyzing”, “checking”, or the like, can refer to operation(s) and / or process(es) ofa computer, a computing platform, a computing system, or other electronic computing device, that manipulates and / or transforms data represented as physical (e.g., electronic) quantities within the computer’s registers and / or memories into other data similarly represented as physical quantities within the computer’s registers and / or memories or other information non-transitory storage medium that can store instructions to perform operations and / or processes.
[0104] Reference is made to Fig. 7, which is a block diagram of an avatar generation platform 700, according to embodiments of the invention. It should be understood in advance that the components and functions shown in Fig. 7 are intended to be illustrative only and embodiments of the invention are not limited thereto. While in some embodiments the system of Fig. 7 is implemented using systems as shown in Fig. 1, in other embodiments other systems and equipment can be used. Avatar generation platform 700 may include enrolment module 210, avatar generation module 230, video generation module 250 and complaint resolution module 270, shown in Fig. 2, and complaint resolution module 380 shown in Fig. 3 and Identity log database 132, video log database 134 and avatar database 136 shown in Fig. 1.
[0105] Although embodiments of the embodiments are not limited in this regard, the terms “plurality” and “a plurality” as used herein can include, for example, “multiple” or “two or more”. The terms “plurality” or “a plurality” can be used throughout the specification to describe two or more components, devices, elements, units, parameters, or the like. The term set when used herein can include one or more items. Unless explicitly stated, the method embodiments described herein are not constrained to a particular order or sequence. Additionally, some of the described method embodiments or elements thereof can occur or be performed simultaneously, at the same point in time, or concurrently.
Claims
CLAIMS1. A method for secure generation of avatars using a computerized avatar generation platform comprising: requesting a user to record a selfie video with a dynamic challenge; verifying the selfie video against the dynamic challenge; enrolling the user on the computerized avatar generation platform and generating an identity log of the user if the verification of the selfie video is successful, and refusing to enroll the user on the computerized avatar generation platform otherwise, wherein the identity log comprises biometric identification data extracted from the selfie video of the enrolled user; receiving a request from a requester to generate an avatar of the requester, the request comprising biometric identification data of the requester; verifying the biometric identification data of the requester against the biometric identification data in the identity log of the enrolled user; and generating an avatar for the requester if the verification of the identity of the requester against the identity log of the enrolled user is successful, and denying the request otherwise.
2. The method of claim 1, wherein enrolling the user on the computerized avatar generation platform further comprises: authenticating the user by performing at least one of: a shared authentication scheme, email verification, two-factor authentication (2FA), multiple factor authentication (MFA), automated know your customer (KYC) verification and credit card verification; and receiving user consent.
3. The method of claim 1, wherein the dynamic challenge comprises at least one of: text that the user is requested to say and an action that the user is requested to perform.
4. The method of claim 3, wherein the biometric identification data in the identity log comprises at least one of: data indicative of the facial image of the enrolled user extracted from the selfie video and data indicative of the voice of the enrolled user extracted from the selfie video.
5. The method of claim 4, wherein verifying the biometric identification data of the requester against the biometric identification data in the identity log of the enrolled user comprises verifying a selfie of the requester against the data indicative of the facial image of the enrolled user.
6. The method of claim 5, wherein verifying the selfie of the requester against the data indicative of the facial image of the enrolled user is performed using an ML model.
7. The method of claim 4, wherein verifying the biometric identification data of the requester against the biometric identification data in the identity log of the enrolled user comprises verifying a voice recording of the requester against data indicative of the voice of the enrolled user.
8. The method of claim 7, wherein verifying a voice recording of the requester against data indicative of the voice of the enrolled user is performed using an ML model.
9. The method of claim 1, comprising: generating a video comprising the avatar; and inserting a watermark into the generated video, the watermark comprising data identifying the requester.
10. The method of claim 9, comprising: receiving a complaint associated with an examined video presenting an examined avatar; analyzing content of the examined video to determine whether the content of examined video is legitimate or illegitimate; in case of illegitimate content, extracting from the examined video data identifying a user that generated the examined video; and banning the user that generated the examined video.
11. The method of claim 1 , comprising: generating a video comprising the avatar; andinserting a watermark into the generated video, the watermark comprises data identifying the computerized avatar generation platform.
12. The method of claim 11, comprising: receiving a complaint associated with an examined video presenting an examined avatar; determining whether the examined video was generated by the computerized avatar generation platform based on the watermark; and handling the complaint only if the examined video was generated by the computerized avatar generation platform.
13. The method of claim 1, further comprising: generating a video comprising the avatar; generating video identifying data from the generated video; and storing the video identifying data in the computerized avatar generation platform.
14. The method of claim 13, wherein the video identifying data comprises at least one of: a video identification number embedded in a watermark inserted into the generated video, a hash of the generated video and an embedding of the video.
15. The method of claim 13, comprising: receiving a complaint associated with an examined video presenting an examined avatar; determining whether the examined video was generated by the computerized avatar generation platform based on the video identifying data; and handling the complaint only if the examined video was generated by the computerized avatar generation platform.
16. A computer-based method for secure generation of avatars, comprising: enrolling a user by: providing a dynamic challenge to the user; requesting the user to record a selfie video with the dynamic challenge; verifying the selfie video against the dynamic challenge; andenrolling the user and generating an identity log of the user if the verification of the selfie video is successful, and refusing to enroll the user if the verification of the selfie video fails; and receiving a request from a requester to generate an avatar of the requester; verifying the identity of the requester against an identity log of the enrolled user; and generating an avatar for the requester if the verification of the identity of the requester against the identity log of the enrolled user is successful, and denying the request if the verification of the identity of the requester against the identity log of the enrolled user fails.
17. A device comprising: a memory; and one or more processors configured to: request the user to record a selfie video with a dynamic challenge; verify the selfie video against the dynamic challenge; enroll the user on the computerized avatar generation platform and generating an identity log of the user if the verification of the selfie video is successful, and refusing to enroll the user on the computerized avatar generation platform otherwise, wherein the identity log comprises biometric identification data extracted from the selfie video of the enrolled user; receive a request from a requester to generate an avatar of the requester, the request comprising biometric identification data of the requester; verify the biometric identification data of the requester against the biometric identification data in the identity log of the enrolled user; and generate an avatar for the requester if the verification of the identity of the requester against the identity log of the enrolled user is successful, and denying the request otherwise.
18. The device of claim 17, wherein the one or more processors are configured to enroll the user on the computerized avatar generation platform by: authenticating the user by performing at least one of: a shared authentication scheme, email verification, two-factor authentication (2FA), multiple factor authentication (MFA), automated know your customer (KYC) verification and credit card verification; andreceiving user consent.
19. The device of claim 17, wherein the dynamic challenge comprises at least one of: text that the user is requested to say and an action that the user is requested to perform.
20. The device of claim 19, wherein the biometric identification data in the identity log comprises at least one of: data indicative of the facial image of the enrolled user extracted from the selfie video and data indicative of the voice of the enrolled user extracted from the selfie video.
21. The device of claim 20, wherein the one or more processors are configured to verify the biometric identification data of the requester against the biometric identification data in the identity log of the enrolled user by verifying a selfie of the requester against the data indicative of the facial image of the enrolled user.
22. The device of claim 21, wherein the one or more processors are configured to verify the selfie of the requester against the data indicative of the facial image of the enrolled user using an ML model.
23. The device of claim 20, wherein the one or more processors are configured to verify the biometric identification data of the requester against the biometric identification data in the identity log of the enrolled user by verifying a voice recording of the requester against data indicative of the voice of the enrolled user.
24. The device of claim 23, wherein the one or more processors are configured to verify a voice recording of the requester against data indicative of the voice of the enrolled user using using an ML model.
25. The device of claim 17, wherein the one or more processors are further configured to: generate a video comprising the avatar; and insert a watermark into the generated video, the watermark comprising data identifying the requester.
26. The device of claim 23, wherein the one or more processors are further configured to: receive a complaint associated with an examined video presenting an examined avatar; analyze content of the examined video to determine whether the content of examined video is legitimate or illegitimate; in case of illegitimate content, extract from the examined video data identifying a user that generated the examined video; and ban the user that generated the examined video.
27. The device of claim 17, wherein the one or more processors are further configured to: generate a video comprising the avatar; insert a watermark into the generated video, the watermark comprises data identifying the computerized avatar generation platform.
28. The device of claim 27, wherein the one or more processors are further configured to: receive a complaint associated with an examined video presenting an examined avatar; determine whether the examined video was generated by the computerized avatar generation platform based on the watermark; and handle the complaint only if the examined video was generated by the computerized avatar generation platform.
29. The device of claim 17, wherein the one or more processors are further configured to: generate a video comprising the avatar; generate video identifying data from the generated video; and store the video identifying data in the computerized avatar generation platform.
30. The device of claim 29, wherein the video identifying data comprises at least one of: a video identification number embedded in a watermark inserted into the generated video, a hash of the generated video and an embedding of the video.
31. The device of claim 29, wherein the one or more processors are further configured to: receive a complaint associated with an examined video presenting an examined avatar;determine whether the examined video was generated by the computerized avatar generation platform based on the video identifying data; and handle the complaint only if the examined video was generated by the computerized avatar generation platform.
Citation Information
Patent Citations
System to provide virtual avatars having real faces with biometric identification
CA2658174A1
Avatar management system, avatar management method, program, and computer-readable recording medium
US20230136394A1
Apparatus, system, and method for generating a video avatar
US20230237722A1
Techniques for verifying user identities during computer-mediated interactions
US20240046687A1
Avatar authenticity registration method, avatar authenticity registration system, expression data management system, and expression data management method
WO2023243623A1