Computer human assessment systems, devices, and methods for recreating characters with human actions
A system with real-time, randomly generated prompts and computer vision verification secures against deepfakes and personalizes content, addressing limitations of existing facial verification systems.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- KOYAL AI INC
- Filing Date
- 2025-11-11
- Publication Date
- 2026-05-15
AI Technical Summary
Existing facial verification systems are deterministic and can be spoofed by deepfakes or require expensive infrared cameras, limiting their accessibility and effectiveness in distinguishing humans from bots and personalizing generated content.
A system that uses real-time, randomly generated prompts for users to perform actions, verified by computer vision and machine learning models, ensuring live interaction and personalizing generative AI models with user data.
Enhances security against deepfakes by requiring real-time interaction, personalizes generated content, and is accessible on various devices without specialized hardware, improving human verification and personalization.
Smart Images

Figure US2025055026_15052026_PF_FP_ABST
Abstract
Description
COMPUTER HUMAN ASSESSMENT SYSTEMS, DEVICES, AND METHODS FOR RECREATING CHARACTERS WITH HUMAN ACTIONSCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit under 35 U.S.C. § 119 of U.S. Provisional Application Serial No. 63 / 718,984, filed on November 11, 2024, which is incorporated herein by reference.FIELD OF THE INVENTION
[0002] This invention relates to computer human assessments for recreating characters with human actions (“CHARCHA”).BACKGROUND OF THE INVENTION
[0003] The present invention comprises systems, devices, and methods that relate to computer human assessments for recreating characters with human actions. The systems, devices, and methods of the present invention comprise embodiments of a facial identity verification protocol that serve several purposes: (1) to verify that the user is human; (2) to verify the user's identity (confirm the user is who they claim they are) by having them perform a list of specific actions live on camera in a random order; and (3) to simultaneously collect these images to finetune generative artificial intelligence (“Al”) models to personalize generated images, videos, or other multimodal content. Other purposes will be apparent as well.
[0004] Existing facial verification and liveness checking systems, such as ID .me®, usually involve taking a photo of a person through a webcam and having the person rotate their head to capture a three-dimensional (“3D”) scan to ensure the person is live in front of the camera. But this is deterministic and can be spoofed by recording a video of another person rotating their head or a deepfake of a person trained on a video. However, the present invention introduces real-time randomness, which adds protection against these attacks since the person verifying their identityis doing so with informed consent. Additionally, a malicious actor cannot predict the next random action with the present invention, which makes it more difficult (if not impossible) for a prerecorded video to pass the present invention’s test. Successfully passing the test normally will require submitting a live camera feed of a real person taking the test with informed consent (in many embodiments of the invention.)
[0005] Other facial verification systems, such as Apple Face ID®, use infrared to generate a 3D facial map of the user during verification to make sure the user is live in front of them. However, this may be expensive to incorporate for many systems since it uses special cameras that are only available on selected phones, such as iPhones®, and are not readily / easily accessible on PCs, laptops, and many other phone brands. Therefore, technologies like Face ID® are not universal and only work on certain types of devices. Moreover, these systems also do not help capture diverse head and face poses for fine tuning generative Al models for personalization.
[0006] Other sources of deepfake detection often are done post-facto to check if a video is Al -generated or real. This is often a cat-and-mouse game because, as deepfake detection gets better, deep fakes generation also becomes better to beat the systems. However, embodiments of the present invention perform real-time checks for detecting and verifying a person’s identity and that are hard to beat using current deepfake generation since good deepfakes take time to generate.
[0007] Thus, embodiments of the present invention stand out through its use of live-action prompts that require the user to respond in real time, reducing the risk of impersonation. While the present invention differentiates humans from bots based on tasks and differs from face identification which relies on infrared technology available only on select devices, the invention can be configured to be device-agnostic and accessible via standard cameras on personal computers (“PCs”), smartphones, tablets, or other electronic devices. This adaptability providesthe invention with a broader scope of uses and greater scalability across diverse devices and platforms. Additionally, while deepfake detection methods often work reactively to detect manipulated content, embodiments of the present invention proactive thwart manipulation by incorporating randomness in its prompts, making it difficult for Al -driven impersonations to keep pace in real time.BRIEF SUMMARY OF THE INVENTION
[0008] One embodiment of the present invention is a method of verifying whether a user is human, comprising the steps of: (1) assigning a unique identifier to a user of an electronic device wherein the user has a head and a body; (2) verifying that the user’s head is within the frame of a camera connected to the electronic device; (3) providing a plurality of prompts to the user for the purpose of having the user comply with the prompts; (4) recording the user’s compliance with each prompt; (5) comparing the recorded compliance to a trained model to determine whether the user correctly complied with each prompt and to develop a score for the user’s compliance; (6) comparing the score to a predetermined threshold score; and (7) creating a file of the recordings to train an Al model on the user.
[0009] Another embodiment of the present invention is a system for verifying whether a user is human, comprising: a user electronic device connected to a video camera and to a web browser; a server-side processing unit configured to verify a user’s identification and to manage verification workflow, wherein the server-side processing unit incorporates at least one computer vision module and at least one machine learning model for real-time facial detection and pose estimation; a verification engine using trained neural networks for pose recognition and comparison and configured with computational resources for real-time video processing; and a database configured to store unique identifiers and maintain user verification status, wherein theelectronic device is configured to have video processing capabilities record, compress, and store image / video data, with storage infrastructure for potential Al model training and wherein the system integrates through API endpoints facilitating client-server communication, with security protocols protecting sensitive user data during transmission and storage.
[0010] One embodiment of the present invention is a method of verifying whether a user of an electronic device is a live human, wherein the human has a head and a body and wherein the electronic device comprises at least one processor and at least one memory, the method comprising the steps of: (i) verifying that at least the user’s head is within a frame of a camera connected to the electronic device; (ii) providing at least one prompt to the user for a purpose of having the user comply with the at least one prompt; (iii) recording a user’s compliance with the at least one prompt by capturing an image of the user with the camera; (v) evaluating the captured image with a trained model to determine the user’s compliance with the prompt and to generate a compliance score; (v) comparing the compliance score to a predetermined threshold score; and (vi) creating a file of the captured image. Another embodiment of a method also comprising repeating the steps of providing at least one prompt, recording a user’s compliance, evaluating the captured image, comparing the compliance score, and creating a file for a predetermined number of times. For one embodiment of a method the at least one prompt is randomly chosen from a library of poses. For one embodiment of a method the trained model is a vision-language model. One embodiment of a method also comprises performing a liveness check on the captured image and evaluating the captured image for spoofing signals. One embodiment of a method also comprises an initial step of assigning a unique identifier to the user and a final step of using the file to train an Al model (the same or a different Al model) on the user and to generate a personalized avatar of the user. One embodiment of a method entails thestep of verifying that at least the user’s head is within a frame also comprising verifying that the user’s body is within the frame.
[0011] Another embodiment of the present invention is a system for verifying that a user of an electronic device is a human and for generating an avatar. For this embodiment, the system comprises the following: (i) an electronic device having a processor, an interconnected memory, and a camera connected to the electronic device, wherein the process is configured to perform the following steps (a) assigning a unique identifier to a user of an electronic device, wherein the user has a head and a body; (b) verifying that the user’s head is within a frame of a camera connected to the electronic device; (c) providing at least one randomly selected prompt to the user for a purpose of having the user comply with the prompt; (d) recording a user’s compliance with the prompt by capturing an image of the user with the camera; (e) evaluating the captured image by a trained model to generate a compliance score; (f) comparing the compliance score to a predetermined threshold score; (g) creating a file of the captured image to train an Al model on the user; and (h) generating an avatar of the user from the trained Al model. The embodiment also comprises at least one non-transitory storage medium connected to the processor. Another embodiment of a system of the present invention also comprises repeating the steps of providing at least one prompt, recording a user’s compliance, comparing the captured image, comparing the compliance score, and creating a file for a predetermined number of time. Another embodiment of a system also has the at least one prompt chosen from a library of poses. For another embodiment the trained model is a vision-language model. For another embodiment the trained model checks for at least one liveness indicators and for at least one anti- spoofing signal. For another embodiment the Al model trained on the user is configured to generate a personalized avatar of the user.
[0012] Another embodiment of the present invention is a method of creating a personalized video from an audio file comprising generating a personalized avatar with an electronic device connected to a camera and having a processing unit and memory, the processing unit configured to execute the steps of: (i) assigning a unique identifier to a user, having a head and a body, of the electronic device; (ii) verifying that the user’s head is within a frame of the camera; (iii) providing a plurality of random prompts to the user for a purpose of having the user comply with the prompts; (iv) recording a user’s compliance with each prompt by capturing an image of the user with the camera; (v) comparing the captured images to a trained model and scoring the recorded compliance to generate a compliance score; (vi) comparing the compliance score to a predetermined threshold score; and (vii) creating a file of each of the captured images to train an Al model on the user, wherein the Al model is configured to generate the personalized avatar of the user based upon the captured images. Also, this embodiment comprises inputting an audio file into at least one model trained to generate visual content from the audio file that is contextually aligned with and synchronized to the audio file to create an animated reproduction of the audio file and inserting the personalized avatar into the animated reproduction of the audio file. For another embodiment of a method the at least one model trained to generate visual content is selected from the group consisting of audio transcription, text-to-image diffusion, linear spherical interpolation, and music emotion recognition. Another embodiment comprises music emotion recognition for predicting valence and arousal of the audio file. Another embodiment comprises using a large language model to extract emotion from the audio file. For another embodiment one of the at least one model trained to generate visual content is a latent diffusion model. For another embodiment the audio file is selected from the group consisting of music, podcast, narration, dialog, and a voicerecording. For another embodiment the captured images are processing using low -rank adaptation.
[0013] One embodiment of the present invention is a system for verifying whether a user is human, comprising: (i) a user electronic device connected to a video camera and to a web browser; (ii) a server-side processing unit configured to verify a user’s identification and to manage verification workflow, wherein the server-side processing unit incorporates at least one computer vision module and at least one machine learning model for real-time facial detection and pose estimation; (iii) a verification engine using trained neural networks for pose recognition and comparison and configured with computational resources for real-time video processing; and (iv) a database configured to store unique identifiers and maintain user verification status. For this embodiment the electronic device is configured to have video processing capabilities record, compress, and store image / video data, with storage infrastructure for potential Al model training and wherein the system integrates through API endpoints facilitating client -server communication, with security protocols protecting sensitive user data during transmission and storage.BRIEF DESCRIPTION OF THE SEVERAL VIEWS OF THE DRAWINGS
[0014] For the purpose of facilitating understanding of the invention, the accompanying drawings and description illustrate embodiments thereof, from which the invention, various embodiments of its structures, construction and method of operation, and many advantages, may be understood and appreciated. The accompanying drawings are hereby incorporated by reference.
[0015] Figure 1 is one embodiment of an interface for a system of the present invention;
[0016] Figures 2A through 2D illustrate various captured frames for different actions of one embodiment of the present invention;
[0017] Figure 3 illustrates one embodiment of one application of the invention for personalized Al music-video creation;
[0018] Figure 4 illustrates one embodiment of an audio-to-video pipeline according to the invention;
[0019] Figure 5 illustrates example images generated based on select lyrics according to one embodiment of the present invention;
[0020] Figures 6A and 6B illustrate one valence / arousal emotion spectrum and one serenemelancholy spherical interpolation for use with the present invention;
[0021] Figure 7 illustrates the results of a survey of an experiment on the present invention;
[0022] Figure 8 is a table of face verification metrics for one embodiment of the present invention;
[0023] Figure 9 is a graph of the CLIP Similarity between images generated by one embodiment of the present invention and generated video frames of one embodiment of the present invention;
[0024] Figure 10 is one embodiment of an arousal / valence prediction model according to one embodiment of the present invention;
[0025] Figures 11A through 11B combine to illustrate is one embodiment of a music-to- video model architecture;
[0026] Figure 12 shows one embodiment of a method of user verification of the present invention;
[0027] Figure 13 shows one schematic of a system of the present invention; and
[0028] Figure 14 show another embodiment of a user interface with prompts according to the present invention.DETAILED DESCRIPTION OF THE INVENTION
[0029] The following describes example embodiments in which the present invention may be practiced. This invention, however, may be embodied in many different ways, and the descriptions provided herein should not be construed as limiting in any way. Among other things, the following invention may be embodied as methods, systems, or devices. The following detailed descriptions should not be taken in a limiting sense. The accompanying drawings are hereby incorporated by reference.
[0030] Before the example embodiments of the systems, devices, non-transitory storage media, and methods according to the present disclosure are disclosed and described below, it is to be understood that embodiments are not limited to those described within this disclosure. Numerous modifications and variations therein will be apparent to those skilled in the art and remain within the scope of the disclosure. It also is to be understood that the terminology used herein is forthe purpose of describing specific embodiments only and is not intended to be limiting. Some embodiments of the disclosed technology will be described more fully hereinafter with reference to the accompanying drawings. This disclosed technology, however, may be embodied in many different forms and should not be construed as limited to the embodiments set forth therein.
[0031] If the specification states a component, element, part, or feature “may,” “can,” “could,” or “might” be included or have a characteristic, then that particular component or feature is not required to be included or have the characteristic.
[0032] In the following description, numerous specific details are set forth. However, it is to be understood that embodiments of the disclosed technology may be practiced without these specific details. In other instances, well-known methods, structures, and techniques have not beenshown in detail in order not to obscure an understanding of this description. References to “one embodiment,” “an embodiment,” “example embodiment,” “some embodiments,” “certain embodiments,” “various embodiments,” etc., indicate that the embodiment(s) of the disclosed technology so described may include a particular feature, structure, or characteristic, but not every embodiment necessarily includes that particular feature, structure, or characteristic. Further, repeated use of the phrase “in one embodiment” does not necessarily refer to the same embodiment, although it may.
[0033] Unless otherwise noted, the terms used herein are to be understood according to conventional usage by those of ordinary skill in the relevant art. In addition to any definitions of terms provided below, it is to be understood that as used in the specification and in the claims, the terms “a” or “an” are used, as is common in patent documents, to include one or more than one. In this document, the term “of” is used to refer to a nonexclusive “of” such that “A or B” includes “A but not B,” “B but not A,” and “A and B,” unless otherwise indicated. Furthermore, all publications, patents, and patent documents referred to in this document are incorporated by reference herein in their entirety, as though individually incorporated by reference. In the event of inconsistent usages between this document and those documents so incorporated by reference, the usage in the incorporated reference(s) should be considered supplementary to that of this document; for irreconcilable inconsistencies, the usage in this document controls.
[0034] The following description has set forth aspects of computer system or computer- implemented devices and / or processes / methods via the use of block diagrams, flowcharts, and / or examples, which may contain one or more functions and / or operations. As used herein, the term or graphic of a “block” in the block diagrams and flowcharts refers to a step of a computer- implemented process executed by a computer system, which may be implemented as a machinelearning (“ML”) system, an assembly of machine learning systems, artificial intelligence (“Al”) models or Al systems. Each block can be implemented as either a machine learning system or as a nonmachine learning system, according to the function described in association with each particular block. Furthermore, each block can refer to one of multiple steps of a process embodied by computer-implemented instructions or programs executed by a computer system (which may include, in whole or in part, a machine learning system) or an individual computer system (which may include, e.g., a machine learning system) executing the described step, which is in turn connected with other computer systems (which may include, e.g., additional machine learning systems) for executing the overarching process described in connection with each figure or figures.
[0035] The terms “connected”, “interconnected”, “in communication”, or “coupled” and related terms are used in an operational sense and are not necessarily limited to a direct physical connection or coupling. As an example, two or more devices, databases, websites, or platforms may be coupled directly, or via one or more intermediary channels, communication channels, or devices. They may be hardwired to each other or connected without hardwiring, such as by wi-fi, Bluetooth®, or cellular service. As another example, devices, databases, websites, or platforms may be coupled in such a way that information can be passed between them, while sharing or not sharing any physical connection with one another. Based on the disclosure provided herein, one of ordinary skill in the art will appreciate a variety of ways in which connection or coupling exists in accordance with the aforementioned definitions.
[0036] The various embodiments of the present invention can incorporate or be configured to run on one or more computing systems, which can include one or more processors) (e.g., central processing units (“CPUs”), graphical processing units (“GPUs”), holographic processing units(“HPUs”), etc.)(collectively, a “processing unit 20.”) Processors can be a single processing unitor multiple processing units in a device or distributed across multiple devices (e.g., distributed across two or more of computing devices).
[0037] The various methods and systems of the present invention can be configured to run on a processor-based computing system that includes one or more central processing units, each including one or more processors. The CPU(s) can be a master device and can have a cache memory coupled to the processors) for rapid access to temporarily stored data. The CPU(s) can be coupled to a system bus and can intercouple master and slave devices included in a processorbased system. As known in the art, the CPU(s) can communicate with other devices by exchanging address, control, and data information over the system bus. For example, the CPU can communicate bus transaction requests to a memory controller as an example of a slave device. Additionally, multiple system buses can be provided, wherein each system bus constitutes a different fabric.
[0038] Computing system(s) can include one or more input devices that provide input to the processors, notifying them of actions. The actions can be mediated by a hardware controller that interprets the signals received from the input device and communicates the information to the processors using a communication protocol. Each input device can include, for example, a mouse, a keyboard, a touchscreen, a touchpad, a wearable input device (e.g., a haptics glove, a bracelet, a ring, an earring, a necklace, a watch, etc.), a camera (or other light-based input device, e.g., an infrared sensor), a microphone, or other user input devices.
[0039] Processors can be coupled to other hardware devices, for example, with the use of an internal or external bus, such as a PCI bus, SCSI bus, or wireless connection. The processors can communicate with a hardware controller for devices, such as for a display. Display can be used to display text, images, and graphics. In some implementations, display includes the inputdevice as part of the display, such as when the input device is a touchscreen or is equipped with an eye direction monitoring system. In some implementations, the display is separate from the input device. Examples of display devices include the following: an LCD display screen , an LED display screen , a projected, holographic, or augmented reality display (such as a heads-up display device or a head-mounted device), and so on. Other input / output (“I / O”) devices can also be coupled to the processor, such as a network chip or card, video chip or card, audio chip or card, USB, firewire or other external device, camera, printer, speakers, CD-ROM drive, DVD drive, disk drive, etc.
[0040] Computing system can include a communication device capable of communicating wirelessly or wire-based with other local computing devices or a network node. The communication device can communicate with another device or a server through a network using, for example, TCP / IP protocols. Computing system can utilize the communication device to distribute operations across multiple network devices. A “communication channel,” as referred to herein, is any communication device, hardware, and / or software (including but not limited to a cable, Wi-Fi, internet, cloud, network, server, or combination thereof) that facilitates or enables a first computing system to communicate with a second computing system. The first computing system and the second computing system are referred to herein generally as one or more computing systems and are configured with the components necessary to run the programs, method, and / or modules identified for a particular computing system to achieve a predetermined output or goal.
[0041] The processors can have access to a memory or storage medium (generally, “storage 26), which can be contained on one of the computing devices of computing system or can be distributed across of the multiple computing devices of computing system or other external devices. A memory or storage medium includes one or more hardware devices for volatile or non-volatile storage and can include both read-only and writable memory. Storage 26 includes non- transitoiy computer-readable storage media. For example, a memory can include one or more of random-access memory (“RAM”), various caches, CPU registers, read-only memory (“ROM”), and writable non-volatile memory, such as flash memory, hard drives, floppy disks, CDs, DVDs, magnetic storage devices, tape drives, and so forth. Memory can include a non-transitory computer-readable storage medium storing one or more programs for generating a 3D avatar model of a user, the one or more programs comprising instructions, which, when executed by at least one processor of an electronic system, cause the electronic system to perform the methods and processes described herein. A memory is not a propagating signal divorced from underlying hardware; a memory is thus non-transitory. Memory can include program memory that stores programs and software, such as an operating system, a local physical environment modeling application, and other application programs. Memory can also include data memory that can include eyeprint content, preconfigured templates for password generation, hand gesture patterns, configuration data, settings, user options or preferences, etc., which can be provided to the program memory or any element of the computing system.
[0042] Some implementations can be operational with numerous other computing system environments or configurations. Examples of computing systems, environments, and / or configurations that may be suitable for use with the technology include, but are not limited to, virtual reality headsets, personal computers, server computers, handheld or laptop devices, cellular telephones, wearable electronics, gaming consoles, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments that include any of the above systems or devices, or the like.
[0043] Network can be a local area network (“LAN”), a wide area network (“WAN”), a mesh network, a hybrid network, or other wired or wireless networks. Network may be the Internet or some other public or private network. Computing devices can be connected to network through a network interface, such as by wired or wireless communication. While the connections between parts, components, modules, and servers are shown as separate connections, these connections can be any kind of local, wide area, wired, or wireless network, including network or a separate public or private network.
[0044] In some implementations, an analysis engine executed by the virtual reality device or a remote system that is receiving images from the virtual reality device can automatically identify features of interest in the primary user's environment andean identify them for the primary and / or second user. For example, the analysis engine can include machine learning models trained to identify damage to particular types of objects, where the models can be trained using pictures from previously verified insurance claims. As another example, the analysis engine can automatically compare images previously submitted by the primary user (e.g., pictures of particular objects) to new images to identify differences (e.g., that may indicate damage). The indications from the analysis engine can include directions to the primary user to focus on the identified locations in the primary user's local environment or indications to the second user, for the second user to provide the instructions to the primary user.
[0045] Other master and slave devices can be connected to the system bus. These devices can include a memory system, one or more input devices, one or more output devices, one or more network interface devices, and one or more display controllers, as non-limiting examples. The input device(s) can include any time of device, including but not limited to input keys, switches, voice processors, etc. The output device(s) can include any type of output device including, butnot limited to, audio, video, other visual indicators, etc. The network interface device(s) can be configured to support any type of communications protocol desired. The memory or memory system can include one or more memory units.
[0046] The CPU(s) can be configured to access the display controllers) over the system bus to control information sent to one or more displays. The display controller(s) sends information to the display(s) to be displayed via one or more video processors, which process the information to be displayed into a format suitable for the display(s). The display(s) can include any type of display, including, but not limited to, a cathode ray tube, a liquid crystal display, a plasma display, a light emitting diode display, a virtual reality device, and / or an augmented reality device. The information sent by the display can include information on a viewer’s selected viewpoint from which to display the 3D avatar model and related information for displaying the 3D avatar model to the viewer.
[0047] The processor-based system(s) or processors can be provided in an integrated circuit. The memory system or memory may include a memory array(s) and / or memory bit cells. The processor-based system can be provided in a system-on-a-chip.
[0048] Those of skill in the art will further appreciate the various illustrative logical blocks, modules, circuits, and algorithms described in connection with the aspects disclosed herein may be implemented as electronic hardware, instructions stored in memory or in another computer readable medium and executed by a processor or other processing device, or a combination of both. The master devices and slave devices described herein may be employed in any circuit, hardware component, integrated circuit, or integrated circuit chip, as examples. Memory disclosed herein may be any type and size of memory and may be configured to store any type of information desired. To clearly illustrate this interchangeability, various illustrative components, blocks,modules, circuits, and steps have been described above generally in terms of their functionality. How such functionality is implemented depends upon the particular application, design choices, and / or design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure.
[0049] The various illustrative logical blocks, components, modules, and circuits described in connection with the aspects disclosed herein may be implemented or performed with a processor, a digital signal processor, an application specific integrated circuit, a field programmable gate array, or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A processor may be a microprocessor, any conventional processor, controller, microcontroller, or state machine. A processor can be implemented as a combination of computing devices. “Component” and “module” are used herein to refer to the hardware and the software, respectively, to achieve a goal and are used interchangeably herein. It will be obvious to one stilled in the art that, a “module” is a part of a process or method defined by its goal our output and includes the software, code, programs, etc. to achieve that goal or output. A “component” generally includes all necessary hardware configured to run or execute a “module”.
[0050] The aspects disclosed herein can be embodied in hardware and in instructions that are stored in hardware, and can reside, for example in random access memory, flash memory, read only memory, electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of computer readable medium. In the alternative, the storage medium can be integral to the processor. The processor and the storage medium can reside in an application-specific integrated circuit (“ASIC”). TheASIC can reside in a remote station. Alternatively, the processor and the storage medium can reside as discrete components in a remote station, base station, or server.
[0051] Those of skill in the art will understand that information and signals can be represented using any of a variety of different technologies and techniques. For example, data, instructions, commands, information, signals, bits, symbols, and chips that are referenced throughout this description can be represented by voltage, currents, electromagnetic waves, magnetic fields or particles, optical fields or particles, or any combination thereof.
[0052] Various systems, non-transitory storage media-based systems, and methods (also referred to as “process(es)”) of the present invention accept as input a set of RGB images capturing the subject’s appearance and are further compatible with RGB-D data comprising both color and depth information. The present invention relates to systems and methods for real-time generation of drivable, photorealistic three-dimensional digital twins or avatars of human subjects using monocular or sparse multi-view visual inputs. The systems and methods accept as input a set of RGB images capturing the subject’s appearance and are further compatible with RGB-D data comprising both color and depth information.
[0053] In the context of photography and imaging, an "image" refers to a single visual representation captured by a camera. A "frame", on the other hand, typically denotes a single still image or a single image within a sequence of images, such as those that make up a video stream. For example, a photograph is an individual image, while a video consists of multiple frames shown in rapid succession to create the illusion of motion. For simplicity, the term "image" is used herein to refer to either a standalone photograph or a single frame extracted from a video stream. In one embodiment, the captured data is comprised of RGB or RGB-D images or frames and, optionally,depth information, calibration information, and / or other information gathered or provided by the imaging device.
[0054] The present disclosure encompasses all changes, substitutions, variations, alterations, and modifications to the example embodiments herein that a person having ordinary skill in the art would comprehend. Similarly, where appropriate, the appended claims encompass all changes, substitutions, variations, alterations, and modifications to the example embodiments herein that a person having ordinary skill in the art would comprehend. By way of example, while embodiments of the present invention have been described as operating in connection with a social networking website, the present invention can be used in connection with any communications facility that supports web applications. Furthermore, in some embodiments the term “web service” and “website” may be used interchangeably and additionally may refer to a custom or generalized software on a device, such as a mobile device (e.g., cellular phone, smart phone, personal GPS, personal digital assistance, personal gaming device, etc.), that makes calls directly to a server.
[0055] The present invention comprises various embodiments of systems 1000, methods 2000, and non-transitoiy storage media 3000 for implementing Computer Human Assessment for Recreating Characters with Human Actions (“CHARCHA”) that comprise facial identity verification protocols that serve multiple purposes: (1) verifying that the user is human; (2) verifying the user's identity (confirming the user is who they claim they are) by having them perform a list of specific actions Eve on camera in a random order; and (3) simultaneously collects these images to finetune generative artificial intelligence (“Al”) models to generated personalized avatars, images, videos or other multimodal content.
[0056] Various embodiments of CHARCHA systems 1000, methods 2000, and non- transitory storage media 3000 not only safeguard individual privacy but also enhancepersonalization by directly integrating user-provided data (the user’s images and / or voice.) Various embodiments of CHARCHA systems 1000, methods 2000, and non-transitory storage media 3000 verify the user's identity by prompting real-time actions which include a variety of facial expressions and head poses but can also include body poses and actions as well. Various embodiments of CHARCHA systems 1000, methods 2000, and non-transitory storage media 3000 are designed to be accessible to consenting humans yet challenging for machines and malicious actors, which can prevent deepfakes, stealing someone’s likeness, spoofing, and other malicious behaviors.
[0057] These novel authentication systems 1000, methods 2000, and media 3000 move beyond previous verification systems to take advantage of the benefits of personalization while minimizing risk of identity manipulation and misuse. However, embodiments of the present invention not only help distinguish humans from robots / machines but also distinguish consenting humans from malicious actors who might be stealing a person’s likeness for misuse with generative Al technology.
[0058] While traditional identity checks can be static and vulnerable to spoofing, the multipurpose protocols of various embodiments of CHARCHA systems 1000, methods 2000, and non- transitory storage media 3000, with random live-action prompts, add a layer of complexity and adaptability. Each verified interaction collects rich facial, body, and / or voice data in real time, which is then used to fine tune generative Al models to recreate the user’s likeness or voice, promoting both privacy and personalization. An interaction is “verified” when a user 1 successfully completes the prompts 13 at the threshold scores 16. By merging identity verification with personalization, the invention provides a novel solution for secure, consent-based personalization in sectors such as entertainment, e-commerce, and media arts.
[0059] Previous authentication protocols present users with tasks that are challenging for machines but simple for humans. They have evolved from simple human-machine differentiation to sophisticated variants like reCAPTCHA & hCAPTCHA. The advent of generative Al further blurred the lines between human and machine-generated content. While generative Al offers unprecedented personalization opportunities, it also facilitates impersonation through deepfakes and other text-to-image models and text-to-image models finetuned to generate celebrities through Low Rank Adaptations (“LoRAs”). However, CHARCHA comprises novel authentication methods 2000, media 3000, and systems 1000 to take advantage of the benefits of personalization while minimizing risk of identity manipulation and misuse.
[0060] One embodiment of a CHARCHA method 2000 of verifying a human user comprises the following steps (as shown in Figures 12 and 14.) First, for embodiments that involve the creation of an avatar 36, a unique identifier is assigned 200 to a user 1 (having a head 2 and body 3) of an electronic device 10. Then, the method 2000 verifies 210 that the user’s head 2 is within the frame 11 of a camera 12 connected to the electronic device 10 (also referred to herein as “calibration” or the “calibration step 210”.) The method 2000 then involves randomly selecting 211 one or more prompts 13 to present to the user 1. The user 1 is given 220 one or more prompts 13 for the purpose of having the user 1 comply with the prompts 13. The user’s compliance 15 with each prompt 13 is recorded 230. The recorded compliance is compared to an Al trained model 240 to determine whether the user 1 correctly complied with each prompt 13 and to develop a score 14 for the user’s compliance 15. The score 14 is compared 250 to a predetermined threshold score 16. Then a file 17 is created 260 of the recordings to train an Al model 18 on the user 1. Consecutive captured images 33 can be compared as a liveness check 270 to ensure that the user 1 is live and in front of the camera 12. It is known in the art to compare the capturedimage of the user and the background 37 for consistency as one method of verifying liveness 270. The user 1 can be prompted a predetermined number of times, which number is determined by several factors including but not limited to the number of images desired to create an avatar 36; the level of confidence necessary for the liveness check 270, the evaluation of antispoofing signals 35 (or to look for indicators of spoofing), or to confirm the user’s identity 2200.
[0061] Various embodiments of CH ARCH A methods 2200 verify the user's identity by randomly prompting 220 them to enact real-time actions 13 in front of a webcam / camera / video recording device 12 which includes a variety of facial expressions, head poses (such as turn your head left or right, tilt head up or down, squint your eyes, open your mouth, smile with teeth, look straight), and / or body poses to ensure diversity and authenticity (collectively, prompts 13 or actions 13.) After the user 1 performs the action 13, a vision language model finetuned to recognize users performing the actions receives both the original text prompt 13 and the captured video frames 33, and then analyzes whether the user 1 actually performed that specific action 13 while simultaneously checking for liveness indicators 34 (nonlimiting examples include natural muscle movements, proper depth, skin texture, temporal consistency) and anti-spoofing signals 35 (nonlimiting examples include detecting printed photos, screen replays, masks, or deepfakes). In one embodiment, the images 33 obtained from CHARCHA systems 1000, methods 2000, and media 3000 can then be used to fine tune Al model(s) 18 to generate a personalized generative Al content 36 of the consenting user 1. To clarify, a method that verifies a user’s identity 2200 is configured according to any of the methods 2000 herein but comprises a vision language model finetuned to recognize the user 1. This is an optional step in the methods discussed herein.
[0062] In some embodiments, CHARCHA systems 1000, methods 2000, and non- transitory storage media 3000 protocols are designed to be simple and fast (a minute or so in oneembodiment). Following a simple calibration 210 of the head position, users 1 are asked to perform an action or pose 220 (used interchangeably herein) picked at random from the set of actions for a specified length of time (in one example, a few seconds.) An image is taken 230 of the user 1 executing the action or prompt 13. This process is repeated for a predetermined number of actions 13 or prompts 13 (ideally, without duplicates) with a few seconds of gap for users 1 to prepare before the next action 13 (see Figure 12.)
[0063] In one embodiment, while the user 1 is performing the actions 13, the user’s performance of the action is verified in real-time and a compliance score 14 is recorded for each action 13. To pass, a user 1 must successfully complete all actions 13 within the appropriate threshold score 16. If a user 1 fails to perform the actions 13 successfully, CHARCHA verification is considered to have failed (see Figure 12.)
[0064] Additionally, in some embodiments, CHARCHA systems 1000, methods 2000, and non-transitory storage media 3000, there also is a liveness check 270 running in the background to make sure the user 1 is live in front of the camera 12. In one embodiment, a liveness check or liveness detection 270 entails checking that the background 37 and the user 1 remain consistent across frames 11, verifying that the background 37 does not change unnaturally, for example, ensuring there are no edge artifacts where a photo or screen meets the real background, and confirming that movements create appropriate spatial relationships between the person and their surroundings. Vision-Language Models can significantly improve this by receiving both the challenge prompt ("smile and tilt head left") and video frames, then performing comprehensive analysis that goes beyond basic background checks: simultaneously verifying the specific prompted action was performed correctly while also analyzing skin texture patterns to detect printed photos, natural micro-movements like eye blinks that occur 15-30 times per minute,temporal frame consistency to catch video replays, subtle blood flow patterns through rPPG analysis (remote photoplethysmography analysis), depth cues from head rotation, proper lighting reflections on living tissue versus flat surfaces, and environmental coherence including parallax shifts, uniform lighting effects, and absence of manipulation artifacts. The verification and the liveness checks 270 are done and can be improved with the use of advanced computer vision algorithms. Furthermore, the images 33 collected from the user 1 can be randomly verified to ensure accuracy and to finetune the machine learning detection, more specifically, captured images can be used to better train and tune the machine learning models for the various embodiments of the present invention.
[0065] If the liveness verification 270 is successful, various embodiments of the present invention use the collected photos / videos 33 from user 1 performing different poses 13 in front of the camera 12 as input to a low rank adaptations technique of fine tuning generative Al model(s) 18 to create personalized high-quality and more realistic generated content for the user 1, which can include but is not limited to avatars and voice representations.
[0066] While traditional verification systems help distinguish humans from robots and machines, systems, devices, and methods incorporating CHARCHA are specifically designed as a robust identity verification system in an age where generative Al has made identity spoofing and deep fake generation more universally accessible. Various embodiments of systems 1000, methods 2000, and media 3000 of CHARCHA go beyond user verification done by traditional, existing technologies by exploiting generative Al's current inability to convincingly replicate randomized, real-time physical micro-movements. According to current statistics, with deepfake fraud surging 1,100% and human detection accuracy at just 24.5%, static biometric checks and passive liveness detection are easily fooled by real-time face-swapping tools and pre-recordedvideos. CHARCHA's algorithmic analysis of unpredictable movements during a short interactive session ensures the user 1 is physically present and responding in real-time, not a simulation or replay. This creates a consent-first barrier that prevents unauthorized likeness capture before content generation, rather than trying to detect deepfakes after they are created: addressing the fundamental problem that deepfake generation is advancing 900% annually while detection capabilities consistently lag behind.
[0067] In some embodiments, CHARCHA systems 1000, methods 2000, and non- transitory storage media 300 comprise the following four steps. First, CHARCHA assigns a unique identifier 200 to the user being verified. This step of assigning a unique identifier 200 is not required if the method 2000 is being used to verify that a user 1 is human nor if the method 2200 is being used to verify the identity of the user 1. However, assigning a unique identifier 200 is used in one embodiment in which the method 2000 is used to generate a personalized avatar 36. Second, CHARCHA calibrates 210, which, in various embodiment, means CHARCHA verifies that the user’s head is in the camera frame. Third, CHARCHA verifies the user is live (liveness check 270) by recording the user following prompts for poses or actions and then checking the pose or action for compliance (steps 212, 220, 230, 240, and 250.) The liveness checks 270 verify, among other things, that a real person is in front of the camera at that moment in time and, as such, consents to be captured as an image. The combination of random prompts and short time frames between prompts deter or prevent a user from manipulating the system to capture or use the image of another person without consent. This verification comprises comparing the user’s pose or action to a trained model. Fourth, CHARCHA creates a file of the images or videos captured that can (in some embodiments) be used to train an Al model on the user 250.
[0068] One embodiment of the CHARCHA verification system 1000 comprises an electronic device 10 with video camera 12 (input device 12 such as a webcam, smartphone camera, integrated laptop camera, etc.) connected to a web browser 19 accessing the CHARCHA method 2000 (see Figure 13.) A server-side processing unit 20 handles user identification and manages verification workflow, incorporating computer vision modules 21 and machine learning models 18 for real-time facial detection 210 and pose estimation during the calibration phase 210. In one embodiment, the system's verification engine 22 uses trained neural networks for pose / action recognition and comparison, requiring computational resources for real-time video processing. A secure database 23 stores unique identifiers and maintains user verification status. Video processing components 24 record, compress, and store image / video data, with storage infrastructure 26 forpotential Al model training. In one embodiment, the system 1000 integrates through API endpoints 25 facilitating client-server communication, with security protocols protecting sensitive user data during transmission and storage 26.
[0069] For some embodiments of the present invention, thecalibration step 210 determines the position of a user’s head 2 in front of the camera 210 (the position of the body 3 can be ascertained for some embodiments as well.) An image 33 of the user 1 with their head 2 in neutral position 31 is captured 212. Following the calibration, users 1 are asked toperform an action 13 picked at random from the set of actions 13 within a relatively short period of time 220. Various embodiments of the CHARCHA systems 1000, methods 2000, and non-transitory storage media 300 of the present invention record an image 33 of the user 230. Then the recorded image 33 is analyzed to determine if the user 1 successfully executed it 240. This recorded image 33 also can be a video for themodel 18 to see if the user 1 is successfully following the action 13. This process is repeated for a plurality of actions 13 (often without duplicates) with a short gap in between eachaction 13 for users 1 to prepare before the next action 13. While the user 1 is performing the actions 13, the actions 13 are verified in real-time 240 and a score 14 is recorded for each action13 out of ten (or any assigned value). To pass, a user 1 must successfully complete all specified actions 13 within the bounds of a threshold score 16 (which can vary by application) 250. This threshold score 16 is determined experimentally. In one embodiment, if the user 1 fails, they can retake the test once more (or a specified number of times). If this fails, CHARCHA verification is considered to have failed and the face cannot be used for generation for this embodiment.
[0070] Figure 1A illustrates one embodiment of a CHARCHA interface 30 that can be accessed or used on any electronic device 10. Again, nonlimiting examples of electronic devices 10 include mobile phones, computers (desktop or laptop), electronic tablet devices, smart TVs, game consoles or any device alone or when connected to a video recorder or camera enables a user to capture video or photographic images of themselves for verification by the systems, devices, and methods of the present invention (collectively, an “electronic device 10”). Figure 1A also illustrates one embodiment of the verification method 2000 of the present invention in which three random poses 13 are captured, (i) turn head left, (ii) open your mouth, and (iii) turn head right. Various embodiments, CHARCHA systems 1000, methods 2000, and non-transitoiy storage media 300 of the present invention take the captured images 33 in Figure 1 A and map them against a library or model of poses 32 (see Figure IB) to determine accuracy or compliance (a true / false determination in one embodiment). Figures 2A through 2D illustrate four images 33 captured by the present invention in which the user 1 is attempting to comply with or execute the prompts 13. As mentioned previously, some embodiments of the invention compare the captured image 33 to the corresponding pose 31 from a pose library 32. However, other embodiments can be configured to use one or more trained Al models to answer a question of whether the user 1 complied withthe prompt 13 (i.e., did the user 1 turn his / her head to the left 13) without having to map the captured image 33 to the corresponding pose 32 from the library 33.
[0071] Robust Generation. Moreover, in some embodiments, CHARCHA systems 1000, methods 2000, and non-transitory storage media 300, in terms of generating an avatar, the photos 33 and videos 33 taken during verification also help train an Al model 18 on the user’s likeness offering a diverse set of face and head poses 31. Unlike other generative Al tools like "Imagine Me" from Meta Al, which captures three selfies of a person from different angles, the larger number of screenshots 33 captured by embodiments of the present invention provide for more diverse features of a user 1 like their smile or stance while standing and to generate expressive videos. As shown in the results, this allows the present invention to create videos with greater resemblance to the user and more diversity in the generation process. Importantly though, existing technologies do not combine the random selecting of a prompt 211 and the liveness check 270 (or use of a liveness indicator 34) to both ensure that the avatar 36 is based upon a Eve, consenting human user 1 and to generate an avatar 36 that resembles the user 1.
[0072] Problems Solved and Advantages. To summarize, some embodiments of CHARCHA systems 1000, methods 2000, and non-transitory storage media 300 address several pressing challenges in digital identity verification and content personalization. They mitigate risks associated with impersonation, revenge pornography, and misuse of personal likenesses by adding an informed -consent layer to generative Al model training. In an embodiment, CHARCHA's novelty lies in reversing how consent works with Al training data. Currently, Al models scrape billions of pieces of data from the internet without asking permission, forcing people to discover misuse after it happens and try to opt out retroactively. CHARCHA flips this by requiring active, physical verification before someone's likeness can be used to train Al models. Users mustperform a predetermined number of randomized physical actions in front of a camera to prove they are really consenting, not just clicking a checkbox buried in terms of service. This prevents bad actors from using scraped photos or videos to create deepfakes or train models on someone's appearance without their knowledge. The system protects against impersonation and revenge pornography by making it technically impossible to use someone's likeness without their conscious, real-time participation, while also creating a verifiable record that can be used as legal proof in right of publicity cases. Rather than relying on legal agreements that are hard to enforce, CH ARCH A makes consent a technical requirement that cannot be bypassed. This proactive measure allows users 1 to control how their likeness is used and ensures that any Al -generated content accurately reflects their consent and participation. Additionally, the CHARCHA systems 1000, methods 2000, and non-transitory storage media 300 various methods of collecting diverse facial expressions and head poses enhance the quality of digital avatars, providing a more authentic and expressive representation for users. Various embodiments of the present invention also differ from existing verification tools because they do not rely on static puzzles or behavior tracking, which modem Al can already bypass. Instead, these embodiments verify a real human in real time through simple live prompts (like small facial or head movements), and then uses that verified, user-consented identity to enable safe personalization. This makes embodiments of the present invention useful in sectors like entertainment, education, healthcare, e-commerce, and social media where both privacy and personal expression matter. Unlike traditional CAPTCHAs that only check “human vs bot,” CHARCHA also ensures the person chooses to share their likeness, so the system can personalize experiences (like characters in videos or avatars in learning tools) without storing or misusing personal data.
[0073] Additional advantages of various embodiments of the present invention that combat impersonation and revenge pornography include that the actions 13 are randomly given to the user 1 in runtime and verified in a short time (in one embodiment, 7 seconds), so it is hard for anyone to pull up images in that time period to substitute for their own likeness. Moreover, there are checks for liveliness (whether the person is Eve in front of the camera - checking for static background etc.) and it is the same person doing all actions.
[0074] Furthermore, by overcoming the limitations of standard facial verification methods - such as susceptibility to static image-based spoofing or the high-cost barrier of 3D scanning equipment - embodiments of thepresent invention democratize access to advanced, secure identity verification. Its real-time liveness checks 270 and dynamic prompts 13 help maintain high security standards while offering a unique advantage in creating consent -based, personalized Al -generated content across sectors like virtual gaming, online retail, and personalized media.
[0075] Uses and Preferred Embodiments of the Invention. Various embodiments of the CHARCHA systems 1000, methods 2000, and non-transitoiy storage media 300 comprise multi-function protocol that allow for seamless integration into sectors requiring high privacy standards and sectors that value user personalization, including but not limited to entertainment, education, healthcare, e-commerce, and social media.
[0076] This advanced form of facial verification allows users to provide informed consent for user-personalization in generative Al applications. These applications can include creation of digital avatars in business or recreational settings (Al generated headshots or videogame characters as non-limiting examples). This can also be used to create photo or video content based on the person’s likeness for a variety of applications.
[0077] Following this protocol before the creation of personalized Al generated content will ensure that the content generated of the user is directly approved by the person. Any content created without use of identity verification for example with CHARCHA protocol can deftly be considered improper and fake and therefore subject to scrutiny. This can drastically reduce the rampant public misinformation campaigns especially for sensitive issues and involving public personalities.
[0078] Moreover, CHARCHA can also replace existing facial identity and liveness detection such as id .me for sensitive applications (such as banking) which require users to confirm they are who they claim they are, due to the added protection of real-time randomness.
[0079] CHARCHA allows users to provide informed consent for personalization in a variety of applications and industries. Some specific use cases where CHARCHA can enable secure and ethical Al generated content are as mentioned below:Entertainment and Media• Digital avatars for gaming and virtual reality: Game designers and creators can use embodiments of CHARCHA systems 1000, methods 2000, and storage media 3000 to ethically create lifelike avatars for immersive experiences in video games or virtual environments.• Virtual actors and digital doubles: Studios can create digital versions of actors, for stunts, or de-aging in films, while ensuring the consent -based capture of likenesses for ethical use with embodiments of CHARCHA systems 1000, methods 2000, and storage media 3000.Personalized content creation: By capturing a user’s unique expressions, embodiments of CHARCHA systems 1000, methods 2000, and storage media3000 facilitate the creation of customized video and photo content, allowing users to see their own likeness in interactive storytelling or music videos.E-commerce and RetailVirtual Try-Ons and Fitting Rooms:• Embodiments of CHARCHA systems 1000, methods 2000, and storage media 3000 s ability to produce personalized digital avatars provides users with realistic virtual try-on experiences for clothing, makeup, or accessories, enhancing online shopping and reducing returns.• Creating virtual fitting rooms for a more immersive online shopping experience using embodiments of CHARCHA systems 1000, methods 2000, and storage media 3000.Personalized Product Recommendations: Retailers can use digital likenesses to provide personalized shopping suggestions, allowing customers to visualize products as they would appear on themselves, enhancing user engagement and conversion rates.Personalized Learning Experiences:• images captured by embodiments of CHARCHA systems 1000, methods 2000, and storage media 3000 can be used to create Al tutors or instructors resembling familiar figures to enhance engagement, especially in settings requiring a sense of trust and familiarity.• Developing tailored educational content featuring the learner's likeness (as created by embodiments of CHARCHA systems 1000, methods 2000, and storage media 3000) for improved retention.Simulation and Role-Playing for Training:• For industries like healthcare or military, embodiments of CHARCHA systems 1000, methods 2000, and storage media 3000 can generate realistic avatars for simulation-based training, enhancing situational readiness by mimicking real- world scenarios with user likenesses.Healthcare• Patient Education: avatar creation by embodiments of CHARCHA systems 1000, methods 2000, and storage media 3000 can provide personalized visual aids that use a patient’s likeness to explain medical procedures or treatments, improving comprehension and patient outcomes.• Mental Health and Therapy: Customizable avatars created by embodiments of CHARCHA systems 1000, methods 2000, and storage media 3000 can create familiar and comfortable digital spaces for mental health support or counseling, fostering better engagement in virtual therapy sessions.Marketing and Advertising• Personalized Advertisements: Advertisers can develop targeted campaigns using avatars tailored to each consumer and do so by ethically sourcing the user’s images through embodiments of CHARCHA systems 1000, methods 2000, and storage media 3000, allowing for greater interaction and emotional resonance in marketing materials.• Interactive Campaigns: embodiments of CHARCHA systems 1000, methods 2000, and storage media 3000’s consent-based, likeness-preserving avatars are ideal for interactive marketing, enabling users to see themselves in different campaign contexts and engaging them in a novel, immersive way.Social Media and Communication• Enhanced Video Calls and Virtual Meetings: embodiments of CHARCHA systems 1000, methods 2000, and storage media 3000 can improve video call experiences by generating high-resolution or animated avatars, making virtual meetings more interactive and reducing camera-related fatigue.• Personalized Emojis and Stickers: Social media platforms can use embodiments of CHARCHA systems 1000, methods 2000, and storage media 3000 to create custom emojis and stickers that reflect users’ actual expressions, adding a new level of personalization to digital communication.Gaming and Entertainment• Character Customization: Gamers can use embodiments of CHARCHA systems 1000, methods 2000, and storage media 3000 to create lifelike, detailed avatars based on their own appearance, enhancing the in-game experience and making it more personal.• Dynamic NPC (Non Player Characters) Creation: Game developers can populate environments with NPCs resembling real individuals, adding realism and engagement for a better immersive gaming experience through character creation using embodiments of CHARCHA systems 1000, methods 2000, and storage media 3000.Customer Service: Personalized Virtual Assistants: Businesses can create highly personalized Al chatbots or virtual assistants that use avatars created by embodiments of CHARCHA systems 1000, methods 2000, and storage media 3000 to increase customer comfort and engagement, blending Al functionality with a personal touch.Tourism and Hospitality: Virtual Tours and Hotel Previews: avatars created via embodiments of CHARCHA systems 1000, methods 2000, and storage media 3000 can enable travelers to visualize themselves in various locations, such as hotels or tourist destinations, through virtual tours or immersive digital previews with the guest’s digital twin, enhancing the travel booking experience.
[0080] Various embodiments of CHARCHA systems 1000, methods 2000, and storage media 3000 can address issues such as:1. Privacy and Consent IssuesUnauthorized Use of Likeness:• Creating deepfakes or Al-generated content using someone's likeness without their permission can violate privacy rights.• Misusing personal data to generate highly personalized content without proper consent.Identity Theft and Impersonation: Al-generated content could be used to impersonate individuals for fraudulent purposes or to spread misinformation.2. Misinformation and ManipulationPolitical Propaganda:• Using Al to create convincing fake videos or speeches of political figures to spread disinformation or manipulate public opinion.Fake News and Hoaxes• Generating realistic but false news stories or images featuring real people to mislead the public.3. Legal and Ethical ConcernsCopyright and Intellectual Property: Recreating a person's likeness in Al -generated content may infringe on their intellectual property rights or violate publicity rights.4. Economic ImpactJob Displacement: In industries like entertainment or modeling, Al -generated likenesses could potentially replace human actors or models, leading to job losses.
[0081] Various embodiments of CHARCHA systems 1000, methods 2000, and storage media 3000’s unique combination of real-time identity verification with Al -driven avatar creation are positioned to meet both immediate and emerging needs across various industries. Its capability to balance security, personalization, and user consent is what we believe will make it a cornerstone technology in the evolving landscape of digital identity and personal representation in virtual spaces.
[0082] Commercial Market for the Invention and Potential Competitors andCustomers. Various embodiments of CHARCHA systems 1000, methods 2000, and storage media 3000 are poised to address a wide-ranging commercial market that spans industries focused on digital identity verification, personalized Al content creation, and enhanced user experiences in virtual environments. The demand for secure, consent-based facial verification is growing rapidly, especially as sectors like entertainment, education, e-commerce, healthcare, and government seek robust solutions to counter the risks of identity fraud, deepfake impersonation, and privacy concerns.
[0083] Primary Commercial Markets:Digital Identity Verification and Security• Financial Institutions and E-commerce: Banks, insurance companies, and online retailers increasingly need secure onboarding and anti-fraud mechanisms for digital transactions.• Government Agencies and Law Enforcement: Government sectors requiring secure access control, digital voting, and sensitive document verification• Healthcare Providers: For secure patient identification, record access, and telemedicine• Personalized Content and Media Creation• Entertainment and Media Companies: In industries like film, television, and gaming• Marketing and Advertising Firms: Personalized marketing campaigns and targeted advertisements• Virtual and Augmented Reality Platforms• Gaming and Metaverse Companies: Virtual reality (VR) and augmented reality (AR) platforms• Social Media and Communication Platforms: Social media applications that incorporate virtual avatars, custom emojis, and stickers• Retail and E-commerce• Virtual Fitting and Product Visualization: virtual outfits, accessories, and makeup, realistic product previews, & personalized product recommendations.• Education and Training• Interactive Education: Al-powered tutors or avatars resemble familiar figures in digital education settings.Simulation-based Training: military, healthcare, and customer service training for improved skill development.• Competitors / Future Customers:• Facial Verification and Identity Security Providers: Companies like ID .me®, Cognitec, Sensory, and Apple (Face ID®) offer facial recognition and liveness detection. However, CHARCHA differentiates itself by incorporating real-time action prompts and dual functionality for avatar generation, adding an additional layer of privacy and consent -focused verification. CHARCHA can thus also augment then- existing verification methods, and they can serve as customers too.• Deepfake Detection Solutions: Companies developing deepfake detection algorithms play a role in identifying manipulated media but are often limited to post -facto detection. CHARCHA, however, proactively combats impersonation by using live- action prompts that make it challenging for deepfakes to mimic in real time. CHARCHA can thus also augment their existing verification methods, and they can serve as customers too.
[0084] Potential Customers. Media and entertainment: Video production companies, music companies, and gaming studios can benefit from the various embodiments of CHARCHA systems 1000, methods 2000, and storage media 3000 for secure and ethical digital avatar creation. Generative Al and avatar creation platforms: Platforms like Meta®, Google®, OpenAI, Pika® Labs, Runway®, Synesthesia®, and klingai® focus on generating generative Al based text-to-image and video content. These applications could expand into personalization through the use of embodiments of CHARCHA systems 1000, methods 2000, and storage media 3000. CHARCHA’s advantage lies in its combination of secure identity verification and rich,consent-driven personalization for avatar creation, addressing privacy and ethical concerns in ways traditional generative platforms do not. Thus, CHARCHA can be integrated into these platforms. Social media and virtual reality companies can leverage embodiments of CHARCHA systems 1000, methods 2000, and storage media 3000 to enhance personalization and security within their immersive environments. E-commerce retailers can integrate embodiments of CHARCHA systems 1000, methods 2000, and storage media 3000’s avatar features for virtual try-ons and personalized shopping experiences, helping to improve user satisfaction and reduce returns.
[0085] Secure & Personalized Audio-to-Video Generation via CHARCHA. One non-limiting use case for the CHARCHA systems 1000, methods 2000, and media 3000 is for secure and personalized audio-to-video generation. This can include any audio-to-video constructions including, but not limited to, the creation of videos from music, podcasts, narration, dialog, any sort of recording reading or speaking. Various types of audio, including music, are deeply personal experiences and one goal is to enhance these experiences with a fully automated pipeline for personalized audio-to-video generation. One use case for CHARCHA systems 1000, methods 2000, and media 3000 enables listeners to not just be consumers but cocreators in the video generation process by creating personalized, consistent, and context-driven visuals based on lyrics, rhythm, and emotion in the audio. One non-limiting use case for CHARCHA systems 1000, methods 2000, and media 3000 is an audio-to-video generation pipeline 4000 that combines multiple modalities- audio, visual, and language and efficiently generates personalized and secure audio-based videos. This approach offers a new way to experience audio, allowing users 1 to participate in creating videos while protecting against impersonation in the age of generative AL
[0086] Traditionally, producing videos requires significant resources, but one embodiment of a CHARCHA audio-to-video pipeline 4000 extends the existing multimodal diffusion-based models beyond text and image, aligning audio with visual synchronization. By exploring the temporal connections between these modalities, CHARCHA audio-to-video pipeline 4000 aims to provide novel techniques and evaluation metrics that could have broader applications for multimodal video generation. Unlike previous works that have only explored the relationships between audio and video or text and video, embodiments of CHARCHA audio- to-video pipeline 4000 take a more integrated approach.
[0087] Figure 3 illustrates one nonlimiting, example embodiment of this application of a CHARCHA audio-to-video pipeline 4000 for the creation of a music video in which image stills and lyrics from generated music videos for Rick Astley’s "Never Gonna Give You Up" are matched with character references from CHARCHA. As shown in Figure 3, a character reference is a single captured image 33 used as the basis to create the avatar 36. Large language models can respond to prompts to change the background, tone, style, and emotion of the scene around the avatar 36. Various embodiments of an audio-to-video pipeline 4000 of the present invention comprise taking a character reference or captured image 33 and creating a personalize avatar 36. As shown in Figure 3, the avatar 36 can be inserted into an animated reproduction 38 of an audio file (for this example, a song). As shown in Figure 3, the Al models used can prepare animated reproductions 38 of the audio file 39 and modify the animation style according to the user’s preference. The videos use Queratogray Sketch, Western Animation Diffusion, and Realistic Vision V5.1 checkpoint models. The embodiment illustrated in Figure 5 shows image generation by one embodiment of the present invention based on the lyric "I just wanna tell youhow I’m feeling", which progressively incorporates LLM conditioning, negative prompting, style prompting, and emotion prompting.
[0088] Figure 4 illustrates one embodiment of an audio-to-video pipeline 4000 of the present invention that can be implemented with a variety of known software programs and Al models, as explained more fully herein. Initially, an audio file 39 is the input to start the pipeline 400. The audio file 39 is then processed to create a speech-to-text transcript 405, to detect vocal emotion 410, and to identify the speakers) 415. The information from those three processes enables a scene breakdown 420 and the creation of a narrative outline 425. Optionally, the narrative outline 425 can undergo manual review and revisions 430 and then be sent back to the scene breakdown step 420. Characters are then generated for the roles in the narrative outline 425. The characters can be Al generated 435 or, at least one character, can be generated via the CHARCHA systems 1000 and methods 2000, 2200 of the present invention. The CHARCHA systems 1000 and methods 2000 can be used to generate a personalize avatar 36 of the user 1, which can be inserted with the Al generated characters into the character roster 445. The characters can undergo styling 450 according to an Al model and / or the user’s preferences. The pipeline 4000 then generates visual storyboards 455. The user 1 can edit the videos and the composition 460 until they are satisfied and a final rendered video 38 is generated 470. This rendered video 38 contains the personalized avatar 36 created by way of the CHARCHA systems and methods.
[0089] With only the music audio file as input, music video pipeline 4000 (“MVP 4000 ”) uses a zero-shot approach to extract the rhythm, melody, lyrics, and emotional context from the audio to generate visual content that is contextually aligned and synchronized with the music. Various embodiments of CHARCHA ’s automated pipeline 4000 integrate a range ofpretrained Al models and APIs for tasks such as audio transcription, text-to-image diffusion, linear spherical interpolation, and music emotion recognition. Additionally, various embodiments of CHARCHA audio-to-video pipeline 4000’s framework also are designed to be versatile, accommodating various music genres and languages, and incorporating region-specific visual styles through diffusion based models. Thus, various embodiments of CHARCHA audio- to-video pipeline 4000’s make it easy for users, even with minimal technical expertise, to get started with generating their own music videos.
[0090] A key innovation in this approach is the inclusion of personalization (by way of the avatar 36), which allows listeners to become co-creators in the music video generation process. By finetuning image generation models on the user’s images, an embodiment of CHARCHA ’s audio-to-video pipeline 4000 can incorporate their likeness into the videos, creating a more immersive and personalized experience. By combining these multimodal techniques, an embodiment of CHARCHA’s audio-to-video pipeline 4000 creates a framework that not only advances the state of the art in music video generation but also prioritizes security and personalization.
[0091] Music (or audio) to video generation currently is done in closed source projects employing text-to-image models in the backend. Neural Frames’ Al music generator uses close source image generation models, trained on 2.7 billion images, to generate video frames from text prompts. The platform employs stem extraction and audio-reactive features to synchronize visuals with music, allowing up to ten visual parameters to be modulated by audio stems. None of these closed source solutions, however, have automatic semantic understanding and consider emotional information. Various embodiments of the present invention run separate trainedmodels to understand valence and arousal from music for understanding emotion and LLMs for understanding semantic context.
[0092] Music Emotion Recognition (“MER”) and Speech Emotion Recognition (“SER”) technologies developed a deep neural network for music mood detection using audio spectrograms and lyric embeddings from 18,000 annotated tracks. They found mid -level fusion optimal for bimodal valence and arousal prediction. Various embodiments of the present invention employ a similar model that has been trained in MER and SER to evaluate spoken words for tone and emotion, predicting valence and arousal from openSMILE features (see Figures 6A and 6B.) An embodiment of CHARCHA’s audio-to-video pipeline 4000 substitutes the emotion extraction from lyric embeddings by directly feeding the lyrics to an LLM in image prompt generation. Figure 6 A illustrates a valence / arousal emotion spectrum. Figure 6B illustrates a serene-melancholy spherical interpolation.
[0093] With respect to the embodiments of CHARCHA’s audio-to-video pipeline 4000 that perform text to image and video modeling, Latent Diffusion Transformer Models (“LDTMs”) uses pretrained autoencoder’s latent space to train diffusion models for image synthesis, reducing computational costs while maintaining quality. LDTMs also achieve competitive performance in text-to-image tasks. Models such as Kontext extend this by allowing flexible stylization through finetuning and LoRA training.
[0094] Text-to-video advancements have been driven by closed -source companies like Pika, Luma Labs, RunwayML, and OpenAI, using diffusion transformer models. Open-source alternatives like Wan 2.5 and Hunyuan Video exist but have limitations in video length (16 seconds) and quality. Thus, one embodiment of CHARCHA’s audio-to-video pipeline 4000 uses a combination of open source and closed source models, as are known in the field, togenerate image frames to form a complete storyboard and agentically orchestrates different models based on the context of each scene to generate video.
[0095] For various embodiments of CHARCHA’s audio-to-video pipeline 4000 to be used with music to video uses, CHARCHA can use an ensemble of pretrained models such as OpenAI’s automatic speech recognition (“ASR”) model Whisper to obtain the lyrics at specific timestamps. Some embodiments also comprise passing the audio into a trained music emotion recognition model.
[0096] Given the lyrics and the corresponding emotion at that timestamp, one embodiment of CHARCHA’s audio-to-video pipeline 4000 uses a pretrained text-to-image model finetuned on the user’s images to generate contiguous images based on the lyrics syncing them in a video.
[0097] An embodiment of CHARCHA’s audio-to-video pipeline 4000 can condition the image prompts using Large Language Model (“LLM”) conditioning, and additional negative prompts and style prompts. A few of these modeling considerations are discussed in more detail herein.
[0098] Some embodiments of CHARCHA’s audio-to-video pipeline 4000 comprise emotion extraction. One way of doing this is to use Music Emotion Recognition to analyze music’s emotional content using the circumplex model of affect, which captures the ebb and flow of feeling through two key dimensions: arousal (the intensity of emotion) and valence (its positivity or negativity). This creates four emotional quadrants: Melancholy, Serene, Tense, and Euphoric (see Figure 6A). One embodiment of CHARCHA’s audio-to-video pipeline 4000 accomplishes this as follows:
[0099] 1. Extract openSMILE features from audio.
[0100] 2. Use a neural network trained on the DE AM dataset to predict arousal and valence.
[0101] 3. Track emotional changes by monitoring position shifts in the arousal-valence space.
[0102] 4. Combine emotional data with lyrics to generate image prompts via an LLM.
[0103] One embodiment of CHARCHA’s audio-to-video pipeline 4000 comprises LLM Conditioning wherein the system, device, and / or method captures the timestamps with every change in lyric or emotion. Such an embodiment can leverage a language model to transform song lyrics and rhythm sections into rich, nuanced image prompts. Then the set of timestamps with the corresponding lyric and emotion is fed to the LLM, prompting it to come up with storybased image prompts that preserve the narrative (see Figure 5). This method overcomes the limitations of literal lyric / emotion interpretation, creating visually evocative descriptions that capture the song’s essence. For example, the lyric “I just wanna tell you how I’m feeling” becomes “The speaker stands beneath a night sky, gazing into their lover’s eyes, conveying the overwhelming magnitude of their feelings” (see Figure 5).
[0104] For on-set strength guided spherical interpolation, one embodiment of CHARCHA’s audio-to-video pipeline 4000 comprises seamless videos generated using the latest image-to-video models such as Seedream v4 and Google® Veo 3.1, which synchronizes visual changes with emotional and lyrical shifts. For audio-reactive animation, an embodiment of CHARCHA’s audio-to-video pipeline 40000 employs an intelligent agent that autonomously selects the most suitable generative model based on contextual evaluation across multiple categories and sub-categories of the audio input 39. For example, in one embodiment, the agent can identify slow-motion sequences and dynamically route them through a model optimized fortemporal coherence (such as a variant of Wan 2.5), while high-energy or rapid -transition segments may be processed using models specialized for fast motion synthesis. This adaptive orchestration enables rapid transitions during beat -heavy segments and smooth, deliberate evolution during quieter passages, resulting in a cohesive and rhythmically synchronized video output.
[0105] In one embodiment, fine-tuning stable diffusion for specific styles can be achieved by CHARCHA’s audio-to-video pipeline 4000 ability to leverage pre-existing style customization models from Civitai fine-tuned on specific art / photos to capture the different styles like (i) sketch, (ii) western animation, and (iii) hyper realistic (see Figure 5).
[0106] One Non-Limiting Example of Audio-to-video Model Architecture. One embodiment of the CHARCHA’ audio-to-video pipeline 4000 runs on one input, an MP3 file of the music. This is then broken down into multiple modalities for the purpose of generating a music video. This embodiment works as per the model diagram illustrated in Figures 11 A and1 IB, which illustrate a single model but could not be fit to show detail on a single page. Style of the video is changed by loading a different fine-timed Flux Kontext checkpoint. The CHARCHA audio-to-video pipeline 4000 images are used to finetune latest image generation models to generate a specific avatar 36 of the user 1.
[0107] For various embodiments of the stylization through checkpoint models, the present invention displays image stills from generated music video of the song "Flowers" by Miley Cyrus (see Figure 4). There is significant difference from baseline with the checkpoint models along with good adherence to the lyric. Thus, the qualitative evaluation supports the consensus that style consistency is strongly dependent on checkpoint models rather than simple prompt engineering.
[0108] For one evaluation of CHARCHA Images and Generated Video Frames, the same song (“Never Gonna Give You Up” by Rick Astley), can be used with identical prompts to generate the music video, only changing the loaded Character finetuned on an individual’s CHARCHA captured images. The Realistic Vision checkpoint model can be used to aim for photo-realism.
[0109] For a similarity analysis of one embodiment of the present invention, for each participant, the CLIP image-to-image similarity between the reference images and the generated video frames for every second of a minute of video was calculated. The scores were averaged across each CHARCHA image captured for that individual (termed the character similarity score). In Figure 9, the average similarity represents the mean of seven different character similarity scores at each timestep (averaged over seven individuals). The character similarity scores of the participants with maximum and minimum deviation also is shown.
[0110] The dynamic nature of this graph (Figure 9), with its fluctuations, indicates that the video does not simply display images from the training data (the CHARCHA images). If that were the case, the graph would be very close to 1 and not exhibit such variability. However, there is evidence that the character is faithfully represented in the generated video, as the character similarity score ranges from 0.6 to 0.9. This is notably higher than what would be expected when comparing the CHARCHA images to a video generated using a different character.
[0111] For one embodiment of the present invention, openSMILE features were extracted from the raw audio of the music and train a neural network model on the DEAM dataset to predict arousal and valence values from openSMILE features (see Figure 10). A validation mean-squared-error of 0.0206 was achieved. For dynamically changing emotion, aunique arousal- valence pair for each 5-second window of the music and corresponding openSMILE features was predicted. These values were used to continually update a position in the arousal-valence space by keeping a running sum and register a change in emotion whenever the position crosses into a new quadrant.
[0112] To achieve beat alignment for one embodiment of the present invention, the lyrictimestamp pairs used to generate images in the music video are taken from an ensemble of different ASR models, which may not be a 100% accurate and lacks consideration for the music’s rhythm and beats. To enhance this, the aim was to align the images with the song’s beats by extracting Predominant Local Pulse (“PLP”) information from the music using a sinusoidal kernel using the librosa library. This alignment not only creates a more immersive experience but also conveys shifts in mood and emotion, making the music video more impactful. By harmonizing visuals with the music’s rhythm, we estab Esh a sense of continuity and flow throughout the video, enhancing viewer engagement and meaning.
[0113] Confidence Testing. Figure 7 illustrates the results of a survey of CHARCHA experiment with n=16 participants. In this test, participants were asked to get creative and to break CHARCHA. Some examples of what the participants tried including: trying to perform several different actions at once, wearing sunglasses or head coverings, occlude the webcam or using someone else’s images in front of the webcam to try and pass the test. Figure 7 provides the results of this testing.
[0114] To measure the reproducibility of characters, Al -generated videos were evaluated to determine how accurately they could reproduce participants’ appearances from CHARCHA protocol images. The process involved training a Dreambooth LoRA model using CHARCHA user images and then generating videos with the Stable Diffusion 1.5 Realistic Vision modelusing the same prompts but different character LoRAs. To analyze reproducibility, the images from seven participants’ video frames were compared to original CHARCHA images. The comparison utilized OpenCV for face detection and VGG-Face for face verification (which shows competitive accuracy using the DeepFace library). This method allowed quantitative assessment of Al’s ability to recreate participant likenesses in generated videos (the average of the results for all 7 generated videos is shown in the table in Figure 8).
[0115] Here, % of face frames with participants face = no. of frames with participants face total frames - no. of frames with no face x 100CLIP image similarity scores were compared between one embodiment of CHARCHA ‘s audio- to-video pipeline 4000 images and video frames in the appendix to verify that LoRA training did not simply replicate training data. The observed score fluctuations throughout the video support our qualitative findings of diverse character settings, as seen in Figure 8.
[0116] In summary, the various embodiments of the present inventions’ commercial market spans industries that value secure, personalized, and consent -based identity verification and digital content creation. Its ability to address privacy concerns while facilitating secure, personalized digital experiences makes it a unique and versatile tool with widespread applicability across sectors where trust, personalization, and privacy are paramount.
[0117] While the disclosure has been described in detail and referring to specific embodiments thereof, it will be apparent to one skilled in the art that various changes and modifications can be made without departing from the spirit and scope of the embodiments. Thus, it is intended that the present disclosure covers the modifications and variations of this disclosure provided they come within the scope of the appended claims and their equivalents.
Claims
CLAIMS1. A method of verifying whether a user of an electronic device is a live human, wherein the human has a head and a body and wherein the electronic device comprises at least one processor and at least one memory, the method comprising: verifying that at least the user’s head is within a frame of a camera connected to the electronic device; providing at least one prompt to the user for a purpose of having the user comply with the at least one prompt; recording a user’s compliance with the at least one prompt by capturing an image of the user with the camera; evaluating the captured image with a trained model to determine the user’s compliance with the prompt and to generate a compliance score; comparing the compliance score to a predetermined threshold score; and creating a file of the captured image.
2. The method of Claim 1, also comprising repeating the steps of providing at least one prompt, recording a user’s compliance, evaluating the captured image, comparing the compliance score, and creating a file for a predetermined number of times.
3. The method of Claim 1, wherein the at least one prompt is randomly chosen from a library of poses.
4. The method of Claim 1, wherein the trained model is a vision -language model.
5. The method of Claim 1, also comprising performing a liveness check on the captured image and evaluating the captured image for spoofing signals.
6. The method of Claim 1 also comprising: an initial step of assigning a unique identifier to the user; anda final step of using the file to train an Al model on the user and to generate a personalized avatar of the user.
7. The method of Claim 1, wherein the step of verifying that at least the user’s head is within a frame also comprising verifying that the user’s body is within the frame.
8. A system for verifying that a user of an electronic device is a human and for generating an avatar, the system comprising: an electronic device having a processor, an interconnected memory, and a camera connected to the electronic device, wherein the process is configured to assigning a unique identifier to a user of an electronic device, wherein the user has a head and a body; verifying that the user’s head is within a frame of a camera connected to the electronic device; providing at least one randomly selected prompt to the user for a purpose of having the user comply with the prompt; recording a user’s compliance with the prompt by capturing an image of the user with the camera; evaluating the captured image by a trained model to generate a compliance score; comparing the compliance score to a predetermined threshold score; and creating a file of the captured image to train an Al model on the user; generating an avatar of the user from the trained Al model; and at least one non-transitoiy storage medium connected to the processor.
9. The system of Claim 8, also comprising repeating the steps of providing at least one prompt, recording a user’s compliance, comparing the captured image, comparing the compliance score, and creating a file for a predetermined number of time.
10. The system of Claim 8, wherein the at least one prompt is chosen from a library of poses.
11. The system of Claim 8, wherein the trained model is a vision-language model.
12. The system of Claim 8, wherein the trained model checks for at least one liveness indicators and for at least one anti-spoofing signal.
13. The system of Claim 8, wherein the Al model trained on the user is configured to generate a personalized avatar of the user.
14. A method of creating a personalized video from an audio file, comprising: generating a personalized avatar with an electronic device connected to a camera and having a processing unit and memory, the processing unit configured to execute the steps of: assigning a unique identifier to a user, having a head and a body, of the electronic device; verifying that the user’s head is within a frame of the camera; providing a plurality of random prompts to the user for a purpose of having the user comply with the prompts; recording a user’s compliance with each prompt by capturing an image of the user with the camera; comparing the captured images to a trained model and scoring the recorded compliance to generate a compliance score; comparing the compliance score to a predetermined threshold score; andcreating a file of each of the captured images to train an Al model on the user, wherein the Al model is configured to generate the personalized avatar of the user based upon the captured images; inputting an audio file into at least one model trained to generate visual content from the audio file that is contextually aligned with and synchronized to the audio file to create an animated reproduction of the audio file; and inserting the personalized avatar into the animated reproduction of the audio file.
15. The method of Claim 14, wherein the at least one model trained to generate visual content is selected from the group consisting of audio transcription, text-to-image diffusion, linear spherical interpolation, and music emotion recognition.
16. The method of Clain 14, also comprising music emotion recognition for predicting valence and arousal of the audio file.
17. The method of Claim 14, also comprising using a large language model to extract emotion from the audio file.
18. The method of Claim 14 wherein one of the at least one model trained to generate visual content is a latent diffusion model.
19. The method of Claim 14, wherein the audio file is selected from the group consisting of music, podcast, narration, dialog, and a voice recording.
20. The method of Claim 14, wherein the captured images are processing using low -rank adaptation.
21. A system for verifying whether a user is human, comprising: a user electronic device connected to a video camera and to a web browser;a server-side processing unit configured to verify a user’s identification and to manage verification workflow, wherein the server-side processing unit incorporates at least one computer vision module and at least one machine learning model for real-time facial detection and pose estimation; a verification engine using trained neural networks for pose recognition and comparison and configured with computational resources for real-time video processing; and a database configured to store unique identifiers and maintain user verification status, wherein the electronic device is configured to have video processing capabilities record, compress, and store image / video data, with storage infrastructure for potential Al model training and wherein the system integrates through API endpoints facilitating client -server communication, with security protocols protecting sensitive user data during transmission and storage.