Personalized digital visual representation system and method

By processing the facial geometry of user and character facial models to generate a fusion model, the problem of existing tools being unable to integrate user and character appearances is solved, and personalized digital visual representations are generated and rendered.

CN122497981APending Publication Date: 2026-07-31URUS ENTERTAINMENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
URUS ENTERTAINMENT CO LTD
Filing Date
2025-01-27
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing digital visual representation creation tools struggle to automatically integrate the user's visual appearance with the typical character appearance designed by the artist, while providing fine-grained control over the degree to which the visual representation resembles the user or typical character and specific areas.

Method used

The generation system utilizes first and second computing devices to process the facial geometry of user and character facial models, generates a fused facial model, and creates a personalized digital visual representation based on the fused model. It uses a hybrid weight graph to control geometric similarity and combines it with pre-configured rendering assets for rendering.

Benefits of technology

It achieves seamless integration of user visual appearance with artist-designed character appearance, provides fine-grained control over visual representation, and generates personalized digital visual representations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122497981A_ABST
    Figure CN122497981A_ABST
Patent Text Reader

Abstract

This paper discloses a computing system and method configured to generate one or more personalized digital visual representations. An example computing system may be configured to receive one or more photographs or videos of a user from a computing device, process the photographs or videos to at least determine the facial geometry of a user's facial model, obtain the facial geometry of at least several typical characters from a character facial model, generate a fused facial model to preserve the facial geometry of both the user's facial model and the character's facial model, generate at least one or more personalized digital visual representations based on the fused facial model, and transmit one or more personalized digital visual representations to the computing device.
Need to check novelty before this filing date? Find Prior Art

Description

Cross-references to related applications

[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 625,719, filed January 26, 2024, entitled “SYSTEM AND METHOD FOR CREATINGSTYLIZED DIGITAL AVATARS,” the entire contents of which are incorporated herein by reference. Technical Field

[0002] This disclosure generally relates to digital image processing, interactive media and software applications, and more specifically to face reconstruction and manipulation, geometric processing and rendering for creating personalized digital visual representations from user-generated photographs or video self-portraits. Background Technology

[0003] Digital visual representations of individuals or characters can be used in a variety of online and offline scenarios. Current digital visual representation creation tools often lack the ability to automatically integrate the user's visual appearance with that of an artist-designed typical character, while also allowing fine-grained control over the degree and specific areas of the visual representation's resemblance to the user or typical character. The term "typical character" can generally refer to any fictional character, including but not limited to native characters from specific books, television programs, movies, comic books, video games, and other story-driven universes, as well as characters synonymous with a specific brand, product, and group (such as brand mascots, promotional characters, and company icons).

[0004] Therefore, there is a need for an advanced computing system and method for creating personalized digital visual representations that automatically combines the visual appearance of a user with that of a typical character designed by an artist, while providing fine-grained control over the degree and specific areas to which the visual representation resembles the user or typical character. Summary of the Invention

[0005] Among other features, in one embodiment, this disclosure relates to a system deployed within a communications network for generating one or more personalized digital visual representations. The system may include a first computing device comprising: a first non-transitory computer-readable storage medium configured to store an application; and a first processor coupled to the first non-transitory computer-readable storage medium and configured to execute instructions of the application to obtain one or more photos or videos of a user. The system may further include a second computing device comprising: a second non-transitory computer-readable storage medium; and a second processor coupled to the second non-transitory computer-readable storage medium and configured to: receive one or more photos or videos of a user via a first application programming interface (API) call; process the one or more photos or videos to at least determine facial geometric features of a user's facial model; obtain facial geometric features of at least several typical characters from a character facial model; generate a fused facial model to preserve the facial geometric features of the user's facial model and the character's facial model; generate one or more personalized digital visual representations based at least on the fused facial model; store parameters related to the fused facial model on the second non-transitory computer-readable storage medium; and transmit one or more personalized digital visual representations to the first computing device via the second API call. The first processor of the first computing device may be further configured to execute instructions of an application program to receive and display one or more personalized digital visual representations on a display interface of the first computing device.

[0006] In an embodiment, the facial geometry of a user's facial model may include multiple three-dimensional (3D) meshes, each of which includes at least data defining the connectivity between vertices, edges, and the face in each of the multiple 3D meshes.

[0007] In another embodiment, the facial geometry of the user's facial model may include parameters related to multiple facial expressions.

[0008] In another embodiment, the second processor may be further configured to process one or more photos or videos to determine parameters related to facial texture.

[0009] In some embodiments, each of the fused facial model, the user facial model, and the character facial model may include multiple 3D meshes, and the second processor is further configured to allow adjustment of the degree of geometric similarity between a selected region of the fused facial model and a corresponding region of the user facial model or the character facial model.

[0010] In another embodiment, the second processor may be further configured to determine geometric similarity based at least on a local geometric descriptor including a facial deformation descriptor, and to use a hybrid weight map to control the geometric similarity between selected regions of the fused facial model and corresponding regions of the user's facial model or the character's facial model. The facial deformation descriptor can measure the local stretching and bending of each patch on the surface of each 3D mesh.

[0011] In another embodiment, the second processor may be further configured to determine multiple vertex positions of the fused facial model by minimizing the sum of the blended target facial deformation descriptors on the facial surface.

[0012] Furthermore, the second processor can be further configured to generate parameters representing at least one of stylized textures, shadows, and lines based on at least one of a plurality of pre-configured rendering assets, and to integrate at least one of the stylized textures, shadows, and lines with one or more background layers selected from the plurality of pre-configured rendering assets to render one or more personalized digital visual representations.

[0013] According to other aspects, this disclosure relates to a computing server system deployed within a communication network for generating one or more personalized digital visual representations. The computing server system may include a non-transitory computer-readable storage medium and a processor; the processor is coupled to the non-transitory computer-readable storage medium and configured to: receive one or more photos or videos of a user from a computing device deployed within a cloud-based communication network via a first application programming interface (API) call; process the one or more photos or videos to at least determine facial geometric features of a user's facial model; obtain facial geometric features of at least several typical characters from a character facial model; generate a fused facial model to preserve the facial geometric features of the user's facial model and the character's facial model; generate one or more personalized digital visual representations based at least on the fused facial model; store parameters associated with the fused facial model on the non-transitory computer-readable storage medium; and transmit one or more personalized digital visual representations to the computing device via a second API call.

[0014] In one embodiment, the computing device can be configured to render one or more personalized digital visual representations on the display interface of the computing device.

[0015] In some implementations, the facial geometry of a user's facial model may include multiple three-dimensional (3D) meshes, each of which includes at least data defining the vertices, edges, and connectivity between the face in each of the multiple 3D meshes.

[0016] Furthermore, the facial geometry of the user's facial model can include parameters related to multiple facial expressions.

[0017] In another embodiment, the processor may be further configured to process one or more photos or videos to determine parameters related to facial texture.

[0018] In some embodiments, each of the merged facial model, the user facial model, and the character facial model may include multiple 3D meshes, and the processor may be further configured to allow adjustment of the degree of geometric similarity between selected regions of the merged facial model and corresponding regions of the user facial model or the character facial model.

[0019] In other embodiments, the processor may be configured to determine geometric similarity based at least on a local geometric descriptor that includes a facial deformation descriptor.

[0020] The processor can be further configured to use a hybrid weight map to control the geometric similarity between selected regions of the fused facial model and corresponding regions of the user's facial model or the character's facial model.

[0021] In another embodiment, the facial deformation descriptor can measure the local stretching and bending of each patch on the surface of each 3D mesh.

[0022] The processor can be further configured to determine multiple vertex positions of the fused facial model by minimizing the sum of the blended target facial deformation descriptors on the facial surface.

[0023] According to another aspect, this disclosure relates to a computer implementation method for generating one or more personalized digital visual representations using different aspects of the system disclosed herein.

[0024] According to another aspect, this disclosure relates to a non-transitory computer-readable storage medium that stores one or more programs or instructions thereon, which, when executed by at least one computing device, generate one or more personalized digital visual representations to perform the computer-implemented methods disclosed herein.

[0025] The simplified overview of the examples above is intended to provide a basic understanding of this disclosure. This overview is not a comprehensive summary of all anticipated aspects and is neither intended to identify key or essential elements of all aspects nor to depict the scope of any or all aspects of this disclosure. Its sole purpose is to present one or more aspects in a simplified form as a prelude to the more detailed description that follows. To achieve the foregoing, one or more aspects of this disclosure include the described features and examples pointed to in the claims. Attached Figure Description

[0026] The accompanying drawings, which are incorporated in and form part of this specification, illustrate one or more exemplary aspects of this disclosure, together with detailed embodiments, to illustrate their principles and implementation.

[0027] Figure 1 An illustration shows a system deployed within a computing environment according to exemplary aspects of this disclosure and configured to generate personalized digital visual representations from user-generated photographs or video self-portraits.

[0028] Figure 2 Exemplary aspects according to this disclosure are shown Figure 1 A block diagram of the computing server system.

[0029] Figure 3 The overall workflow for generating personalized digital visual representations from user-generated photos or video self-portraits according to exemplary aspects of this disclosure is demonstrated.

[0030] Figure 4 A workflow for a facial reconstruction process according to an exemplary aspect of this disclosure is shown.

[0031] Figure 5 The workflow of a role fusion process according to an exemplary aspect of this disclosure is shown.

[0032] Figure 6 The workflow of a geometric fusion process according to an exemplary aspect of this disclosure is shown.

[0033] Figure 7 The workflow of a rendering process according to an exemplary aspect of this disclosure is shown.

[0034] Figure 8 A first example graphical user interface (GUI) according to an exemplary aspect of this disclosure is shown.

[0035] Figure 9 A second example GUI is shown according to an exemplary aspect of this disclosure.

[0036] Figure 10 It displays a front view image of the user's face.

[0037] Figure 11 It shows an image of the user's face taken from a left-hand perspective.

[0038] Figure 12 It shows an image of the user's face taken from the right-hand perspective.

[0039] Figure 13 An example user face model generated by a face reconstruction module according to an exemplary aspect of this disclosure is shown.

[0040] Figure 14 An example facial geometry specified by a character facial model is shown according to an exemplary aspect of this disclosure.

[0041] Figure 15 An example facial geometry (neutral expression) obtained from a user's face module is shown according to an exemplary aspect of this disclosure.

[0042] Figure 16 An example fused facial geometry (neutral expression) generated by the character fusion module according to an exemplary aspect of this disclosure is shown.

[0043] Figure 17 An example fused facial geometry (in a template expression) generated by the character fusion module according to an exemplary aspect of this disclosure is shown.

[0044] Figure 18 A first example of a personalized digital visual representation generated by a rendering module according to an exemplary aspect of this disclosure is shown.

[0045] Figure 19 A second example of a personalized digital visual representation generated by a rendering module according to an exemplary aspect of this disclosure is shown. Detailed Implementation

[0046] Various aspects of this disclosure will be described with reference to the accompanying drawings, wherein the same reference numerals are used throughout to refer to the same elements. In the following description, numerous specific details are set forth for purposes of explanation to facilitate a thorough understanding of one or more aspects of this disclosure. However, in some or all instances it will be apparent that any of the aspects described below may be implemented without employing the specific design details described below.

[0047] See Figure 1According to various aspects of this disclosure, a computing system 100 deployed within a computing environment and communication network can be configured to acquire certain image signals (e.g., user-generated photographs or video self-portraits) from at least one user 102a, 102b...102n. The computing system 100 can also be configured to automatically generate personalized digital visual representations that seamlessly integrate the visual appearance of each user 102a, 102b...102n with a typical character designed by an artist. Furthermore, the computing system 100 implements fine-grained control that can be directed individually or in combination by a user, artist, or other entity to adjust the degree and specific areas of the digital visual representation's resemblance to the user or typical character. As will be fully described below, each artist-designed typical character of this disclosure may include a data object representing a structured digital representation of the character, the data object including attributes such as appearance (3D / 2D models, textures), behavior (animations, scripts), speech (audio files or synthesized speech data), and metadata.

[0048] In one embodiment, the application may be a mobile application or a web-based application (e.g., a native iOS application or an Android application) that can be downloaded and installed on selected computing devices or systems 104, 106, or 108 to interact with each user 102a, 102b…102n and exchange data and information, as well as other characteristics, with other computing devices deployed within computing system 100. For example, users 102a, 102b…102n may include end users, subscribers, content creators, players, members, customers, system administrators, network administrators, software developers, and artists. Automated agents, scripts, and playback software representing one or more individual actions may also be users 102a, 102b…102n. Such user-oriented applications of computing system 100 may include multiple modules (e.g., a camera or any suitable optical sensor) executed and controlled by the processor of the hosting computing device or system 104, 106, or 108 to obtain input such as user-generated photographs or video self-portraits. Each computing device 104, 106, or 108 hosting a mobile or web-based application can be configured to connect to a remote backend computing server system 114 using a suitable communication protocol 110 and a communication network 112. In this document, the communication network 112 typically includes a collection of geographically distributed computing devices or data points interconnected by communication links and segments for transmitting signals and data between them. The communication protocol 110 generally includes a set of rules defining how computing devices and networks can interact with each other, such as Frame Relay, Internet Protocol (IP), Transmission Control Protocol (TCP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), etc. It should be understood that the computing system 100 of this disclosure can use any suitable communication network, ranging from local area networks (LANs), wide area networks (WANs), cellular networks to overlay networks and software-defined networks (SDNs), packet data networks (e.g., the Internet), mobile phone networks (e.g., cellular networks such as 4G or 5G), Plain Old Telephone Service (POTS) networks, and wireless data networks (e.g., known as Wi-Fi). ® WiGig® The Institute of Electrical and Electronics Engineers (IEEE) 802.11 series of standards, known as WiMax. ® The IEEE 802.16 series of standards, the IEEE 802.15.4 series of standards, the Long Term Evolution (LTE) series of standards, the Universal Mobile Telecommunications System (UMTS) series of standards, peer-to-peer (P2P) networks, virtual private networks (VPNs), Bluetooth, Near Field Communication (NFC), or any other suitable network.

[0049] In some embodiments, the computing server system 114 may be a cloud-based or on-premises server. The term "server" generally refers to a computing device or system, or a collection of computing devices or systems, including processing hardware and process space, associated storage media (such as memory devices or databases), and in some cases at least one database application well known in the art. The computing server system 114 may provide functionality for any connected device, such as sharing data or provisioning resources among multiple client devices, or performing computations for each connected client device. According to a preferred embodiment, within a cloud-based computing architecture, the computing server system 114 may use shared resources to provide different cloud computing services. Cloud computing generally includes internet-based computing, where computing resources are dynamically provisioned and allocated on demand to each connected computing device or other device via a network or cloud from an available set of resources. Cloud computing resources may include any type of resource, such as computing, storage, and networking. For example, resources may include service devices (firewalls, deep packet inspectors, traffic monitors, load balancers, etc.), computing / processing devices (servers, central processing units (CPUs), graphics processing units (GPUs), random access memory, caches, etc.), and storage devices (e.g., network-attached storage, storage area network devices, hard drives, solid-state devices, etc.). Furthermore, such resources can be used to support virtual networks, virtual machines, databases, applications, etc. As used herein, the term "database" may refer to a database (e.g., a Relational Database Management System (RDBMS) or a Structured Query Language (SQL) database), or any other data structure such as, for example, comma-separated values ​​(CSV), tab-separated values ​​(TSV), JavaScript Object Notation (JSON), extensible markup language (XML), text (TXT) files, flat files, spreadsheet files, and / or any other widely used or proprietary format. In some embodiments, one or more databases or data sources may be implemented using one of the following: relational database, flat file database, entity relational database, object-oriented database, hierarchical database, web database, NoSQL database, and / or record-based database.

[0050] Cloud computing resources accessible via any suitable communication network (e.g., the Internet) can include private clouds, public clouds, and / or hybrid clouds. In this document, a private cloud can be cloud infrastructure operated by an enterprise for its own use, while a public cloud can refer to cloud infrastructure that provides services and resources for public use over a network. In a hybrid cloud computing environment, which uses a mix of enterprise-on-premises, private clouds, and third-party public cloud services that coordinate between the two platforms, data and applications can move between private and public clouds for greater flexibility and more deployment options. Some example public cloud service providers may include Amazon (e.g., Amazon Web Services). ® AWS, IBM (e.g., IBM Cloud), Google (e.g., Google Cloud Platform), and Microsoft (e.g., Microsoft Azure) ® These providers use computing and storage infrastructure at their respective data centers to deliver cloud services, which are typically accessed via the Internet. Some cloud service providers (e.g., Amazon AWS Direct Connect and Microsoft Azure ExpressRoute) offer direct connection services, and such connections typically require users to purchase or rent a private connection to a peer provided by these cloud providers. In some implementations, the computing server system 114 of this disclosure (e.g., a cloud-based or on-premises server) may additionally be configured to connect to various data sources or facilities, such as multiple computing systems 116a, 116b, 116c, ... 116n.

[0051] According to some embodiments, the user-facing applications of computing system 100 may include multiple modules and libraries, which are executed and controlled by a microcontroller or processor of a managed computing device or system 104, 106, 108, for performing functions locally on each computing device and / or making remote calls (e.g., application programming interface (API) calls) to computing server system 114 to access specific functions. The division of labor between local execution and server-side operation depends on how each module or library is designed and its functional requirements.

[0052] Depending on the implementation, one or more libraries downloaded on each selected computing device or system 104, 106, 108 can be configured to perform all their operations locally without relying on computing server system 114. That is, once a library is installed, it can access the resources and computing power available on each computing device 104, 106, 108 to perform tasks. For example, some libraries can be configured to perform computations locally using the CPU / GPU of each computing device. Furthermore, file processing libraries can be configured to process files stored on local devices. In one aspect, as will be fully described below, the rendering of personalized digital visual representations can be performed locally on each computing device 104, 106, 108, or by computing server system 114, or partially by both computing devices and computing server system.

[0053] According to other embodiments, remote execution (server-side processing) can be implemented, and libraries downloaded on each computing device 104, 106, 108 can remotely invoke (e.g., API calls) the computing server system 114 to access certain functions, such as when the functions provided by the library are too resource-intensive for local execution or require access to constantly updated data (e.g., real-time services, large-scale models, or databases). In this case, the library acts as a client-side interface for making API calls or requests to the computing server system 114 to perform specific tasks.

[0054] In one example, the library can connect to service interfaces such as OpenAI's GPT, Google Cloud AI, Claude Sonnet, or Amazon S3, where computations can be performed on computing server system 114, and selected computing devices 104, 106, and 108 send requests and receive results. In another example, libraries such as the AWS SDK and Google Cloud SDK allow interaction with cloud storage to upload, retrieve, and manipulate data in the cloud.

[0055] Server-side processing can offload a large amount of computation to powerful servers, provide access to real-time data and update services, and achieve device independence by operating even on resource-constrained devices (smartphones, tablets, etc.). In one embodiment, computing server system 114, or at least one of a plurality of computing systems 116a, 116b, 116c, ... 116n that can be accessed by computing server system 114 via, for example, API calls, may be configured to provide server-side processing.

[0056] According to another embodiment, the libraries implemented on each selected computing device 104, 106, 108 may adopt a hybrid model, where some operations or computations can be performed locally, while more complex or resource-intensive tasks are offloaded to the computing server system 114. For example, basic computations, preprocessing, or user interface elements may be processed locally, while complex processing, data retrieval, or heavy computations (e.g., interacting with a database or running large AI (Artificial Intelligence) / ML (Machine Learning) models) may be performed via the computing server system 114.

[0057] See now Figure 2 According to various aspects of this disclosure, at least one processor 204 of the computing server system 114 may be configured to control and execute multiple modules, including a transceiver module 206, an interface 208, a face reconstruction module 210, a character fusion module 212, and a rendering module 214. In one embodiment, the face reconstruction module 210 may include an image processing module 210a, a geometry fitting module 210b, and a photometric fitting module 210c. The character fusion module 212 may include a mesh retopology module 212a, a geometry fusion module 212b, a face model fitting module 212c, a face animation module 212d, and a deformation module 212e. Further, the rendering module 214 may include a stylized texture generation module 214a, a shading module 214b, a line generator 214c, and a compositing module 214d. As used herein, the terms "module" or "generator" refer to a real-world device, component, or arrangement of components and circuits implemented using hardware (such as by an Application Specific Integrated Circuit (ASIC) or Field-Programmable Gate Array (FPGA)) or as a combination of hardware and software (e.g., by a microprocessor system and a set of instructions for implementing the module's functionality), which, when executed, transforms the microprocessor system into a dedicated device. A module or generator can also be implemented as a combination of both, where some functionality is facilitated solely by hardware and other functionality by a combination of hardware and software. Each module or generator can be implemented in a variety of suitable configurations and should not be limited to any of the exemplary implementations illustrated herein.

[0058] Memory 216 coupled to processor 204 may be configured to store at least a portion of information acquired and generated by computing server system 114. In one aspect, memory 216 may be a non-transitory machine-readable medium configured to store at least one set of data structures or instructions (e.g., software) embodying or utilized by at least one of the techniques or functions described herein. It should be understood that the term "non-transitory machine-readable medium" may include a single or multiple media (e.g., one or more caches) configured to store at least one instruction. The term "machine-readable medium" may include any medium capable of storing, encoding, or carrying instructions that are executed by all modules of the computing device and cause those modules to perform at least one of the techniques disclosed herein, or may include data structures capable of storing, encoding, or carrying data used by or associated with such instructions. Examples of non-limiting machine-readable media may include solid-state memory as well as optical and magnetic media. Specific examples of machine-readable media may include: non-volatile memory, such as semiconductor memory devices (e.g., electrically programmable read-only memory (EPROM) and electrically erasable programmable read-only memory (EEPROM)) and flash memory devices; disks, such as internal hard disks and removable disks; magneto-optical disks; random access memory (RAM); solid-state drives (SSDs); and CD-ROM (Compact Disc Read-Only Memory) and DVD-ROM (Digital Versatile Disc Read-Only Memory) disks.

[0059] The transceiver module 206 of the computing server system 114 can be configured by the processor 204 to exchange various information and data with other computing devices on which the computing system 100 is deployed. For example, the transceiver module 206 can receive data from devices such as electronic cameras or image capture devices commonly used by users, or smartphones with built-in cameras (e.g., [missing information]). Figure 1 Input of a self-portrait photograph or short video taken by at least one of computing devices or systems 104, 106 or 108.

[0060] Interface 208 can be configured by processor 204 to provide various communication and interaction functions between various software components, hardware components, or users. For example, interface 208 can provide a set of functions or protocols for other components to interact with a specific system or service, or it can be a physical device or circuit connecting different electronic components or systems. In one embodiment, the mobile application or web-based application of this disclosure can be a thin client device / terminal / application deployed within computing system 100 and can be configured to perform some preliminary processing of data. The preprocessed data can then be transferred to computing server system 114 for further processing. In one implementation, interface 208 may include an API interface configured to make one or more API calls. According to another implementation, computing server system 114 may include an API gateway device (not shown) configured to receive and process API calls from various connected computing devices deployed within computing system 100 (e.g., operating system, libraries, device drivers, APIs, applications, software, or other modules). Such an API interface or gateway device may specify one or more functions, methods, classes, objects, protocols, data structures, formats, and / or other characteristics of the computing server system 114 that can be used by mobile applications or web-based applications. For example, the API interface may define at least one calling convention that specifies how a function associated with the computing server system 114 receives data and parameters from the requesting device / system and how the function returns results to the requesting device / system. It should be understood that the computing server system 114 may include additional functions, methods, classes, data structures, and / or other characteristics not specified through the API interface and not available to the requesting computing device.

[0061] Figure 3 A general workflow 300 for rendering a personalized digital visual representation from a user-generated photograph or video self-portrait, according to various aspects of this disclosure, is demonstrated. For example, a computing server system 114 can be configured to take a collection of user photographs or videos 302 as input and then sequentially invoke a face reconstruction module 210, a character fusion module 212, and a rendering module 214 to generate a rendering of the personalized digital visual representation. The user photographs or videos 302 can be transmitted via, for example, a computing device (e.g., Figure 1 The images are obtained using a camera or image capture device built into or externally connected to the computing device or system 104, 106, or 108. In another embodiment, one or more photos and videos stored on the computing device or system 104, 106, or 108 may be uploaded to the computing server system 114. User photos or videos 302 may capture various features of the user's face from one or more angles.

[0062] In implementing the face reconstruction process 304, the face reconstruction module 210 may be configured by the processor 204 to take a set of user photos or videos 302 as input, process the input 302 using the image processing module 210a, and estimate and output the user face model 306 using the geometry fitting module 210b and the photometric fitting module 210c. In some aspects, the user face model 306 may include at least the user's facial geometry and may also include texture maps encoding the user's facial appearance attributes, such as albedo, reflectivity, and roughness parameters. Additionally, the user face model 306 may include a representation of face segmentation that assigns values ​​to points on the facial geometry based on predefined semantic categories. These categories may include, but are not limited to, left eye, right eye, nose, lips, facial hair, left ear, right ear, and other facial features. The face segmentation representation may be implemented as a semantically labeled texture map, where each texture element (texel) is assigned a discrete semantic label, or as a multi-channel texture map, where each channel stores the probability that a given texture element belongs to a specific semantic category. A texel (short for texture element) is the smallest unit of a texture map, representing a 2D image used to apply visual details (such as color, pattern, or material) to a 3D object. In some implementations, a texture element can be equivalent to a pixel in a texture image, but it is used specifically in the context of textures applied to 3D surfaces. Furthermore, the user facial model 306 may include expression-related information, which enables the calculation of the user's facial geometry and surface appearance attributes corresponding to different facial expressions.

[0063] In implementing the character fusion process 310, the character fusion module 212 can be configured to take the user's facial model 306 determined by the facial reconstruction module 210 and the artist-designed character model 308 as inputs to generate a fused facial model 312. In some aspects, the character model 308 may include at least one or more facial geometries of a typical character, each facial geometry corresponding to a different pre-configured facial expression. The fused facial model 308 may include at least facial geometry that can combine local and global geometric features from both the user's facial geometry and the artist-designed character's facial geometry. As will be fully described below, the character fusion process 310 can be configured to allow the end user or artist to select and adjust the degree of geometric similarity between each local patch on the output facial mesh and the two input meshes representing the facial geometry of the user and the selected artist-designed character. In one embodiment, geometric similarity can be measured by local curvature and / or other local geometric descriptors. The fused facial model 312 may also include facial appearance attributes, such as texture maps encoding albedo, reflectivity, and roughness parameters. Furthermore, the fused facial model 312 may include a representation of facial segmentation similar to that included in the user facial model 306. Additionally, the fused facial model 312 may include expression-related information that enables the calculation of facial geometry and surface appearance attributes corresponding to different facial expressions, wherein the calculated facial geometry may combine local and global geometric properties from both the user's facial geometry corresponding to a given facial expression and the facial geometry of an artist-designed character.

[0064] In implementing rendering process 316, rendering module 214 may be configured to generate a personalized digital visual representation 318 for the user, at least in part, based on the blended facial model 312 calculated by character blending module 212 and pre-configured rendering assets 314 generated by the artist. The personalized digital visual representation may include one or more images or videos, including but not limited to the rendering of the blended facial model 312. In some embodiments, the personalized digital visual representation 318 may be rendered on a local computing device (such as a smartphone, tablet, personal computer, or game console, etc.) (e.g., via an application downloaded and installed on selected computing devices or systems 104, 106, or 108 for interaction with each user 102a, 102b…102n), thereby enabling a real-time interactive user experience. In such an embodiment, the rendering module 214 may be further configured to receive and process input from a local computing device (e.g., one of computing devices or systems 104, 106, or 108), including but not limited to keyboard input, mouse input, game controller input, touchscreen input, camera feed, accelerometer, and gyroscope, to render a personalized digital visual representation 318.

[0065] In one embodiment, the pre-configured rendering asset 314 may include various data and parameters associated with icon elements and components of a typical character designed by an artist. This data and parameters may be stored, for example, in memory 216 or on any suitable data storage computing system (e.g., one of multiple computing systems 116a, 116b, 116c, ... 116n) deployed within computing system 100 and accessible by computing server system 114. For example, multiple 3D models of a typical character designed by various artists (e.g., 3D meshes, topology, and levels of detail) may define the geometry of each 3D model, including: vertices, edges, and polygons (e.g., triangles or quadrilaterals), the arrangement of polygons (which may determine deformation, animation, and optimization), and versions of each mesh at different resolutions to ensure performance optimization for different rendering distances. Texture data and parameters, such as texture maps and resolution information, may also be included. For example, texture maps may include, but are not limited to: diffuse / base color maps for defining the color and visible details of a surface; normal maps configured to simulate surface details (such as bumps) without increasing mesh complexity; specular or glossy or metallic or roughness maps that can define reflectivity or how light interacts with the surface; ambient occlusion maps that can be configured to add shadows to cracks and fine details for a more realistic appearance; and emission maps for controlling the self-emission of specific areas. Resolution information can be measured, represented, and stored in pixels (e.g., 1024×1024, 2048×2048).

[0066] In another embodiment, the pre-configured render asset 314 may include different shaders. A shader can refer to a small program or script designed to run on a CPU / GPU to control how surfaces, textures, and materials appear on 3D objects within a rendered scene. Shaders can be used to create visual effects, add realism, or stylize the appearance of objects in a 3D environment. For example, the pre-configured render asset 314 may include data and parameters related to various material properties to define how each asset interacts with light (e.g., metallic, matte, translucent, subsurface scattering of skin, etc.). Different shader models (such as pre-configured emissivity and rendering rules) can be used to mimic physically or stylized effects. Custom shaders can be incorporated to achieve specific artistic or unique effects tailored to a selected design.

[0067] Furthermore, in some embodiments, the pre-configured render asset 314 may include rigs and animation data. For example, rigs may include a layered skeletal structure within each asset for animation. Joint number and placement data can be used to generate natural movement. Inverse kinematics data can simplify animations such as walking or hand movement. Weighted drawing data can assign the influence of each bone to mesh vertices for smooth deformation. Blend shape or deformable target data can be used for facial expressions, lip synchronization, or deformation. Additionally, various animation clips may be included to define pre-built motions such as walking, running, or gestures.

[0068] The pre-configured rendered assets 314 can also include lighting data. For example, pre-configured lighting settings can define how each asset will look under different lighting conditions. High dynamic range imaging can be used to provide an environment for dynamic lighting settings.

[0069] Audio metadata can also be included in pre-configured render assets 314. For example, lip-sync data can define parameters that align facial animation with audio. Trigger points can synchronize specific actions or animations with sound.

[0070] Interoperability can be provided by pre-configured rendering assets 314. For example, common file formats processed by computing system 100 can involve models (e.g., FBX, OBJ, GLTF / GLB), textures (e.g., PNG, JPEG, TIFF), and animations (e.g., BVH, FBX). Compatibility can also be provided. For example, certain elements of the pre-configured rendering assets 314 can be designed to work with specific engines (e.g., Unity, Unreal Engine) or platforms.

[0071] Depending on the implementation, performance optimization can be provided by pre-configured rendering assets. For example, the number of polygons can balance visual quality with performance requirements. Atlas textures can be used to combine multiple textures into a single piece to reduce painting calls. Bake maps can be used to reduce rendering computations by pre-compiling certain effects.

[0072] Custom parameters can be included in pre-configured rendered assets. For example, different deformation options allow end users to adjust features such as body shape, facial features, or clothing details. Material libraries enable quick switching of looks and styles. Various color palettes provide predefined or customizable colors for different parts of the asset.

[0073] Pre-configured render assets 314 can additionally include documentation and metadata. For example, various specifications can provide guidance on using and integrating each asset. Metadata tags can be used to provide information such as asset name, type, version, and dependencies.

[0074] Figure 4 A workflow 400 is illustrated for a facial reconstruction process for generating a user facial model 416, according to several aspects of this disclosure. Depending on some implementations, the facial reconstruction module 210 may be configured to take a collection of user photographs or videos 402 as input. The user photographs or videos 402 may contain different facial expressions of the user and may be captured under different lighting conditions.

[0075] According to one embodiment, see Figure 2 and Figure 4 Generating a user's facial model 416 via the facial reconstruction module 210 typically includes three steps: image preprocessing 404 via the image processing module 210a, geometric fitting 410 via the geometric fitting module 210b, and photometric fitting 414 via the photometric fitting module 210c.

[0076] Image processing 404 can take a user photograph or video 402 as input and extract a facial image and facial segmentation map 406, as well as facial landmark points 408, associated with each facial image. If a photograph is included in input 402, it is loaded as an image, which in this context means a rectangular grid of colored pixels. When video is included in input 402, it can first be converted into a sequence of images corresponding to a subset of all video frames. In one embodiment, the input video can be compressed by a video codec associated with a computing device, which uses intra-frame encoded frames or I-frames. Each I-frame can be fully encoded and can be decoded independently of any other frame. In other words, each I-frame can contain all the information needed to reconstruct the image of that particular frame. I-frames can serve as reference points for decoding subsequent frames, such as P-frames (predictive frames) and B-frames (bidirectional predictive frames). Because they are compressed independently and do not depend on adjacent frames, I-frames can have a larger data size compared to other frame types. In one embodiment, to achieve compression efficiency, I-frames can typically be used slightly at the beginning of the video sequence or at regular intervals (e.g., every 2-5 seconds). A subset of all I-frames can then be extracted from the input video to create an image sequence. In one embodiment, if the total number of extracted I-frames exceeds a predefined threshold, a subset equal to the threshold size can be selected by picking I-frames at fixed intervals. In another embodiment, all I-frames can be extracted from user video 402. Therefore, image processing module 210a can be configured to ensure that both photos and videos in input 402 are converted into image sets and stored, for example, in memory 216 or any suitable data storage computing system (e.g., one of multiple computing systems 116a, 116b, 116c, ... 116n) deployed within computing system 100 and accessible by computing server system 114.

[0077] Subsequently, a pre-configured facial landmark detector (such as off-the-shelf facial landmark detectors available in Dlib, OpenCV, and MediaPipe) can be applied to each image in the image set transformed from input 402, generating a list of facial landmark points 408 for each image. These facial landmark points 408 are the image locations of predefined facial landmarks, which may include, but are not limited to, the corners of the eyes, the root of the nose, the boundary points of the lips, and the jawline contour. These facial landmark points 408 can be used to identify approximate facial regions in the images. In one embodiment, each image can be filled and cropped using bounding boxes calculated based on these facial landmark points 408 to approximately enclose the user's face in each image. The cropped images can be resized to a configurable predefined size (e.g., 512 × 512 pixels in one embodiment) to produce a final facial image 406. The coordinates of the facial landmark points 408 can be affine transformed to match the new coordinate system of the cropped and resized facial image 406.

[0078] Furthermore, a pre-configured face parser can be executed to generate a face segmentation map 406 for each face image. A face parser is a computational model or algorithm designed to analyze a face image and segment it into different semantic regions such as eyes, nose, mouth, ears, and facial skin. The segmentation map can assign discrete labels to each pixel in the image, or alternatively, it can be represented as a multi-channel map, where each channel can store the probability that a given pixel belongs to a specific semantic category. Example face parsers can use MediaPipe Face Mesh and FaceParsing-PyTorch.

[0079] In summary, image processing step 404 can generate a set of predefined dimensional facial images 406, and a corresponding segmentation map 406 and facial landmark points 408 for each facial image. These outputs 406, 408 can be passed to geometric fitting step 410 and photometric fitting step 414. Figure 4 As shown, for further processing.

[0080] In some embodiments, the geometry fitting step 410 may estimate facial geometry and camera parameters 412 for each facial image. This step roughly aligns the facial geometry with detected facial landmarks 408 and provides a basis for subsequent photometric fitting 414.

[0081] In some embodiments, the estimated camera parameters for the i-th face image may include camera rotation. R i Camera panning t iand field of view (FOV) angle θ i The estimated facial geometry can be defined using a predefined 3D Morphable Face Model (MFM) (also referred to as a 3D Face Morphable Model (3DMM) in the literature). MFM is a statistical model that generates a 3D facial mesh from a fixed set of coefficients, enabling parametric control over facial shape and appearance. In one embodiment, the FLAME (Faces Learned with an Articulated Model and Expressions) model can be used as a predefined MFM. MFM coefficients can be divided into identity coefficients and expression coefficients. Identity coefficients encode object-specific facial structures, while expression coefficients encode expression-specific facial structures and other object-independent facial structures. Since all facial images can correspond to the same individual, a single vector of identity coefficients can be estimated for all facial images. C id Each facial image can have its own expression coefficient vector. (in, i (Indexed facial images). In this way, the estimated facial geometry of all facial images is encoded as an MFM coefficient vector. C id and .

[0082] The geometry fitting step 410, performed by the geometry fitting module 210b, can also utilize a predefined landmark-mesh correspondence table (not shown). This table can be used to determine 3D points on the facial geometry corresponding to each facial landmark point 408 detected in image processing step 404. For clarity, a 3D point on the facial geometry corresponding to a given facial landmark point 408 can be referred to as a 3D landmark point corresponding to facial landmark point 408. Since the 3D facial meshes generated by MFM share the same mesh connectivity information, the landmark-mesh correspondence table can be represented as a list of facial indices specifying which triangle in the 3D mesh generated by MFM each 3D landmark point is located in and its centroid coordinates relative to the indexed face. Centroid coordinates are a coordinate system used in geometry to represent the position of a point relative to the vertices of a triangle (or a simplex in a higher dimension).

[0083] The geometric fitting step 410 can be achieved by minimizing at least one geometric fitting error. E geo To estimate by C id and Encoded facial geometry and camera parameters 412; geometric fitting error measurement: the difference between the facial landmark point 408 detected in image processing step 404 and its corresponding 3D landmark point projected onto the 2D image plane by the corresponding camera; in: i An index on a facial image; j For indexes on facial landmarks predefined by the facial landmark detector; It is in the first i The first detected on the facial image j Facial landmarks; It is the first j The 3D position of each 3D landmark is calculated using estimated identity coefficients. C id and expression coefficient Estimate the MFM, then use the landmark-grid correspondence table for lookup; finally for Based on estimated camera parameters , and The projected 2D position.

[0084] Minimization can be performed using gradient-based solvers such as Adam or LBFGS, where the gradient is estimated using automatic differentiation. Minimization may include additional regularization terms to ensure the coefficient vector and camera parameters remain reasonable. These terms may include L2 regularization of the identity and expression coefficients, as well as prior knowledge-based regularization of the camera parameters.

[0085] See also Figure 4 The photometric fitting step 414 can optimize the initial facial geometry and camera parameters 412 generated by the geometric fitting step 410, and estimate the lighting parameters and facial albedo color by comparing the input facial image 406 with the currently estimated rendering based on the facial geometry, camera parameters, lighting parameters and facial albedo color.

[0086] In some embodiments, the facial geometry estimated during photometric fitting step 414 can be determined by a single identity coefficient vector. C id Facial expression coefficient vector for each face image And an additional per-vertex scalar displacement field representation, wherein the coefficient vector is initialized from the output of the geometry fitting step 410. For example, the per-vertex scalar displacement field can be applied by first estimating the MFM using the estimated identity coefficients and expression coefficients to obtain a 3D mesh, then calculating their per-vertex normals, and finally shifting the mesh vertex positions along the per-vertex normals with associated scalar displacements. The same per-vertex scalar displacement field can be shared across all face images. As in the geometry fitting step 410, the estimated camera parameters for each face image can be represented using camera rotation, camera translation, and FOV angle. The estimated illumination parameters 412 can be represented using spherical harmonic coefficients in a pre-configured order (e.g., order 3 in one embodiment). The estimated facial albedo color can be represented as a 2D color texture map defined relative to the 3D mesh generated by the MFM. The same facial albedo texture can be shared across all face images.

[0087] The photometric fitting step 414, performed by the photometric fitting module 210c, can utilize the photometric fitting error. To measure the difference between the input facial image 406 and the rendering made from various estimated parameters. , in, i An index for facial images; The index of the pixel; It is the first i Facial images; and It is used by the differentiable renderer. i An image is generated by estimating the facial geometry, camera parameters, and lighting parameters from a facial image, as well as the estimated facial albedo color. A differentiable renderer can use a diffuse shading model to calculate pixel colors.

[0088] Photometric fitting step 414 can minimize the total energy function, which includes not only the photometric fitting error. It also includes the geometric fitting error used by the geometric fitting step 410. This involves various regularization terms. This minimization can be performed to estimate the identity coefficient vector and expression coefficient vector, the scalar per-vertex displacement field, the facial albedo color texture, camera parameters, and lighting parameters. Regularization terms may include norm constraints on the optimization variables, image-space smoothness constraints on the facial albedo color texture, and geometric smoothness constraints on the estimated facial geometry. Optimization can be performed using gradient-based optimizers such as Adam or LBFGS, which have gradients computed via automatic differentiation.

[0089] After optimization, photometric fitting step 414 can output an identity coefficient vector and a per-vertex displacement field as part of the user's facial model 416. The identity coefficient vector and the per-vertex displacement field together define the user's facial geometry. The user's facial geometry corresponding to different facial expressions can be estimated by using at least the identity coefficient vector in combination with the expression coefficient vector encoded for the selected facial expressions, and then computed by applying the per-vertex scalar displacement field.

[0090] The photometric fitting step 414 can also output an estimated facial albedo color texture as part of the user's facial model 416, which specifies the user's facial surface appearance.

[0091] Furthermore, the photometric fitting step 414 can utilize the facial geometry and camera parameters estimated during the photometric fitting step to unproject the face segmentation map 406 from image space to texture space, thereby generating a face segmentation texture. Individual face segmentation textures can be computed for each camera view by projecting the corresponding 3D facial geometry onto the camera's image plane. These face segmentation textures derived from multiple camera views can then be combined using a weighting scheme. The weights can be determined based on the visibility of each texture element point in the respective camera view, such as by using a visibility test (e.g., z-buffering), or by averaging contributions from all views, weighted by the projected regions of texture elements in different views. In this way, face segmentation maps 406 from different face images 406 can be unified into a single face segmentation texture. The resulting face segmentation texture can be included as part of the user face model 416 to represent face segmentation.

[0092] It should be recognized that the facial reconstruction module 210 can be implemented using any chosen technology. For example, the computing server system 114 can host, train, and operate at least one deep learning and neural network (e.g., Figure 1 The computational system comprises at least one of 116a, 116b, 116c, ..., 116n. In one embodiment, one or more convolutional neural networks (CNNs) can be used to recognize and process facial features from an image. CNNs can segment and recognize detailed facial landmarks with high accuracy. In yet another embodiment, a CNN can be trained to predict facial geometry, texture, lighting parameters, and camera parameters directly from 2D input. Furthermore, a neural radiance field (NeRF) can be used to reconstruct a 3D scene containing a face from a photograph, and 3D facial meshes and textures can be extracted from the NeRF-based reconstruction.

[0093] If multiple angles are available (e.g., from user photos or videos 402), the computational server system 114 can use structure-from-motion (SfM) to reconstruct 3D geometry by matching points across frames and estimating camera positions. Furthermore, for images with known stereo pairs or multiple views, stereo photogrammetry can be used to provide accurate depth estimates for each point in the face.

[0094] According to other embodiments, the computational server system 114 can perform a highly accurate face reconstruction process using depth estimation and shape-from-X techniques. For example, monocular depth estimation can be performed using a deep network trained on a large dataset, thereby predicting the depth map directly from a single image. In one embodiment, shape-from-shading techniques can be combined to estimate surface normals and reconstruct fine facial details by leveraging variations in illumination. The computational server system 114 can use photometric stereo techniques to analyze multiple images under varying illumination directions to estimate surface details by resolving complex lighting interactions.

[0095] In some implementations, the computational server system 114 can use physics-based modeling. For example, soft tissue simulation can be configured to model the biomechanics of facial skin, muscles, and underlying structures to achieve realism, particularly in dynamic expressions. Facial biomechanical techniques can be used to integrate knowledge of muscle anatomy and skeletal structure to constrain the reconstruction within reasonable limits.

[0096] In some embodiments, cross-modal fusion techniques can be implemented. For example, computing server system 114 can use voice-driven facial dynamics to combine audio and visual inputs to improve the accuracy of facial movements in voice-related expressions.

[0097] Alternatively, the computing server system 114 can use semantic prioritization and AI-assisted techniques. For example, prior knowledge enhancement can be used to incorporate prior knowledge such as average facial shape, statistical facial models, and symmetry constraints to fill in incomplete or occluded areas.

[0098] In a further embodiment, the computation server system 114 may use augmented datasets to improve the model's generalization ability by leveraging datasets with varying poses, lighting, and expressions. Alternatively, the computation server system 114 may train the reconstruction model on diverse datasets (synthetic, real-world, and augmented) to achieve broader applicability.

[0099] Figure 5A workflow 500 of a character fusion process according to several aspects of this disclosure is shown, which generates a fused facial model 312 from a character model 308 and a user facial model 306 generated by a facial reconstruction process 304.

[0100] According to one embodiment, the character model 308 may include a character facial geometry 502, a character template expression 504, and a facial attachment model 503. The character facial geometry 502 may be stored on one of a plurality of computing systems 116a, 116b, 116c, ... 116n deployed within the computing system 100 and accessible by the computing server system 114. The character facial geometry 502 may be configured to include information and data relating to a triangular mesh of a selected typical character. The facial attachment model 503 may include information and data relating to facial attachments such as facial hair, earrings, teeth, etc. The user facial geometry 501 refers to the facial geometry specified in the user facial model 306 generated by the facial reconstruction process 304. The user facial texture 510 refers to various facial textures that may be part of the user facial model 306 generated by the facial reconstruction process, such as facial albedo color textures and facial segmentation textures. In one embodiment of face reconstruction step 304, if the user's facial geometry 501 is represented by the MFM identity coefficient vector as described above, the role fusion module 212 can be configured to use the identity coefficient vector in combination with the MFM to reconstruct a triangular mesh representation of the user's face for a specific facial expression encoded by a given expression coefficient vector.

[0101] If the meshes specified by the user facial geometry 501 and the character facial geometry 502 have different topologies, a mesh retopology step 508, performed by the mesh retopology module 212a, can be performed on one or both of these meshes to convert them to the same vertex ordering and connectivity. According to one embodiment, the retopology step 508 can be performed by resampling the mesh using barycentric coordinates loaded from a retopology graph 509, which is generated offline in the mesh wrapping step 506. According to one implementation, the mesh wrapping step 506 can use a software package (such as R3DS Wrap) to wrap the geometry of one facial model onto another facial model and store the retopology graph in, for example, memory 216. The retopology step 508 can output a neutral expression user facial geometry 513 and store it in, for example, memory 216 or in any suitable data storage computing system (e.g., one of a plurality of computing systems 116a, 116b, 116c, ... 116n) deployed within computing system 100 and accessible by computing server system 114.

[0102] According to other embodiments, multiple facial expressions can be generated by... Figure 2 The character fusion module 212 processes the data. For example, the 3D geometry fitting step 505 performed by the facial model fitting module 212c can be performed by the character fusion module 212 to fit the MFM to an artist-created template facial geometry model corresponding to different expressions 504. In one embodiment, a gradient descent-based method can be applied to determine the MFM coefficients that minimize the vertex distance to the target mesh, and the resulting expression coefficients 507 can be stored in, for example, memory 216. These coefficients 507 can be combined with identity coefficients in the user's facial geometry 501 to reconstruct a triangular mesh of the user's face 512 corresponding to multiple facial expressions.

[0103] If the retopology step 508 is not performed, the user face texture 510 can be directly used as the blended face texture 515. Otherwise, texture migration 514 can be performed by resampling each face texture included in the user face texture 510 on the new user face geometry 513. This can be implemented on a CPU (e.g., processor 204 and / or the processor of the managed computing device or system 104, 106, or 108) using pixel bilinear interpolation, or on a GPU (e.g., processor 204 and / or the processor of the managed computing device or system 104, 106, or 108) by rendering a flattened face mesh using texture coordinates as vertex positions, and storing the output texture 515 in, for example, memory 216 or in any suitable data storage computing system (e.g., one of a plurality of computing systems 116a, 116b, 116c...116n) deployed within computing system 100 and accessible by computing server system 114. In this way, each facial texture included in the user facial texture 510 (e.g., facial albedo color texture and facial segmentation texture) can have a resampled version in 515 that is compatible with the new user facial geometry 513.

[0104] According to key aspects, the geometry fusion module 212b can be configured to perform geometry fusion 516 using the output obtained from the user's facial geometry 512 (multiple expressions) or 513 (neutral expressions) and the character's facial geometry 502 to generate fused facial geometry 517 (one or more template expressions) and 518 (neutral expressions), respectively.

[0105] In yet another embodiment, the geometry blending module 212b can be invoked multiple times to perform geometry blending 516 for different expressions, and the output can be converted into deformable targets by the facial animation module 212d. The facial animation module 212d can calculate the vertex offsets of the mesh corresponding to different expressions relative to the neutral expression mesh to create deformable targets 520. These deformable targets 520 can be stored, for example, in memory 216 or in any suitable data storage computing system deployed within computing system 100 and accessible by computing server system 114 (e.g., one of a plurality of computing systems 116a, 116b, 116c, ... 116n).

[0106] For each facial attachment model 503, the character fusion module 212 can be configured to determine a set of vertices on the attachment mesh whose distance from the surface of the original character facial mesh 502 is within a selected threshold, and project them onto the facial mesh surface as deformation handles 511. For example, after generating the fused facial geometry 518, the character fusion module 212 can calculate the new 3D position of the deformation handle on the new surface mesh, and use the geometry deformation step 519, implemented by the deformation module 212e, to calculate the 3D position of the remaining vertices of the attachment mesh, and store them as fused attachment models 521 in, for example, memory 216 or in any suitable data storage computing system (e.g., one of a plurality of computing systems 116a, 116b, 116c, ... 116n) deployed within computing system 100 and accessible by computing server system 114.

[0107] According to one embodiment, the geometry blending module 212b can output a blended facial model 312, which includes a blended facial geometry of a neutral expression 518, a blended facial geometry in a template expression 517, a facial transformation target 520, a blended facial texture 515, and a blended attachment model 521. The facial geometry corresponding to different facial expressions can be calculated from the blended facial geometry of the neutral expression 518 and the facial transformation target 520 using a selected deformation target animation (also known as blend shape animation) technique.

[0108] See now Figure 2 and Figure 6According to various aspects of this disclosure, an example geometric fusion process 600, which can be performed by the geometric fusion module 212b of the character fusion module 212, can combine the geometric features and idiosyncrasies of two faces to produce a fused facial model. In this document, "idiosyncrasies" can generally refer to unique and individual facial features or characteristics that distinguish one person from another. These can include subtle details such as the shape of facial features, eye position, jawline curve, or distinctive skin patterns.

[0109] According to some embodiments, the geometry fusion module 212b may take two main inputs: two input facial geometry shapes 601 and 602, and two auxiliary inputs: a blending weight map 603 and a set of anchor points 604.

[0110] The two input facial geometries 601 and 602 can be in the form of triangular meshes, each mesh being defined by a set of 3D vertices and a set of edges connecting them to form triangles.

[0111] In one aspect, the hybrid weight map 603 may include a set of values ​​defined on each vertex or face of the 3D mesh or on each pixel in texture space, based on a parameterization of the 3D mesh surface.

[0112] Anchor point 604 is a set of vertices on the triangular mesh of the input facial geometry 601, 602 whose 3D positions remain unchanged during the geometry blending process.

[0113] In one embodiment, the triangular mesh can be a manifold and share the same topology. Otherwise, a retopology step can be performed (e.g., Figure 5 The recurrent topology (508) can be transformed into a common manifold topology.

[0114] In the context of the character fusion implementation disclosed herein, the first input geometry 601 may include a 3D reconstructed mesh from a user photograph, as disclosed above, and stored, for example, on memory 216 or on any suitable data storage computing system (e.g., one of a plurality of computing systems 116a, 116b, 116c, ... 116n) deployed within computing system 100 and accessible by computing server system 114. The second input model 602 may include a character facial model digitally sculpted by an artist, which is converted into a triangular mesh and stored, for example, on memory 216 or on any suitable data storage computing system (e.g., one of a plurality of computing systems 116a, 116b, 116c, ... 116n) deployed within computing system 100 and accessible by computing server system 114.

[0115] Next, the facial deformation descriptor (FDD) can be locally defined on each facet of the triangular mesh as a weighted sum of two terms as shown in Equation (1). Global weights Used to adjust the relative importance of two terms in an FDD. (1)

[0116] First item The amount of stretching on a geometrically shaped surface can be measured and calculated by first identifying the rotation of each deformed triangle to align it as close as possible to its pre-deformation position and orientation, and then summing the differences in the corresponding side lengths. As disclosed above, a 3D mesh can include triangles (or other geometric elements) connected to form a surface. When the 3D mesh is deformed (e.g., bent, stretched, or compressed), the positions of the vertices change, thereby altering the shape and side lengths of the triangles. Stretch deformation can be used to measure how much the surface deviates from its original undeformed state in terms of side lengths. To separate the pure stretching effect (excluding rotation), each deformed triangle can be aligned with its original (pre-deformation) orientation using rotation. This alignment ensures that the measured difference is due to stretching rather than overall rotation. After alignment, the process measures the length difference between the corresponding side of the aligned triangle and its pre-deformation counterpart. Stretch can correspond to the difference in side lengths before and after deformation. To calculate the total stretching across the entire surface, these side length differences (after alignment) can be summed over all triangles in the mesh. This produces a global measurement of how much the mesh has been stretched from its original configuration. Determining the first term effectively separates stretching (e.g., changes in size or scale) from other transformations such as rotation, thereby ensuring accurate measurement of deformation by first aligning the triangles to their original orientation and then comparing the lengths of the corresponding sides.

[0117] According to one implementation method, the first term can be calculated using formula (2): (2)

[0118] In this article, F Triangles can be represented on the surface of a geometric shape. F k yes F All vertices on, , These represent the edges between vertices i and j in the deformed mesh and the original mesh, respectively. R k This indicates the rotation used to align the deformed edge with the original edge. It is the normalization coefficient for each face.

[0119] Second item E b The amount of curvature at local patches on a surface can be measured. This can be calculated by identifying an affine transformation to align the deformed triangle with its original position, then applying the same affine transformation to all its one-ring neighborhood triangles, and summing the differences in their corresponding side lengths.

[0120] In 3D surfaces or meshes, bending deformation generally refers to how much a shape bends or deviates in relative orientation between regions of adjacent triangles or meshes, rather than just focusing on stretching. A local patch or ring neighborhood refers to a set of triangles that share vertices or edges with a particular triangle. For a given triangle, this creates a local "patch" connecting the triangles. Affine transformations (e.g., scaling, rotation, translation, and shearing) can be computed to map the deformed triangle as close as possible to its pre-deformation position. Affine transformations capture how triangles move and deform as units. In some embodiments, the same affine transformation aligned with the first triangle can be applied to its neighboring triangles in a ring patch. This adjustment realigns the neighboring triangles to a corrected reference frame of the deformed triangle. After aligning the neighboring triangles to match the affine transformation of the deformed triangle, to identify the differences caused by the relative bending of the triangles in the patch, the next step is to compare the side lengths of these neighboring triangles with their corresponding portions in the pre-deformation mesh. The difference in side lengths between corresponding triangles (before and after deformation) can be summed across all triangles in the ring neighborhood. This total value can measure the bending deformation of the surface in this local area.

[0121] In other words, the second term measures how much local bending (curvature) occurs on the surface. Specifically, the affine transformation captures the global transformation and separates the relative changes between a triangle and its neighboring triangles. Applying this transformation to the neighborhood ensures that the comparison separates bending, rather than other types of deformation (e.g., stretching or rigid motion). The difference in side lengths quantifies how much the relative positioning of the triangles has changed, which is directly related to bending. The second term provides a local, curvature-sensitive measurement of deformation, complementing the stretching term for a comprehensive understanding of the mesh's behavior.

[0122] According to one implementation method, the second term can be calculated according to formula (3), where, F , , , It has the same meaning as in formula (2). N k Indicates the first k A ring-shaped neighborhood of the face. Tk This represents the affine transformation used to align the deformed edge with the original edge, and c k This represents the number of edges in the neighborhood. (3) like Figure 6 As shown, the FDD calculation module 606 can take the fused face geometry 605, in addition to the two original input meshes 601 and 602, as input. In one embodiment, the FDD calculation module 606 can calculate the deformation of the fused mesh for the two original input meshes and output the FDD for the first face 607 and the FDD for the second face 608, respectively. This calculation can be repeated in each iteration of the optimization process to reflect the updated vertex positions of the fused face mesh. In the first iteration, the fused mesh can be initialized using any reasonable mesh with the same topology. In fact, the FDD calculation module 606 can initialize it to be the same as the first input mesh.

[0123] Then, at each local patch on the surface, the FDDs of the first and second meshes can be combined (609) into a weighted sum, and the weights can be adjusted from the mixed weight map (603). Sampling is performed to generate the target FDD 610, as shown in formulas (4) and (5). In these two formulas, It has the same meaning as in formulas (2) and (3), and is marked with a superscript when appropriate. u and c To indicate whether the variable refers to the user or the typical grid. (4) (5)

[0124] Subsequently, the vertex position solver module 611 can use a quadratic energy minimization method with fixed variables to solve for the optimal vertex positions of the output fused mesh, which is the vertex configuration that minimizes the target FDD.

[0125] Because of the edge vector It is possible To calculate, where Represents the vertices of the mesh i The 3D position of the vertex. Equivalently, the vertex position solver module 611 can solve for the 3D position of the vertex. Minimized P .

[0126] Quadratic energy minimization generally refers to a mathematical method in which "energy" (representing a scalar value of the objective function) is minimized. The energy function is quadratic, meaning it involves squared terms of variables (e.g., distance, length, or other deformable measurements). In one embodiment, minimization can be performed using a direct linear least squares solver or a nonlinear iterative solver based on Lagrange multipliers. In another embodiment, the vertex position solver module 611 can use the input anchor point 604 and its position as fixed variables in the solver to provide constraints and boundary conditions. That is, during optimization, some variables (vertex positions) can remain fixed, serving as boundary conditions or anchor points, thereby ensuring stability and preventing accidental shifts in the mesh.

[0127] As the final step in an iteration, the merged face geometry 605 can be updated using the vertex positions output from the solver. Then, the process continues into the next iteration by feeding the merged face mesh back into the FDD computation module, as... Figure 6 As shown in the image.

[0128] Can be repeated Figure 6 This process continues until a predetermined number of iterations has been completed, or the threshold of the FDD value has been reached, whichever comes first.

[0129] Intermediate outputs during the process (including FDD, fused mesh) can be stored, for example, in memory 216 or in any suitable computing system (e.g., one of multiple computing systems 116a, 116b, 116c, ... 116n) deployed within computing system 100 and accessible by computing server system 114.

[0130] Figure 6 The final output of the process may include a fused facial mesh (e.g., Figure 3 The fused face model 312), the fused face mesh can be stored in, for example, memory 216 or in any suitable computing system deployed within computing system 100 and performing the optimization process (e.g., a cloud server as one of a plurality of computing systems 116a, 116b, 116c, ... 116n).

[0131] See now Figure 2 and Figure 7 The rendering module 214 can be configured to generate one or more personalized digital visual representations 716 based at least on the fused facial model 702. The personalized digital visual representations are generated via previous steps and pre-configured rendering assets 704 in multiple processes including but not limited to stylized texture generation 706, shading 710, line generation 712, and compositing 714.

[0132] The stylized texture generation step 706 can be performed by the stylized texture generation module 214a. This step can take a blended facial model 702 as input and produce a stylized texture 708 that preserves the user's personal characteristics as embodied in the blended facial model 702 while conforming to a predefined artistic direction. As used herein, the term "stylized" refers to conforming to a predefined artistic direction, which can include both realistic and non-realistic directions, such as cartoon coloring. The term "stylized" is not intended to exclude realistic rendering styles. In one embodiment, facial textures (such as facial albedo color textures) in the blended facial model 702 can be output without alteration, suitable for a realistic rendering style. In another embodiment, facial textures in the blended facial model 702 can be blurred to create a simplified stylized appearance, and example blur types can include Gaussian blur, box blur, median blur, and bilateral blur. The blur intensity can vary spatially and is controlled by a weight map stored as part of a pre-configured rendering asset 704. In yet another embodiment, guided texture compositing can be used. The compositing process works by selecting and rearranging pixel patches from a style source texture (stored as part of a pre-configured rendering asset 704) to fill a target texture, which is determined by several guide channels defined for both the style source and target textures. These guide channels can include various facial textures (e.g., facial albedo color textures, facial segmentation textures) that may be part of a blended facial model 702, shading intensities based on pre-configured lighting and material parameters, and other surface signals defined on the facial geometry specified by the blended facial model 702. Furthermore, compositing can be performed in texture space, or first in an image plane defined by pre-configured camera parameters and then back-projected into texture space. Additionally, a blending approach combining these methods can be used to transform the original texture into a desired style.

[0133] Subsequently, shading module 214b can perform shading step 710 to render the blended head geometry using a pre-configured shading style 704, optionally leveraging the stylized texture 708 generated in the previous step 706. In one embodiment, cartoon shading can be used, along with pre-configured color lookup tables, to independently stylize the diffuse and specular components from the shading calculations. These lookup tables can be stored as part of the pre-configured rendering asset 704 and can be used to achieve, for example, a cartoon shading appearance. The stylized texture 708 corresponding to the albedo color texture from the blended face model 702 can be used as the albedo color and blended with the cartoon shading output. In another embodiment, physically-based shading (PBR) can be used to create a more realistic effect. Physically-based PBR model lighting incorporates realistic material properties such as albedo, roughness, and metallicity. A stylized texture 708, corresponding to the albedo color texture from the blending surface model 702, can be applied to the albedo component, while the shading model can use pre-configured roughness and metallicity parameters to produce visually consistent results within a predefined artistic orientation. This can include applying non-physical or amplified lighting models to give a more illustrative or stylized appearance while retaining the benefits of physically based lighting, such as accurate reflections and material interactions. The combination of PBR and stylized textures provides a balanced approach to blending realism and artistic expression.

[0134] The line generation step 712, performed by line generator 214c, can render a line graph from the fused facial data to enhance the rendering. Several embodiments exist for implementing line generation 712. A first embodiment may employ the inverse shell method, which generates lines by rendering an inverted extended version of the fused facial model with backface culling enabled, followed by re-rendering the fused facial model as described in shading step 710. A second embodiment can explicitly generate different contour segments, including silhouette contours, suggestive contours, ridges, valleys, and apparent ridges, by analyzing the geometric and topological features of the fused facial mesh. Another embodiment may use image-space contour generation, which uses image processing techniques such as gradient analysis or depth buffer comparison to directly detect and extract edges and lines from the rendered image of the 3D model. The line width, color, and transparency of the generated lines 712 can be controlled by pre-configured parameter fields defined on the facial mesh of the personalized digital visual representation to achieve the desired stylistic effect. These parameters can be stored as part of a pre-configured render asset 704 and further adjusted at runtime based on factors such as the distance between the rendered 3D model and the camera to enhance visual clarity and style consistency.

[0135] Subsequently, the synthesis module 214b can execute... Figure 7 The compositing step 714 integrates the outputs from shading 710 and line generation 712, compositing them on one or more background layers specified by pre-configured assets 704. For example, the background layers may include static images or animated layers that combine effects such as sub-picture animation, particle effects, vertex animation, and screen-space reflections. Compositing step 714 may also render a pre-configured 3D scene and use the result as a background layer. This supports a 2.5D visual style that combines the depth and parallax effects of a 3D environment with the illustrative texture of a rendered personalized digital visual representation. Compositing step 714 can use different image blending operators (such as darken, screen, overlay, and blending) to merge its individual image layers, allowing flexible control over the final visual style of the rendered personalized digital visual representation 716.

[0136] The aforementioned stylized texture generation step 706, shading step 706, and line generation step 706 can all utilize facial geometry specified by the blended facial model 702. Facial geometry corresponding to different facial expressions can be calculated by the blended facial model as previously described. Thus, the facial geometry used in these steps can represent different facial expressions, which can be specified by a pre-configured rendering asset 704 or determined by other inputs, such as a combination of user input and pre-configured rendering asset 704.

[0137] As previously seen Figure 3 As described, the character fusion module 212 of the computing server system 114 can be selected and adjusted by the end user or artist for each local patch on the output facial mesh to determine the degree of geometric similarity between the two input meshes that respectively represent the facial geometry of the character designed by the user and the selected artist.

[0138] In some implementations, artists can use vertex painting features from 3D modeling software packages, such as Blender, to assign a number x between 0 and 1 to each vertex on the face mesh. In other implementations, artists can use any 2D or 3D painting software, such as Adobe Photoshop, Substance Painter, etc., to paint colors on a texture map of the face mesh, and the system converts each color into a value x between 0 and 1 based on the color map. In both implementations, x closer to 0 indicates that the local patches around that vertex should resemble the user's face more, while x closer to 1 indicates that it should resemble the character's face more. These values ​​of x are stored along with the character's face model as a per-vertex or per-pixel blend weight map.

[0139] According to various aspects of this disclosure, Figure 8The first example graphical user interface (GUI) 800 is shown, which is used in the artist's computing device (e.g., Figure 1 An application downloaded to at least one of computing devices or systems 104, 106, or 108 is associated with controlling the fusion result. In GUI 800, the face can be segmented into multiple predefined regions corresponding to recognizable features such as the mouth 801 and the left eye 802. It can also be a single region corresponding to the entire face. GUI 800 can be configured to generate and display individual UI (User Interface) elements to allow user interaction with each of the predefined regions 801, 802, respectively. For example, an artist can use one of the sliders 803 to control a selected region, making the region more similar to the user's face by moving the slider closer to the left end, or by moving the slider closer to the right end to make the region more similar to the character's face. The slider position can be assigned to all vertices in the corresponding region with values ​​between 0 and 1, and stored together with each character as a per-vertex blending weight map 603.

[0140] According to other aspects of this disclosure, Figure 9 A second example GUI 900 is shown, which is used in conjunction with the user's computing device (e.g., Figure 1 The application is associated with a computing device or system (at least one of 104, 106, or 108) for adjusting and controlling parameters related to the similarity between the facial model and the user's face or the face of a selected character.

[0141] In one embodiment, such as Figure 9 The GUI 900 shown can be configured to display a user's facial model 901, a facial model 902 of a selected character designed by an artist, and UI elements (e.g., a slider 903 or any other suitable UI element) for controlling the similarity between them. For example, received user input can move the slider 903 toward the user's facial model 901 or the selected character's facial model 902 to dynamically align the parameters of the resulting fused model with the parameters of the user or the selected character. In some implementations, the GUI 900 can display the resulting fused model in real time based on continuous user input via the slider 903, which is particularly user-friendly for touch interfaces such as mobile devices. The slider 903 allows for real-time, smooth adjustments, providing immediate feedback on changes. In embodiments, the position of the slider 903 can be converted into a fractional value between 0 and 1 and multiplied by... Figure 6 Hybrid weight graph 603 in the fusion FDD calculation module 609.

[0142] Figure 10 , Figure 11 and Figure 12An example user selfie is shown as the input to the facial reconstruction process 304 described above (e.g., Figure 3 User photos or videos (302). Specifically, Figure 10 It displays a front view image of the user's face. Figure 11 The image shows the same user's face taken from a left-hand perspective, and... Figure 12 The image shows the same user's face taken from a right-hand perspective. Figure 13 An example user face model generated by the face reconstruction module 210 according to several aspects of this disclosure is shown (e.g., Figure 3 User facial model 306).

[0143] See now Figure 14 It showcased the work of Figure 5 The example facial geometry specified in character facial geometry 502. Figure 15 An example facial geometry (neutral expression 513) obtained from user facial model 306 is shown. Figure 16 An example of a blended facial geometry (neutral expression 518) generated by the character blending module 212 is shown. Notably, the blended facial model incorporates the geometric features of the eyes, nose, and lips from the user's facial model. Figure 15 ), while also employing geometric features derived from the character's facial geometry, including the cheeks, jawline, and neck. Figure 14 This demonstrates that the hybrid weight map 603 can be used to control both the region of the facial geometry that resembles the user or character and the degree of similarity within each region. Figure 17 An example of a blended facial geometry (in template expression 517) is shown, specifically corresponding to a fake smile expression generated by character blending module 212.

[0144] Figure 18 and Figure 19 Examples of personalized digital visual representations generated by the rendering module 214 based on different pre-configured rendering assets 314 are shown. Figure 18 Corresponding to a more cartoonish art direction, and Figure 19 This demonstrates a more realistic artistic direction, but with stylized visual elements such as hatching in shadow areas. These examples highlight the ability to finely control various aspects of the rendering, including rendering style, background, lighting conditions, and perspective. It should be noted that... Figure 13 , Figure 15 , Figure 16 , Figure 17 , Figure 18 and Figure 19 All of them showcased based on Figure 10 , Figure 11 and Figure 12 The result shown is generated from a user's selfie. Therefore, Figure 18 and Figure 19 The rendered face in the image can still be recognized as Figure 10 , Figure 11 and Figure 12 The user in the selfie shown.

[0145] Those skilled in the art will recognize that the various embodiments of this disclosure may be implemented in other specific ways than those set forth herein without departing from the scope and essential characteristics of the various embodiments of this disclosure. Therefore, the above embodiments are to be interpreted in all respects as illustrative rather than restrictive. The scope of this disclosure should be determined by the appended claims and their legal equivalents, not by the foregoing description, and all variations falling within the meaning and equivalents of the appended claims are intended to be included therein. It will be apparent to those skilled in the art that claims not expressly referenced in the appended claims may be combined to present embodiments of this disclosure, or may be included as new claims by subsequent amendments after the filing of this application.

Claims

1. A system deployed within a communication network for generating one or more personalized digital visual representations, the system comprising: The first computing device includes: A first non-transitory computer-readable storage medium is configured to store an application; and A first processor, coupled to the first non-transitory computer-readable storage medium and configured to execute instructions of the application to obtain one or more photos or videos of the user; and The second computing device includes: A second non-transitory computer-readable storage medium; and A second processor, coupled to the second non-transitory computer-readable storage medium, is configured to: Receive one or more photos or videos from the user via a first application programming interface (API) call. Process the one or more photos or videos to at least determine the facial geometry of the user's facial model. Obtain at least the facial geometry of a selected typical character from the character's facial model. Generate a fused facial model to preserve the facial geometry of both the user's facial model and the character's facial model. Generate at least one or more personalized digital visual representations based on the fused facial model. The parameters related to the fused facial model are stored on the second non-transitory computer-readable storage medium, and The one or more personalized digital visual representations are transmitted to the first computing device via a second API call. The first processor of the first computing device is further configured to execute instructions of the application to receive and display the one or more personalized digital visual representations on the display interface of the first computing device.

2. The system of claim 1, wherein, The facial geometry of the user's facial model includes multiple three-dimensional 3D meshes, each of which includes at least data defining the connectivity between vertices, edges, and the face in each of the multiple 3D meshes.

3. The system of claim 1, wherein, The facial geometry of the user's facial model includes parameters related to multiple facial expressions.

4. The system of claim 3, wherein, The second processor is also configured to process the one or more photos or videos to determine parameters related to facial texture.

5. The system of claim 1, wherein, Each of the fused facial model, the user facial model, and the character facial model includes multiple 3D meshes, and the second processor is further configured to allow adjustment of the geometric similarity between selected regions of the fused facial model and corresponding regions of the user facial model or the character facial model.

6. The system of claim 5, wherein, The second processor is also configured to determine the geometric similarity based at least on a local geometric descriptor including a facial deformation descriptor.

7. The system of claim 6, wherein, The second processor is also configured to use a hybrid weight map to control the geometric similarity between selected regions of the fused facial model and corresponding regions of the user facial model or the character facial model.

8. The system of claim 6, wherein, The facial deformation descriptor measures the local stretching and bending of each facet on the surface of each 3D mesh.

9. The system of claim 6, wherein, The second processor is also configured to determine multiple vertex positions of the fused facial model by minimizing the sum of the blended target facial deformation descriptors on the facial surface.

10. The system of claim 1, wherein, The second processor is also configured to generate parameters representing at least one of a plurality of pre-configured rendering assets, representing at least one of stylized textures, shadows, and lines, and to integrate at least one of the stylized textures, shadows, and lines with one or more background layers selected from the plurality of pre-configured rendering assets to render the one or more personalized digital visual representations.

11. A computing server system deployed within a communication network for generating one or more personalized digital visual representations, the computing server system comprising: Non-transitory computer-readable storage medium; as well as A processor, coupled to the non-transitory computer-readable storage medium, is configured to: The user receives one or more photos or videos from a computing device deployed within the communication network via a first application programming interface (API) call. Process the one or more photos or videos to at least determine the facial geometry of the user's facial model. Obtain the facial geometric features of at least several typical characters from the character's facial model. Generate a fused facial model to preserve the facial geometry of both the user's facial model and the character's facial model. Generate at least one or more personalized digital visual representations based on the fused facial model. The parameters related to the fused facial model are stored on the non-transitory computer-readable storage medium, and The one or more personalized digital visual representations are transmitted to the computing device via a second API call.

12. The computing server system of claim 11, wherein, The computing device is configured to receive the one or more personalized digital visual representations and render the one or more personalized digital visual representations on the display interface of the computing device.

13. The computing server system of claim 11, wherein, The facial geometry of the user's facial model includes multiple three-dimensional 3D meshes, each of which includes at least data defining the connectivity between vertices, edges, and the face in each of the multiple 3D meshes.

14. The computing server system of claim 11, wherein, The facial geometry of the user's facial model includes parameters related to multiple facial expressions.

15. The computing server system of claim 11, wherein, The processor is also configured to process the one or more photos or videos to determine parameters related to facial texture.

16. The computing server system of claim 11, wherein, Each of the merged facial model, the user facial model, and the character facial model includes multiple 3D meshes, and the processor is further configured to allow adjustment of the geometric similarity between selected regions of the merged facial model and corresponding regions of the user facial model or the character facial model.

17. The computing server system according to claim 16, wherein, The processor is also configured to determine the geometric similarity based at least on a local geometric descriptor including a facial deformation descriptor.

18. The computing server system according to claim 17, wherein, The processor is also configured to use a hybrid weight map to control the geometric similarity between selected regions of the fused facial model and corresponding regions of the user facial model or the character facial model.

19. The computing server system according to claim 17, wherein, The facial deformation descriptor measures the local stretching and bending of each facet on the surface of each 3D mesh.

20. The computing server system according to claim 17, wherein, The processor is also configured to determine multiple vertex positions of the fused facial model by minimizing the sum of the blended target facial deformation descriptors on the facial surface.