Create and consume 3D animation

By employing hierarchical AI/ML models with parametric 3D curves and real-world causality, the limitations of 2D-based solutions are overcome, enabling efficient and cost-effective 3D avatar creation and animation for content creation and human-computer interfacing.

WO2025217682A1PCT designated stage Publication Date: 2025-10-23HUANG DEAN
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/AU2025/050378
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-25
Filing Date
2025-04-16
Publication Date
2025-10-23

AI Technical Summary

Technical Problem

Conventional AI/ML modeling heavily relying on 2D visuals results in high costs and limited usage, with existing solutions for digital avatars in content creation and human-computer interfacing being suboptimal.

Method used

The development of hierarchical AI/ML models utilizing parametric 3D curves and real-world causality for creating and animating avatars, which are trained autonomously and interpretable, allowing for efficient content creation, authenticity verification, and human-computer interfacing.

Benefits of technology

The solution provides high-efficiency, cost-effective, and interpretable 3D animation models that can be deployed across various devices, facilitating realistic and autonomous avatar animation with reduced training costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure AU2025050378_23102025_PF_FP_ABST
    Figure AU2025050378_23102025_PF_FP_ABST
Patent Text Reader

Abstract

Methods, systems, computer algorithms are presented for creating hierarchical information about static and animatable avatars on devices, and consuming the information in various applications. Models may be trained by utilizing parametric 3D curves, real-world causality, and AI / ML, etc. Trained hierarchical models include information representing characteristics and animation of subjects of one kind, and connections or patterns for subjects of different kinds, to offer high levels of efficiency. The training process can be autonomous, and trained models can be interpretable and traceable. The trained hierarchical models can be used in content creation, authenticity verification, and human-computer interfacing. Further 3D static and animatable information described herein can be used for generic AI / ML modelling, as alternatives to conventional AI-modelling that heavily relies on 2D visuals.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] TITLE: CREATE AND CONSUME 3D ANIMATION

[0002] BACKGROUND OF INVENTION - FIELD OF INVENTION

[0003] Methods, systems, computer algorithms are presented in creating and consuming 3D- animation, in the fields of generic AI / ML modelling, content creation, authenticity verification, human-computer interfacing, etc.

[0004] BACKGROUND OF INVENTION

[0005] Conventional AI / ML modelling heavily relying on 2D visuals contributes to high-costs and limited usage. Digital avatars may be used in content creation, human-computer interfacing, etc. Many current solutions are not optimal. Their drawbacks and disadvantages will become apparent from a consideration of the rest of this document.

[0006] BACKGROUND OF INVENTION - OBJECTS AND ADVANTAGES

[0007] This document discloses methods, systems, computer algorithms for creating hierarchical information about static and animatable avatars on devices, and consuming the information in various applications.

[0008] Models may be trained by utilizing parametric 3D curves, real-world causality, and AI / ML, etc. Trained hierarchical models include information representing characteristics and animation of subjects of one kind, and connections or patterns for subjects of different kinds, to offer high levels of efficiency. The training process can be autonomous, and trained models can be interpretable and traceable. The trained hierarchical models can be used in content creation, authenticity verification, and human-computer interfacing. Further 3D static and animtable information described herein can be used for generic AI / ML modelling, as alternatives to conventional Al-modelling that heavily relies on 2D visuals.

[0009] Still other objects and advantages will become apparent from a consideration of the ensuing description and drawings.

[0010] DRAWINGS - FIGURES

[0011] FIG. 1 illustrates a high-level workflow for creating and consuming 3D-animation, by one or more processors, according to some example embodiments.

[0012] FIG. 2 is a workflow for Al-modeling static avatars.

[0013] FIG. 3 is a workflow for Al-modeling animatable avatars.

[0014] FIG. 4 illustrates a coordinate system that the rest of this document is consistent with.

[0015] FIG. 5A illustrates a system of using Artificial Neural -Network (ANN) to create info about 3D-animation.

[0016] FIG. 5B illustrates a sub-system of the controlling-curves layer in FIG. 5A.

[0017] FIG. 6 illustrates a system for utilizing video contents in Al-modelling avatars and other aspects of this document, according to some example embodiments.

[0018] FIG. 7 illustrates a type-hierarchy that can be created and consumed.

[0019] FIG. 8 is a workflow of a method for an individual 3D-subject being fed through the learning phase during Al-modelling.

[0020] FIG. 9A is a flowchart of a method for building up the type-hierarchy.

[0021] FIG. 9B is a flowchart of a method for creating type-wide info for a unit in a low layer of the type-hierarchy.

[0022] FIG. 10 illustrates a system for safeguarding the quality of the animation of avatars.

[0023] FIG. 11 is a flowchart of a method for converting a 3D-subject to an animatable 3D-avatar.

[0024] FIG. 12 illustrates a system for creating visual contents from descriptive info such as texts.

[0025] FIG. 13 is a workflow of a method of converting descriptive info to visual contents.

[0026] FIG. 14 is an architecture of a system for verifying authenticity of visual contents. FIG. 15 illustrates the modules of the content verifier in FIG. 14.

[0027] FIG. 16 is a workflow of a method for adaptively verifying authenticity of visual contents.

[0028] DETAILED DESCRIPTION

[0029] DEFINITIONS of terms in this document

[0030] An avatar means a virtual character. An avatar in this document includes at least the face part of a 3D-character. Avatars are 3D models if not otherwise stated. To model and animate a single avatar, the origin of its local coordinate-system can locate at the geometric centre of the avatar. An avatar can be created by multiple ways, including from scanning a 3D-object or converting from multiple 2D-photos via photogrammetry.

[0031] A static avatar refers to an avatar that can’t be animated yet, due to the lack of a mechanism or ability for being animated. An animatable avatar is an avatar that is ready for animation, since it has a mechanism for being animated. An animatable avatar may be able to animate on contextual events or signals (denoted as signals in this document), such as expressions, moods, speech or messages, with little or non-human intervention. During the learning phase, the AI / ML modeling of the animatable avatars shall be continuously improved; the phase can use intermediateanimation for checking quality of the AI / ML modeling. Intermediate-animation means animated avatars using parameters that are still being trained. Required-quality means quality of animation that can meet the needs of different applications, e.g. realistic animation, or cartoon-like, fantasized animation. Afterwards, an animatable avatar may be able to be animated autonomously and the animation may approach the realism of a real-world subject.

[0032] Corresponding means element-to-element or point-to-point correspondence in 3D-space. The correspondence may exist between shape vertices and texture vertices and across different avatars. For example, a system or software can make a specific shape vertex (X, , T, , Zs) locate at a certain spot (e.g. the tip of the avatar’s nose), while another vertex Si+d(Xi+d, Yi+d, Zi+d) locate at another spot (e.g. a comer of the mouth). This corresponding relationship remains when an avatar is being animated. The correspondence may be automatically achieved by an API, e.g. photogrammetry APIs that build a 3D-model from a photo (See “Photogrammetry" section in References).

[0033] Feature points of an avatar refer to representative 3D-points of the avatar. The full set of feature-points shall be sufficient for AI / ML systems or software program to fulfill functionalities as disclosed in this document. Statistics about the feature-points may refer to statistic relationship among the feature-points. For example, the distance between centroids of two eyes is what fraction (ee / W) of the width of the whole head, assuming the centroids are feature points. The values of ee / W and other similar ones are denoted as Measurements of Avatars in this document. The statistics can be calculated or trained across a common type of subjects. Sometimes an application will need to make an avatar look more attractive than it is originally. Also, an application may change an avatar from one type to another, e.g. from a human character to a fantasized character. Pattern information for creating such changes, whether within one type, or across different types of subjects may be identified. The patterns may be identified on the basis of the above statistics, thus the pattern-information may be added as a component of the statistics information in a dataset. For example, to change a human avatar from a Caucasian to African then to an Asian, an AI / ML learning process will aim to find what patterns across the statistics of these races can be identified.

[0034] DETAILED DESCRIPTION OF EXAMPLE EMBODIMENTS

[0035] Various drawings, descriptions and equations merely illustrate example embodiments of the present disclosures and cannot be considered as limiting its scope. As will be appreciated, various details may be application-dependent (e.g. whether realistic or cartoon-like), device-dependent (e.g. whether running on desktop or mobile devices), and effectdependent (e.g. whether modelling weak or strong emotions).

[0036] FIG. 1 illustrates a high-level workflow for creating and consuming 3D-animation, , by one or more processors, according to some example embodiments.

[0037] The workflow first Al-models static avatars 102. Next, the workflow Al-models animatable avatars 104.

[0038] Then, the results of the Al-modeling can be consumed in various applications 106, e.g. animating avatars and interacting with users on devices, or create videos or images from text descriptions, verifying content authenticity.

[0039] FIG. 2 is a workflow for Al-modeling static avatars (i.e. 102 in FIG. 1). The goal may include building statistics, and identifying patterns.

[0040] The workflow may first collect data of 3D-subjects that shall include both geometry and texture data. Laser 3D-scanners can be used to collect these data. The collected data are saved in database.

[0041] A static modelling trainer may continuously improve Al-modeling of the static avatars until quality requirements are met. In static modelling trainer, the workflow locates feature-points of each scanned subject. Feature-points may be automatically located, e.g. by using the grid described later in this document.

[0042] The workflow then can build statistics among these feature-points across a large number of subjects. The type of the subjects may be chosen so that AI / ML-modeling of such a type converges in a regression, by feeding a number of subjects belonging to that type. For example, choosing a human race (e.g. Caucasians) and limit its subjects to certain age-group, gender or their combinations may be easier to reach regression than choosing all human-subjects. Regression models may be chosen from parametric, semi-parametric or non-parametric regression.

[0043] Next the workflow identifies patterns. The patterns may be identified using regression methods. For example, statistics regarding ee / W, among others, may exhibit regularity when changing from one human-race to another, or from an average looking person to a highly attractive-looking person.

[0044] With the database collected from the 3D-subject-data, the flowchart shall recursively improves statistics and patterns of Al-models, by being fed data of 3D-subjects. The learning may reach a satisfied state (including convergence) when both the following conditions can be met:

[0045] 1. New 3D-subject-data (e.g. subjects not yet used in the training phase) can fit with the statistics and patterns. If a software program supports correspondence, it can quickly find feature-points on new subjects, calculate statistics, and check if their measurements fit well with the statistics information.

[0046] 2. The full set of feature-points are able to determine elements of animatable avatars mathematically, to generate Controlling-information for animatable avatars that can meet quality requirements. More details are given later in this document.

[0047] The workflow may finalise static modeling when the above two conditions can be met, by saving statistics and patterns of Al-models in data files on devices. Static modelling trainer may also receive requests for retraining-static-models from other methods, systems elsewhere in this document, and requests may cause the static modelling trainer to further recursively improve the static models.

[0048] FIG. 3 is a workflow for Al-modeling animatable avatars (i.e. 104 in FIG. 1).

[0049] An animatable modelling trainer 304 may recursively improves Al-modeling of the animatable avatars until quality requirements are met, by building animation mechanism on each static avatar (e.g. one that was fed through FIG. 2). The animatable modelling trainer may autonomously process the following steps: • build curves on a 3D-avatar that may include controlling-hierarchy among the curves;

[0050] • connect controlling curves with the avatar; and

[0051] • modify the curves thus animate the avatar.

[0052] The workflow uses 3D-curves to control the animation of 3D-avatar. A 3D-curve can be represented parametrically. For example, Non-Uniform Rational B-Spline (NURBS) curve is a general parametric curve. A NURBS curve is represented as a function of a free-parameter u , (1) where { P are control points, {w are weights of the control points, and {p(u )] are B-spline basis functions which are defined on a sequence of knots {n . The NURBS curve is determined by its control points and associated weights. Cubic curves may be used in the workflow.

[0053] Among many useful properties of a NURBS curve, the localness (or “local support”) means that modification of a control point (or its weight) only affects a local section of the curve, and affects more substantially a section closer to the control point than another section further away. NURBS curves may be created so that the localness property is satisfied, e.g. by using the averaging method to create knot vectors of the curves. The NURBS curve (including the localness property) is explained in detail in math literature

[0054] Training data for animatable avatars (e.g. 302 in FIG. 3) may be organized on contextual signals that determine the animation of the avatars. The contextual signals may depend on elementary biological processes. For examples, it is understood from anthropology and biology that all humans have six primary expressions. Thus, the workflow may assign an ID to each elementary signal and store the training data matching elementary signals in a relational database accordingly.

[0055] Then the workflow compares intermediate-animation with matching animation of required- quality 312. Intermediate-animation may be from animated avatars that are controlled by curves. Matching animation may be from videos captured live or another application or human modeler. The loss / error is back-fed, thus the system can continuously improve Al-modeling animatable avatars. As will be appreciated, various machine learning techniques may be used to update the animatable modelling, in light of the back-propagated losses.

[0056] In some example embodiments, the comparison can be organized on areas of a 3D-mesh surface that reflect real-world mechanism of movement. For example, an area (of a subject’s face) that is controlled by a muscle can be a unit of such comparisons; and info about real-world anatomy can be consulted to set such areas.

[0057] The animatable modelling trainer may continuously train on data from different 3D-subjects of a type or comparable types. The animation results for an individual avatar and type(s) may all be checked against quality requirements. The workflow may finalize animatable modeling when the quality requirements can be met, by saving information for animatable modeling in data files on devices. The animatable modelling trainer may also receive requests for retraining animatable-models from other methods, systems elsewhere in this document. The requests may cause the animatable modelling trainer to further recursively improve Al-modeling of animatable avatars.

[0058] Detailed implementations for FIG. 1 ~ 3 will be given in the rest of this document.

[0059] FIG. 4 illustrates a coordinate system that that the rest of this document is consistent with.

[0060] The local coordinate system can represent the 3D-head-model of an individual avatar. There shall be a cuboid whose 6-planes are parallel to the x - y, y - z , x - z planes respectively, and be the minimum cuboid that can hold the 3D-avatar within it. The x, y , z values within the cuboid can be limited to a range that is consistent for all avatars of a type or comparable types. Each 3D-model may be modelled inside the minimum cuboid, and the cuboid can have multiple imaginary (auxiliary) lines that are parallel to the x, , z axes respectively. These auxiliary lines can make up an auxiliary grid inside and around the 3d model. The overall setting, including the local coordinate system, the minimum cuboid and the auxiliary grid, can be regarded as an individualised coordinate system of the avatar.

[0061] To make coefficients and vectors consistent and comparable across avatars of comparable types, 3D-meshes of the avatars shall be oriented, arranged, spread out and measured consistently in the individualised coordinate system. Further, control curves can be built consistently across avatars of the types.

[0062] For examples, when Al-modelling avatars that target the same kind of devices, 3D-meshes representing avatars of type(s) may all have the same number of vertices, and the 3D-meshes may orient and spread out in the same fashion when measured relative to their respectively individualised coordinate systems; ee / W, among other necessary Measurements, their featurepoints shall be located and their distances shall be measured, all consistently across avatars of the types; further the 3D-meshes may be controlled by the same set of curves, and corresponding curves may all have the same amount of control points and be sampled the same way, e.g. separated by the same intervals of parameter u in their corresponding sections.

[0063] Afterwards, Measurements of static avatar and Controlling-information of animatable avatars may all be normalised, by taking values of avatars from their respectively individualised coordinate system. Further, the auxiliary grid can be built consistently across avatars of type(s), and the grid can composite of multiple auxiliary points, and these points can be candidates of feature-points, thus the feature-points (e.g. in FIG. 2) can be chosen autonomously from the auxiliary points.

[0064] FIG. 5A illustrates a system of using Artificial Neural -Network (ANN) to create info about 3D-animation, according to some example embodiments.

[0065] There may be theses layers in the ANN: initial mesh, feature points, controlling-curves, and animated mesh.

[0066] The first layer of the ANN may have the details of initial mesh, i.e. 3D mesh-data of an avatar when it is initially unanimated or at a neutral (e.g. with a relaxed, emotionless and silence) pose, including all 3D-coordinates of the mesh vertices. Each node in this layer can have the data of at least 3D-coordinates (e.g. x, y, z) of a single vertex.

[0067] The next layer or layers may have the details of feature points, and feature points can be individualised for each different avatar. Each node in this layer can have the data of 3D- coordinates (e.g. x, y, z) of a single feature point.

[0068] Those skilled in the art will appreciate the nodes in different ANN layers are connected by weights. In this example configuration, the geometric locations of elements at one layer will determines locations of elements at next layer. Thus, via the trained weights, the initial mesh will determine the feature points on or near the surface of the mesh.

[0069] Each control point of controlling-curves may be determined by multiple feature points, as:

[0070] CPi = fcxi * Fx) (2) where fcxi represents a coefficient parameter that determines the influence of feature point Fx on the control point CPi. Equation 2 shows each control point can be determined by multiple feature points, and represents control points and feature points as vectors.

[0071] These control points then create the NURBS curve that has sample points S;, Smand Snetc, which are points separated by intervals of parameter u (see Equation 1).

[0072] The movement of vertices (e.g. Vo, Vp, Vq , on the 3D-mesh representing an avatar) can be controlled by sample points of the NURBS curves. For example, the movement of a vertex V can be represent via this equation:

[0073] T vEE s.s (3) where Snis a sample point on a NURBS curve that controls or influences the movement of a vertex V; Snis the new position of Snafter the curve is modified (Snand Snthus have the same value for parameter u in Equation 1); and V is the updated position of V; coefficient parameter rnrepresents the influence of sample point Sn’s movement on the vertex V. Movement of a vertex V is composited of a sum of vectors since the movement can be affected by a number of sample points that may be on one or more curves.

[0074] The coefficients in Equation 3 may be functions of one or more of elements including time, contextual signals, the kind of devices on which avatars are animated.

[0075] Since feature points may determine control points of NURBS curves (e.g. as in Equation 2), the next layer 506 of the ANN may have the details of the curves, including their control points and sample-points. The NURBS curve may be created via Equation 1 while the avatars are at a unanimated or neutral pose.

[0076] Since the curves can control the animation of 3D avatars, the subsequent layer after the controlling-curves layer may have the information for animated mesh, i.e. 3D vertices of an avatar after being animated, including all 3D-coordinates of the mesh. Each node in this layer can have the information of at least 3D-coordinates (e.g. x, y, z) of a single vertex after the mesh is animated.

[0077] FIG. 5B illustrates a sub-system, or more details of the controlling-curves layer 506 in FIG. 5 A. The controlling-curves layer 506 can further has a hierarchy among different curves. Utilising the hierarchies may create more realistically complex animation, among other benefits.

[0078] The controlling hierarchy may have:

[0079] • Control points of a first curve

[0080] • Curve points of the first curve

[0081] • Control points of a second curve

[0082] • Curve points of the second curve

[0083] • More curves controlled by the above curves if needed

[0084] As illustrated, a first NURBS curve has control points CPU, CPlj, etc. These control points create the NURBS curve that has the sample points Sim, Sin, etc.

[0085] The sample points can control or influence the location of another NURBS curve, e.g. determining control points of a second NURBS curve, via this equation: where S lnis a sample point on the first NURBS curve that influences the movement of a control point CP 2tof a second curve; S lnis the new position of S lnafter the curve is modified. Coefficient parameter ch 12ndetermines the influence of sample point S lnon the control point CP 2fof the second curve. Movement of control point CP 2fis composited of a sum of vectors since the movement may be affected by a number of sample points on other layer(s).

[0086] Afterwards, the control points of a second NURBS curve will create the second curve, which has samples points S 2n, S 2m, etc. This hierarchy between the curves can continue.

[0087] The controlling relationship among the NURBS curves may simulate, or extend the causality of the real world. For example, if a muscle A controls or influences another muscle B, and if muscle A’s effect is simulated by a curve X, and muscle B’s effect is simulated by a curve Y, then the system can use curve X to control curve Y. This controlling relationship may be set by a developer during bootstrap phase of the learning process, or set in a text file, or automatically searched and tested by using techniques such as decision-trees or random forests, etc.

[0088] In some example embodiments, the movement of mesh vertices on a 3D-avatar may be controlled by a curve or curves of the controlling hierarchy, and the 3D-mesh of an avatar is animated after controlling-curves have been modified. For example, sample points such as Snin Equation 3 may be controlled by one or more curves that are situated in the controlling hierarchy. Info about real-world anatomy can be consulted, to decide what curves in the controlling hierarchy may directly animate mesh vertices in their neighbourhood, in consideration of disclosures of this document.

[0089] The details of the hierarchy, including relationship among curve elements, and among the elements and mesh-vertices, may depend on contextual signals and other factors (e.g. application, device and effect-dependent). When these factors vary, the relationship among the elements may also vary, as their statistics and patterns.

[0090] Conventional ANN normally have a large number of connecting weights between layers. This may contribute to high-cost in learning and limit its wide-spread usage. The disclosures in this document may limit, prune, ignore or truncate (denoted generally as truncation in this document) unnecessary connecting-weights, thus may facilitate wide-spread usage. There can be two ways to achieve this goal - neighbourhood and relevance truncations.

[0091] Neighbourhood truncation may connect weights within certain geometric neighbourhood, and truncate weights outside the neighbourhood. For example, some feature points, such as the tops of two out-ears, may only be determined by mesh-vertices near their geometric neighbourhood. Thus the system can set distance thresholds to limit weights connecting feature points to meshvertices in their neighbourhood, and truncate weights connecting mesh-vertices further away.

[0092] Relevance may connect weights if they are relevant to real-world situations or desired effects of animation. For example, if a muscle A controls or influences another muscle B and doesn’t influence muscle C; and if muscle A’s effect is simulated by a curve X, and muscle B’s effect is simulated by a curve Y, and muscle C’s effect is simulated by a curve Z, then the system can consider connecting weights between curve X and curve Y, and truncate weights connecting between curve X and curve Z.

[0093] By making use of the localness of the NURBS curve, the system can trace which upstreamcurve controls or more strongly influences which section of a downstream curve; and the tracing info can be built into the hierarchy in FIG. 5B, until the hierarchic tracing reaches controlled mesh vertices of 3D-avatars. As will be appreciated, the tracing can be created at detailed levels, e.g. among control points and sample points of different curves. Thus, while being trained, the system can build this tracing information and store into suitable data-structures (e.g. a searchtree), to effectively extend the correspondence on avatar-meshes (See Definitions section) to the whole ANN. For example, after the tracing info is built and stored, given any vertex on an avatar mesh, the system can retrieve which upstream control parameter(s) ultimately determine the animation of the vertex, and can retrieve their degrees of influence. The tracing info may be stored in searchable trees in a top-down fashion, e.g. elements of the first curve are placed near the top of a searchable tree; then retrieving the tracing info (from mesh vertices) may use optimal algorithms to traverse the tree bottom-up. Thus, the Al-models can be interpretable and traceable, thus may be trained much more quickly, and at much lower costs, and may be suitable for being deployed to various devices including resource-restrained devices.

[0094] It is noted the ANN (including connecting weights) may be trained on a number of signals, such as elementary expressions, moods and visemes. Thus after the learning phase, when given any one, or sequences, or compositions of the elementary signals, the avatar may be immediately animated, by fetching or blending relevant trained parameters (e.g. parameters as in Equation 4), with little or non-human intervention. Blending of parameters may simulate or extend the causality of the real world, e.g. how elementary expressions, moods and visemes can composite or combine into complex facial movements; and qualitative or quantitative info of such composition may be found in biology or anatomy, or to be precisely quantified via additional computer vision systems.

[0095] FIG. 6 illustrates a system for utilizing video contents in Al-modelling avatars and other aspects of this document, according to some example embodiments. This document uses Generative Adversarial Network (GAN) to create animatable avatars and check the quality of their animation. GAN has Generator and Discriminator networks that can be application, device and effect-dependent.

[0096] The Generator network (or generator) can aim to create realistically animated 3D-avatars if needed. The Discriminator network (or discriminator) may try to tell real-world 3D-subjects (X) from an animated 3D avatar (X ) that is created by the Generator.

[0097] The generator may be implemented using the details disclosed in this document. For example, the system can build a controlling hierarchy detailed in FIG. 5A and 5B; the system may use Equation 2 to get control points of a first curve from feature points; then the system may use Equation 4 to control a second curve in the controlling hierarchy, and may control a third curve using sample points of the second curve, so on and so on if needed; then the system may use Equation 3 to animate an avatar by modifying curves in the controlling hierarchy that connect to avatar 3D-meshes. The discriminator may be pre-trained using photos taken on the same type or kind of subjects, before comparing with data from the generator.

[0098] Videos may be captured and stored in the dataset, and may be continuously in real-time during the comparison. The discriminator may compare the videos (denoted as X) with rendered 2D- images from an animated 3D avatar (denoted as X ); while X and X may be compared in a synchronized way (e.g. at the same time intervals from start of animation). The video can be captured from the same subject that the 3D-avatar has been created from (e.g. the same humansubject scanned). For example, a camera can continuously capture a human-subject genuinely smiling, then the system can label the whole video as “human genuinely smiling”. Then the system asks the generator to animate a 3D-avatar smiling genuinely, and compare the two (X and X ) at their synchronised time instants (e.g. both at G and t2from starting to smile genuinely).

[0099] The generator may transform animated avatar from 3D-space to 2D before being fed to the discriminator by rendering the 3D-avatar with texture and lighting data. More details of this aspect are at least covered by references in the “Photogrammetry" section in References. A 3D- to-2D converter may be between the two networks. The 2D-plane for projecting the 3D-avatar onto can be set either manually, or automatically by analyzing a capture-angle (e.g. relative to a precise front-photo) of a 2D-photo. See “Angle detection" section in References for solutions in this aspect. The video dataset may send X to an angle detector that detects the angle to be used in the 3D-to-2D converter. The 3D-to-2D converter can render the 3D-avatar to a 2D-plane that matches this angle. The discriminator may calculate the distance or loss between Corresponding points in 2D-space, and take the total loss as a sum of all losses between the Corresponding points, and back-propagated to update the two networks.

[0100] Training data may be organized on contextual signals. To synchronise X and X , a pairing controller can send pairing information to both the Generator and the video dataset so the discriminator can get comparable data, and the pairing information can include: an identify of a 3D-subject (e.g. an ID in a relational database); thus the video dataset can retrieve the subject, and the Generator may get an-avatar that has been customised on the same subject, and which contextual signal that the system is currently training, thus the video dataset and the Generator can retrieve data for the signal, and what time instance of the current signal that the system is comparing, thus the video dataset and the Generator can retrieve data at the time instance, i.e. X(t) and X (t).

[0101] In some example embodiments, the discriminator and the generator may be trained sequentially on the time dimension, e.g. fully training Controlling-information at G (e.g. 1- second) from the start of animation before moving to train Controlling-information at G (e.g. 8- seconds). The optimal density and selection of the time instances may be application, device and effect-dependent, and may be adaptively determined during the above process, and Controllinginformation of in-between time instances may be interpolated.

[0102] Both the Generator and Discriminator can be trained using the Discriminator’s losses, which may include classification loss and other losses detailed in this document.

[0103] The total loss may be represented via this equation i=0 i=0 where Tte X , GtX ' ; in other words, Ttis from the training dataset, while Gfis from the generator, and they are Corresponding points; and i denotes either a single vertex ( x, y, z) in 3D- space or a single pixel ( x, y) in 2D-space, if the element contributes to the loss.

[0104] Feed-back to the Generator may also include ordered losses with their Corresponding locations, e.g. via this Equation so the Generator can trace back to the control parameters that contribute to the losses, and their degrees of the contributions. The tracing info is detailed previously in this document. Thus the Generator may focus on adjusting the control parameters, and to an amount that may offset their contributions to the losses. The benefits for feeding-back these details and making use of the tracing may include at least easier to reach equilibrium in GAN training, smaller and more portable Al-models, and lower training costs.

[0105] The system may also be used in training Controlling-information that controls the transitions among different elementary factors or processes, and different intensities of one factor, and their blending thereof. For example, a camera can continuously capture the face of a human-subject’s transition from smiling to surprise, and the system can label the video accordingly. Then the system asks the generator to animate a 3D-avatar’s transition from smiling to surprise, and compare the two (X and X ) at their synchronised time instants, thus the Controlling-information that controls the transition can be trained.

[0106] To update the generator, Controlling-information of animatable avatars can be adjusted. For example, refer to Equation 3, to adjust the movement of a vertex V , sample point Snor coefficient parameter rn, or their both may be changed, e.g. by using a different neighbouring sample point Sm, a different value of parameter rm.

[0107] In some example embodiments, the Controlling-information of animatable avatars can be adjusted by spawning additional controlling curves, or changing an area a curve can control, e.g. a new curve might influence a smaller area of the mesh surface. For example, under some situations such as modelling strong emotions, the controlling hierarchy (of FIG. 5B) may spawn additional curves and one curve may control a surface size (of a 3D-mesh) that is different from before the spawning, or using a different set of parameters. Newly spawned curves may replace an existing curve, or may extend into additional layers in the control-hierarchy (e.g. extending towards right side of FIG. 5B), and might be controlled by curves exist before the spawning. The coefficients and vectors in Equation 4 may be adjusted accordingly. Further, updating the Controlling-information may be aided by qualitative or quantitative info about biological or anatomical processes. For example, at a macro-scale, a curve may be simulating a group of muscles; at a micro-scale, a curve may be simulating one or more muscle fibers. Thus triggering spawning additional curves, and the rations of curves-to-elements may be application, device, and effect-dependent. During consuming Al-model (as 106 in FIG. 1), Controlling-information of in-between triggering spawning curves may be interpolated.

[0108] The training will be iterative until the Generator improves its controlling information so as to make the Discriminator can’t tell the differences between animated avatar and 3D subjects that are needed for a specific application, device and effect. Controlling-information of animatable avatars (e.g. those in Equation 3 and Equation 4) may be represented as functions of contextual signals. The control points of the curve(s) situated at the top of the controlling hierarchy (e.g. as closer to left side in FIG 5B) may have an additional free parameter of time, since they control other curves that then animate the avatars. The functions may be parametric, semi-parametric or non-parametric, depending on trained results.

[0109] For the generator in this document, latent space (an abstract multi-dimensional space that is widely used in GAN literature) may be used to represent and encode 3D-geometry, and their movements, instead of 2D-pixels as widely used in prior art.

[0110] It is noted that training, modelling and deploying 3D and 4D (i.e. 3D + time variations detailed in this document) information makes use of the real-world causality, and may be a much more efficient way in generic AI / ML modelling. Since a same set of 3D static and animtable information can be projected as a large number of images on many different 2D-planes, thus 3D static and animtable information described in this document may be used for generic AI / ML modelling, i.e. alternative to conventional Al-modelling that heavily relies on 2D images.

[0111] In some example embodiments, higher fidelity may be needed (e.g. a celebrity or high-paying customers), the system may create a denser set of the Controlling-information. For example, time intervals (e.g. between the above and t2) can be made smaller; controlling-information (e.g. in Equation 3 and 4) may depend on additional factors (such as intensity of expressions or movements); elementary signals and transitions among them may be created on finer scales; the ratio of controlling curves to a single real-world element may be increased; or a combination thereof.

[0112] FIG. 7 illustrates a type-hierarchy that can be created and consumed.

[0113] The type-hierarchy has a number of layers, e.g. a Super-Type, a Type and a Sub-Type layer, and each layer has one or more units. Each unit in the type-hierarchy may have type-wide information of both static and animatable avatars.

[0114] As illustrated, a Super-Type may be situated at the top of the hierarchy. The Super-Type may be an abstract type for all animatable avatars.

[0115] The next layer below the Super-Type is denoted as a Type layer. For example, this layer may have one type for all humans, one type for all species of animals, and another type for imaginary subjects, etc. Each of the types can be represented as a unit in this layer.

[0116] The next layer below the Type-layer is denoted as a Sub-Type layer. For example, the type for all humans may have subsequent sub-types representing different races; the type for imaginary or fancy subjects may have sub-types for 3D models of plants, and furniture respectively, which may be animated as if they were human-like, or animal-like.

[0117] In some example embodiments, the type-hierarchy may be represented in a suitable data structure, such as a searchable tree. Each unit in the type-hierarchy may have metadata to describe its content (e.g. a unit for “Caucasian human-subjects”), its position and relationship with other units in the tree.

[0118] FIG. 8 is a workflow of a method for an individual 3D-subject being fed through the learning phase during Al-modelling, according to some example embodiments.

[0119] For ease of understanding, a customised static avatar (e.g. one static avatar representing or resembling a particular individual 3D-subject) is denoted as CSA; and a customised animatable avatar (e.g. a CSA with the added ability for animating) is denoted as CAA; and trained ML models (e.g. that may animate different subjects of the trained type) are denoted as TMM. Again, CSA, CAA and TMM may organize on contextual signals that are application, device and effectdependent.

[0120] The method first collects data of an individual 3D-subject. The method then creates a static avatar from the data. The static avatar can be represented as a 3D-mesh that include geometric and texture details that may realistically resemble the 3D-subject. Afterwards the static avatar can be considered a CSA. Next, the method flows to labelling types that the 3D-subject belongs to, e.g. what sub-type, type, and super-type the avatar shall be classified into. The labeling process may be manually, semi-automatically, or fully-automatically, e.g. by using additional computer-vision software that can detect the types that the 3D-subject belongs to.

[0121] Next, the method flows to creating CAA for this individual 3D-subject, by training the subject on different contextual signals that may be application, device and effect-dependent.

[0122] In creating CAA, the method flows to getting feature-points for the CSA. The method can then get the Measurements of static-avatars of the CSA by using normalised local coordinate system detailed previously.

[0123] The method may create a controlling hierarchy on the CSA. The controlling hierarchy is detailed in FIG.5A and 5B. After this step, the elements of the controlling hierarchy, such as feature points, controlling-curves, are customised for this individual 3D-avatar.

[0124] Next, the method flows to generate intermediate-animation that can be checked against required-quality. FIG. 6 have details to Al-model avatars, and FIG. 10 has details to safeguard the quality of animation.

[0125] Until quality requirements are met, the method shall continuously improve the Al-modelling, e.g. by updating the parameters in Equation 3 and 4, etc.

[0126] After animation of the CAA can meet the desired quality -requirements, the method can save into a database the result of training the avatar that may include at least Measurements of the CSA, Controlling-information of the CAA, together with its type information.

[0127] FIG. 9A is a flowchart of a method for building up the type-hierarchy.

[0128] First, the method may collect a large amount of data representing different 3D-subjects. The data shall include at least geometric and texture data representing these subjects.

[0129] Next, each subject is labelled for types it belongs to, e.g. what sub-type, type, and super-type the avatar shall be classified into.

[0130] Next, the method flows to group collected 3D-subjects according to their types. For example, data scanned from Caucasians can be grouped under a unit representing Caucasian-humans.

[0131] Next, the method flows to generating type-wide information for units in the type-hierarchy. The type-wide information may be created bottom-up, by initially creating information at the lowest layer in the type-hierarchy that the avatars belong to, e.g. starting from a unit representing Caucasian-humans in a sub-type layer. Then the method may use generalisation or inference to create type-wide information for units higher in the hierarchy, e.g. for a unit representing all humans in the Type-layer, and up to the Super-Type layer.

[0132] FIG. 9B is a flowchart of a method for creating type-wide info for a unit in the lowest layer (i.e. 9A010 in FIG. 9A) that the avatars belong to, according to some example embodiments.

[0133] The unit shall have the data representing a number of subjects. Each subject can be fed into the training method detailed in FIG. 8. Afterwards, each subject may have at least Measurements of its CSA and Controlling-information of its CAA.

[0134] Then the method in FIG. 9B collects all Measurements of CSA and Controlling-information of CAA, from all subjects of the unit.

[0135] The method may identify statistics among all these Measurements. For example, the method can check if parametric regression of a normal-distribution fits well.

[0136] If the statistics is converging, then the method flows to processing animatable avatars; otherwise the method flows to an adjuster; details of the adjuster will be given shortly.

[0137] The method may identify statistics (with patterns) among all the Controlling-information, which can include coefficients and vectors in Equation 1 ~ 4. The statistics and its patterns may be checked with regression models. For example, the method can check if parametric regression of a normal -distribution, a linear or a polynormial model fits well.

[0138] Statistics and patterns of the Controlling-information (of CAA) may be checked for dependency on units in the type-hierarchy and their Measurements of CSA. For example, the method may check if Controlling-information for animating happiness mood in different units (e.g. Caucasian and Asians, men and women) may, or may not, show different statistics and patterns. If presented, the dependency can be saved as a part of statistics and patterns of the Controlling-information.

[0139] In case a plateau is reached before convergence, the method in FIG. 9B may use an adjuster to adjust training parameters or settings.

[0140] The adjuster may do one or more of (not mutually exclusive) adjustments:

[0141] • further sub-divide units in type-hierarchy (adjustment- A)

[0142] • change basis of normalization (adjustment-B)

[0143] • request re-training (adjustment-C)

[0144] For example, adjustment-A may further extend into lower-layers in the type-hierarchy by subdividing a unit representing a human race into different age-groups, genders or their possible combinations.

[0145] For example, adjustment-B may change the ways the coefficients and vectors are normalised. Example Measurements of static avatars (e.g. ee / W ) may be generalised as this equation:

[0146] Da,x

[0147] SMB where Da xis the distance between any two feature-points a and x; and SMB means static measurement basis. and some Controlling-information of animatable avatars may be generalised as this equation: where cpt jis the / th control point of a zth curve; cpt }( n , t ) is a function of cp on / / th elementary signal at the time instance of / ; CPtmeans the initial 3D-location of cp, , when the avatar is unanimated or neutral; AMB means animatable measurement basis.

[0148] In some example embodiments of the adjuster, to reach or speed up convergence, adjustment- B may change the measurement-basis of normalisation in Equation 7.1 and 7.2 to one of the following:

[0149] • any one of D , W , H of an individual avatar, where D, H are the depth and the height;

[0150] • a weighted value of D , W , H of an individual avatar, e.g. (W+H) / 2;

[0151] • an average-value of D , W , H of all subjects of a unit in the type-hierarchy that avatars belong;

[0152] • a weighted average-value of D , W , H of all subjects of a unit in the type-hierarchy that avatars belong; so to choose optimal measurement-basis to reach convergence, generalise to abstract types, and to safeguard animation quality which will be detailed shortly.

[0153] The branching of the two re-training requests (i.e. adjustment-A or adjustment-B) may depend on how the optimal set of parameters is being approached, during the training process.

[0154] After Measurements (of CSA) and Controlling-information (of CAA) of the unit converges, the method of FIG. 9B may finalise ML-models of both static and animatable avatars, by saving the Measurements and Controlling-information of the unit on a data file on devices.

[0155] The method in FIG. 9A then flows to creating type-wide information for units one-layer up in the type-hierarchy 9A012.

[0156] The method may use generalisaton to create type-wide information higher in the typehierarchy, and the generalisaton may be situation-dependent or application-dependent.

[0157] For example, the method in FIG. 9A may generalise type-wide information for a unit representing all humans. The method may use the adjuster detailed in FIG. 9B to implement the generalisaton; the adjuster may take as inputs the Measurements and Controlling-information of the units in the Sub-Type layer. The adjuster may generalise Measurements of static avatars and Controlling-information of animatable avatars with adjustment-B and adjustment-C. Thresholds can be set for Measurements and Controlling-information, and checks can be made to determine if they converges in a unit representing all humans, respectively. When convergence is uneconomic to achieve, the method may identify statistics and patterns (of Measurements of static avatars and Controlling-information of animatable avatars) among these units. After the checks and adjustments, a sub-tree representing all human-subjects may have type-wide information for static and animatable avatars, or / and statistics and patterns among the units in the sub -tree.

[0158] Afterwards, a unit higher (e.g. than the layer in FIG. 9B) in the type-hierarchy may finalise Al-models of both static and animatable avatars. The process may continue for units further higher in the type-hierarchy, depending on situations or applications.

[0159] Afterwards, the method in FIG. 9A may flow to finalising Al-models of both static and animatable avatars, bysaving the type-wide information for units in the type-hierarchy.

[0160] FIG. 10 illustrates a system for safeguarding the quality of the animation of avatars, according to some example embodiments.

[0161] The system may use a quality -guarder to set and check requirements, and send re-training requests if the requirements are yet to met. The quality of ML-models, which may be measured as losses detailed previously, shall be made visible to the quality-guarder.

[0162] Since there may be many different ML-models after training, the quality-guarder may set conditions including:

[0163] - approach (as conditions such as budget allow) or maximise realism (Requirement 10A),

[0164] - generate as small as possible, or minimise variations among information of applicable different units in the type-hierarchy (Requirement 10B), to choose an optimal ML-model that satisfies or approaches the conditions.

[0165] Requirement 10A aims to make as small as possible the losses, various machine learning techniques may be used to update the AI / ML modelling, e.g. FIG. 6.

[0166] Requirement 10B aims to make TMM information of a unit varies the least from other upperlayer units that the avatars belong. For example, if a unit (in the type-hierarchy) representing Caucasians has a number of different sets of Measurements and Controlling-information after training, the system may choose the set of information that is the closest to an upper-layer unit that represents all human subjects. The adjuster (detailed in FIG. 9B) may be used to get Requirement 10B satisfied or approached.

[0167] Overall systems and software programs (including parameters of the controlling hierarchy of FIG. 5 A) that satisfies Requirements 10A and 10B may be considered the optimal set of parameters, which was previously referred.

[0168] If training is not approaching the optimal set of parameters, or a plateau has been reached before achieving the optimal set of parameters, the system may take one or more the following:

[0169] - requesting re-training Al-modeling of animatable avatars (request 314),

[0170] - requesting re-training Al-modeling of static avatars (request 214), Request 314

[0171] If losses are decreasing during the training, but not yet meet the threshold requirements, the system may request re-training Al-modeling of animatable avatars, in some example embodiments.

[0172] Request 214

[0173] The optimal set of parameters, in combination with other factors, such as overall training costs, can also be used in determining if the systems may be better re-visiting Al-modelling static avatars. The feature points are placed closer to the top of the controlling hierarchy, thus will affect the overall quality of animation and costs of training. Thus the system may request retraining Al-modeling of static avatars, which may re-adjust whole controlling hierarchy. If the system reaches (or approaches as conditions allow) the optimal set of parameters, the learning process may be considered complete, the system may be referred as fully trained, and may be considered as having achieved TMM (i.e. type-wide information of both static and animatable avatars for the control -hierarchy). The information may be saved in retrievable data structures on computing devices. Then the TMM may be used for production.

[0174] FIG. 11 is a flowchart of a method for consuming trained AI / ML models, via an exemplary workflow of a 3D-subject being converted to an animatable 3D-avatar.

[0175] After the learning phase, the AI / ML models shall be ready for production usage, which may create and animate customised 3D-avatars that may haven’t been used in AI / ML learning process, and with little or non-human intervention.

[0176] First, the method creates a static avatar from mesh-data of a subject. The method may then normalise the data by an individualised local coordinate system. The method then creates a static avatar from the data. The static avatar can be represented as a 3D-mesh that include geometric and texture details that may realistically resemble the 3D-subject.

[0177] The method may locate feature points on the static avatar. If 3D-mesh is obtained from 3D- scan or photogrammetry software, the method may be able to retrieve the feature points autonomously from the mesh-data, since some photogrammetry software offers APIs that offers this functionality. Otherwise, the method may automatically pattern-match the static avatar with TMM to determine the feature points.

[0178] The method next calculates all measurements of the CSA that the TMM possesses, e.g. eelW .

[0179] Next, the method analyses the avatar (1108) against the AI / ML models (or TMM) that has been built. With these measurements and their relationship, the method next search the TMM information in the type-hierarchy, e.g. to identify what sub-type, type and super-type that the CSA may fit into. The search may be conducted top-down or bottom-up, e.g. on a type-hierarchy as illustrated in FIG. 7. If measurements of the CSA fit type-wide TMM info of units in the typehierarchy, then the method may successfully identify the sub-type, type and super-type for the CSA.

[0180] After identifying the CSA’s location in the type-hierarchy, the method next retrieve Controlling-information of CAA that the TMM possesses, to obtain type-wide TMM info for the avatar. The method may fine-tune the type-wide Controlling-information, e.g. if the Controllinginformation possesses dependency on factors such as Measurements of CSA. Then, the method creates customised controlling hierarchies. As will be appreciated, elements of the controlling hierarchy (as in FIG. 5 A and FIG. 5B), such as feature points, controlling-curves, shall be individually customised for this 3D-subject by the method. Detailed parameters that can locate a subject within statistics can be calculated and saved in an avatar database. Thus the customised avatar may be considered an animatable avatar (1116) that can react on contextual signals.

[0181] Depend on the applications, the animatable avatar may be able to, when being given signals about expressions, moods, speech or descriptions, text-messages, create 3D-animation that may approach the realism of a real-world subject.

[0182] In some example embodiments, the animatable avatar may interact with users autonomously, and being displayed on user interfaces running on various devices.

[0183] FIG. 12 illustrates a system for creating visual contents (e.g. including video contents) from descriptive information such as texts, according to some example embodiments.

[0184] The system may have an input handler component that is able to accept descriptive information from multiple sources, including texts, chat contents, speeches and other visual or audio inputs.

[0185] The system may have a Natural Language Processing (NLP) component that can breaks down textual contents and extract key elements that are needed for visual representation. The key elements may include story-characters (e.g. humans or other creatures), actions and poses (e.g. characters’ moods, facial expressions, speech contents, body movements, transitions between movements), and settings (e.g. natural or man-made environments, lighting conditions). Relative to animation that is continuous in time, a pose means the status of an avatar at an instance of time.

[0186] The system may send the elements to TMM. TMM may use the elements to create animatable avatars.

[0187] The system may use the animatable avatars to create 3D-animation in an animation generator or a pose generator that may include the fully trained Generator network detailed previously. The 3D-animation or pose may temporarily be saved or buffered in-memory on computing devices.

[0188] The 3D-animation or pose is then sent to a Tenderer that applies a rendering pipeline to create videos or images from the 3D-animation.

[0189] FIG. 13 is a workflow of a method of converting descriptive information (e.g. texts) to visual contents using the system in FIG. 12.

[0190] The method first uses the input handler component to input texts or info that can be converted to texts. The original texts may be from various sources. Non-text signals can be pre-processed, e.g. voice inputs can be converted to textual input via Text-to- Speech (TTS) services, and images or handwriting can be converted to textual input via image analysis services.

[0191] Next the method parses and analyses the text inputs. The analysis aims to understand contexts of the original texts, and can use the NLP component. As will be appreciated, NLP can extract information from unstructured text using text mining techniques such as deep learning. In some example embodiments, the method may use the Transformer models in the NLP component.

[0192] After the analysis, the system may extract key elements for visual contents, including information about characters in the story, their actions and poses, and environments settings.

[0193] The method may use extracted information about story-characters to create animatable avatars. An example of extracted character info may be “a Caucasian man in his 30s”. The method may aim to match this description with metadata of a unit in the type-hierarchy (e.g. of FIG. 7). As stated in this document, metadata of the type-hierarchy shall include description about type contents. The rest for creating an animatable avatar that matches the character-info shall be similar to parts of FIG. 11 (e.g. including 1108 and 1116).

[0194] The method may use extracted information about actions to animate the avatars. An example of extracted action info may include “he smiles warmly and greets his visitor”. Smiling and mouth-shapes for making various speech are examples of contextual signals. The fully trained TMM shall have the Controlling-information that is able to animate avatars to match the contextual signals, from the methods, systems, computer algorithms detailed elsewhere in this document.

[0195] The method may use extracted information about actions to animate transitions among different signals (e.g. moods, mouth shapes, etc). For example, extracted info about actions may include “he then turns to a stern face after the greeting”. The TMM of animatable avatar shall have the Controlling-information for transitions, e.g. from warm-smiling to being stern. Further, the Controlling-information of transitions may be complemented by interpolating the CSA and CAA of different contextual signals, fidelity, etc.

[0196] As will be appreciated, the 3D-animation may be fine-tuned depending on the details of extracted info. If applicable, Controlling-information among different sets of spawned curves may be interpolated, e.g. to control continuous animation from a weak emotion to a stronger emotion. Interpolation may be applied to the whole or a part of the controlling hierarchy. Fine- tuning and spawning curves are detailed elsewhere in this document.

[0197] With the TMM information, the method may use the animation generator component to generate 3D-animation, or poses.

[0198] The method may use extracted information about settings to render the animated avatars. Examples of extracted setting info may be “they are in a crowd meeting room” or “they are on an open ground shined by morning sun”. The method may create other avatars that match their respective descriptions, and render them in a suitable environment accordingly. As will be appreciated, 3D rendering is the act of creating 2D contents from animated 3D scenes, and a render may use techniques such as ray-tracing and rasterization. Thus the method may generate 3D-animation that matches descriptive information (e.g. texts).

[0199] FIG. 14 is an architecture of a system for verifying authenticity of visual contents (e.g. detecting deepfake), according to some example embodiments.

[0200] As shown in FIG. 14, a client may be an individual or an organization who consumes internet contents. A content provider has visual contents that become accessible to the client. The client may use a content verifier’s services to verify the authenticity of the visual contents.

[0201] A content provider can supply contents and a content verifier can offer verification services via their respective server computers, and their server computers can communicate directly via web protocols (e.g. HTTPS), when serving common clients. Thus, after a client subscribing to a content verifier’s services, the client may not need to copy URLs hosting visual contents and send the URLs to a content verifier. The verification may be a background service to clients that does not interrupt their routine work. The visuals and their accompanying info (e.g. titles, texts, audios and links) can be generalised as contents-under-verification and is denoted as CUV in this document. A content verifier may provide online services to a larger number of clients who seek to verify CUV that are supplied by content providers.

[0202] FIG. 15 illustrates the modules of the content verifier in FIG. 14.

[0203] The content verifier may include several modules which may be implemented in hardware, software (e.g. programs), or integrated with the operating system of devices, or a combination thereof, and may be distributed on applications running on client devices, server computers, or a combination thereof. The modules include an avatar database module, a TMM module, a content analyser module, an element varier module, an avatar creator module, an animation generator module, a pose generator module, a Tenderer module, a comparer module, a coordinator module, and may include a challenger module.

[0204] The TMM module has the trained AI / ML models of avatars. The content analyser module analyses CUV, and the element varier module adjusts elements in consistence with the analysis results. The avatar creator module creates avatars on information that are made available to the content verifier. The animation generator module generates animation for the avatars. The pose generator module generates poses for the avatars. The Tenderer module renders avatars from either the animation generator or the pose generator. The comparer module compares rendered avatars to CUV. Example embodiments for creating avatars, generating animation and poses, rendering animated avatars, and comparing avatars have been given previously. The coordinator module may branch possible workflow, depending on results from the comparer module and trends of the comparison. The challenger module can create challenges and send the challenges to content providers depending on comparison results. The challenger module may be an optional module of the content verifier.

[0205] FIG. 16 is a workflow of a method for adaptively verifying authenticity of visual contents (e.g. detecting deepfake), according to some example embodiments.

[0206] Operation 1602 synthesises contents according to info from CUV, and compares synthesised contents with CUV. Operation 1602 further includes operations of 1604, 1606, 1608, 16010, 16012, 16014.

[0207] Operation 1604 is for the content analyser to analyse CUV. The content analyser can analyse information from texts, audios, images and videos that are available or searchable from CUV.

[0208] Operation 1606 is for the content analyser to extract useful information that shall include key elements, including information about real or fancy characters, their actions and poses, and environments settings.

[0209] Operation 1608 is for the avatar creator to create or retrieve avatars on extracted key elements. The extracted key elements shall include character info. Operation 1608 can check if an avatar representing the character already exists in the avatar database. If a check returns positive, operation 1608 can retrieve the avatar from the avatar database; otherwise, operation 1608 may create a 3D-avatar using available and searchable visuals, and the search can use extracted key elements (e.g. photos and videos about the character). The 3D-avatar may be created using photogrammetry on-the-fly by one or more processors. Created 3D-avatars in operation 1608 can be saved into the avatar database module.

[0210] Operation 16010 checks whether CUV has images or videos, or both. By applying extracted information on TMM data, the animation generator can create animation to compare to video contents of CUV; the pose generator can create poses to compare to image contents of the CUV.

[0211] Operation 16012 is for the Tenderer to render the outputs from the animation generator or the pose generator, and save to computer memory or on files as synthesised contents. Operation 16014 is for the comparer to compare the synthesised contents to CUV. Example embodiments for creating and comparing visuals are given previously.

[0212] Operation 16016 is for the coordinator module to check whether to challenge the content provider, or to vary elements, or to proceed to operation 16030. The comparer and the coordinator modules may make use of losses detailed previously (e.g. Equations 5 ~ 6). In determining the branching, the coordinator module may save and analyse trends of the losses during continuous comparison. Possible trends may include increasing, fluctuating, or decreasing in losses, and response times of the content providers. For example, if losses are consistently large, the method may flow to operation 16018 that is for the element varier to vary elements in a manner consistent with extracted key elements. For example, if extracted key elements include general greeting, the element varier may change detailed contents of greeting or a different expression that shall be consistent with extracted action-elements. The coordinator may branch the flow depending on the trends of the comparisons, the importance, severity or nature of CUV, the credibility of the content provider, and the likelihood that verification algorithms are known to the content provider. For example, fluctuation in overall losses might lead to generating challenges 16020.

[0213] In operation 16020, the challenger may adaptively challenge the content provider. For example, the challenger may ask a content provider to supply visuals taken at different angles, or different points of time; or the challenger may ask the content provider to supply video contents if only images are initially presented in CUV.

[0214] The challenges are sent to the content provider via web protocols. After receiving challenges, a content provider may get contents meeting the challenges and also related to the initial CUV, e.g. images taken at different angles, or video contents from which images were taken out as initial CUV. The content provider can send a URL hosting the related contents to the content verifier.

[0215] Operation 16022 is for a content verifier to take a content provider’ response to the challenges to analyse further. There are many ways for a content verifier and a content provider to match the initial CUV, its challenges and responses, e.g. via an encrypted Id in URL strings.

[0216] After getting the responses, operation 16022 sends the related contents of the initial CUV to operation 1602, to have another round of analysis. For example, if the response includes photos taken at different angles, operation 1602 may analyse capture-angles; if the response includes video streams, operation 1602 may compare the streams with synthesised animation (e.g. as detailed previously in FIG. 6). Operation 1602 may compare the related contents in consideration of trends of previous comparisons that have been saved in the coordinator module.

[0217] The above process may iterate until the coordinator decides to proceed to operation 16030.

[0218] Operation 16030 outputs the result of comparison, e.g. whether the CUV are genuine or faked. The result may be sent to clients who use a content verifier’s services, via web protocols. The result of the verification may be displayed or logged on clients’ devices. Content verifier’s services may block CUV if the CUV are found faked. It is noted that a content verifier and a content provider may form challenge - response loops. For example, such loops may include continuous challenges and responses illustrated in FIG. 16. The content verifier may repeatedly challenge the content provider, and the content provider may response each of the challenge. Since the content verifier may create new challenges in a dynamic and adaptive fashion. If the verification algorithms and their dependent details may be known to the content provider, the content verifier may ask the content provider to supply related visuals taken at specific angles, or specific points of time; or the content verifier may ask the content provider to supply video contents taken by a specific camera that might be presented in an important venue.

[0219] As will be appreciated, a content verifier may have resources that are unavailable to the a content provider, such as a denser set of the Controlling-information (covered previously), more training data or pre-proccessed data, secrets (such as camera-configurations in venues), server machines with more computing power, or a combination thereof. Thus the content verifier may dynamically synthesise related contents using details disclosed in this document and challenge a content provider to do the same. For example, the content verifier may be synthesising related contents by setting a virtual camera to the same location of a real-camera, of which the content verifier may have detailed knowledge. If the CUV are fake, the content provider may fail the dynamical and adaptive challenges even if the verification algorithms and their dependent details are known to the content provider.

[0220] Thus, with resources including but not limited to, info and data, secrets and computing power, the verification may not need to rely on obscurity of algorithms.

[0221] Aspects of this document may be implemented using hardware, software, firmware or a combination thereof, and may be implemented in one or more computer systems or other processing systems. In one variation, aspects of this document are directed toward one or more computer systems capable of carrying out the functionality described herein, machine-readable medium on which is stored one or more sets of data structures or instructions (e.g. software) embodying or utilized by any one or more of the methods, systems, algorithms, techniques or functions described in this document.

[0222] Readers are advised to refer to priority documents (i.e. provisional patents) for more definitions (including the scopes of devices) and examples, references and additional embodiments.

[0223] REFERENCES

[0224] Patents:

[0225] Huang, "Parameterization of deformation and simulation of interaction", US patent 8,988,419 (application no. 10 / 416459).

[0226] Huang, "Method for customizing avatars and heightening online safety", US patent 8,555,164 (application no. 10 / 862396).

[0227] Photogrammetry https: / / techcrunch.com / 2008 / 02 / 27 / make3d-turn-a-2d-picture-into-a-3d-model / https: / / www.saashub.com / face-gen-alternatives

[0228] Blanz, et al, "A Morphable Model For The Synthesis Of 3D Faces", https: / / www.face-rec.org / algorithms / 3d_morph / morphmod2.pdf Angle detection

[0229] Kashevnik, Alexey M., Ammar Ali, I. B. Lashkov and Dmitry Zubok. “Human Head Angle Detection Based on Image Analysis.” (2020). https: / / www.semanticscholar.org / paper / Human- Head-Angle-Detection-Based-on-Image-Analysis-Kashevnik-Ali / dfe32a688b23ff dl l47fbcd898f aaclc5b7291

Claims

CLAIMS: I CLAIM:

1. Novel methods for creating hierarchical information about static and animatable avatars and consuming the hierarchical information in generic AI / ML modelling and various applications.

2. Novel systems for creating hierarchical information about static and animatable avatars and consuming the hierarchical information in generic AI / ML modelling and various applications.

3. A machine-readable storage medium including instructions that, when executed by a machine, cause the machine to perform operations including creating hierarchical information about static and animatable avatars, and consuming the hierarchical information in generic AI / ML modelling and various applications.

Citation Information

Patent Citations

  • Method and apparatus for creating 3D face model by using multi-view image information

    US20090153553A1

  • Semantic Rigging of Avatars

    US20120139899A1

  • Generating Machine-Learned Inverse Rig Models

    US20230394734A1

  • System and method for animating secondary features

    WO2023004507A1