Determination of re-aged facial-image based on generative diffusion model

The generative diffusion model in the electronic device efficiently predicts re-aging delta images, addressing the challenges of facial aging in digital media by enhancing latent features and reducing computational complexity, resulting in realistic and precise re-aging effects.

WO2025181772A1PCT designated stage Publication Date: 2025-09-04SONY GROUP CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2025/052229
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-01
Filing Date
2025-02-28
Publication Date
2025-09-04

AI Technical Summary

Technical Problem

Existing image processing technologies struggle to achieve flawless facial aging and re-aging effects in digital images, particularly in high-resolution and high-frame-rate media, while maintaining the likeness and subtleties of the face, due to computational intensity and technical complexity.

Method used

An electronic device employs a generative diffusion model to predict re-aging delta images based on a textual prompt, using a neural network architecture that includes an image encoder, text encoder, and a linear layer to enhance latent features, allowing precise manipulation of facial attributes and minimizing computational overhead.

Benefits of technology

The method provides efficient and realistic re-aging results with fine control over facial features, preserving high-frequency details and maintaining the subject's identity, reducing the labor and errors associated with manual sculpting and animation techniques.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2025052229_04092025_PF_FP_ABST
    Figure IB2025052229_04092025_PF_FP_ABST
Patent Text Reader

Abstract

Provided is an electronic device for determination of a re-aged facial image based on a generative diffusion model. The electronic device receives a source image including a facial region of a user. Further, the electronic device receives a textual prompt indicative of a set of facial attributes of a target image of the facial region of the user. The electronic device applies a generative diffusion model on the received source image, based on the received textual prompt. The electronic device predicts a set of re-aging delta images based on the application of the generative diffusion model. The electronic device determines the target image of the facial region of the user, based on the predicted set of re-aging delta images and the received source image. The determined target image corresponds to the set of facial attributes of the target image.
Need to check novelty before this filing date? Find Prior Art

Description

DETERMINATION OF RE-AGED FACIAL-IMAGE BASED ON GENERATIVE DIFFUSION MODELCROSS-REFERENCE TO RELATED APPLICATIONS / INCORPORATION BY REFERENCE

[0001] This Application also makes reference to Indian Provisional Patent Application Ser. No. 202411015339, which was filed on March 01 , 2024. The above stated Patent Application is hereby incorporated herein by reference in its entirety.FIELD

[0002] Various embodiments of the disclosure relate to image processing. More specifically, various embodiments of the disclosure relate to an electronic device and a method for determination of re-aged facial-image based on generative diffusion model.BACKGROUND

[0003] In media and entertainment, image processing involves manipulating and enhancing digital images (For example, facial aging and re-aging) using various techniques. The facial aging and re-aging of images refers to digital alteration looks of a user within the images to represent the user at various life phases (for example young, old, and the like). In media platforms (such as, social media platforms), aging effects may be applied to various faces of the user to naturally depict younger or older version of the user. This may include showing wrinkle development, textural changes, and structural alterations in the skin of the user. Conversely, re-aging effects may be used to halt the aging process, allowing the user to portray younger characteristics within the image. Application of visual effects (VFX) to simulate facial aging and re-aging may require a comprehensive method that blends technical expertise with artistic skill. Even with the recent significant advancements in VFX technology, flawless facial aging and re-aging effects remain unattainable. It may be difficult to maintain likeness of the face of the userduring a digital alteration process. It can be challenging to faithfully capture the subtleties of the face and expressions, even with advanced algorithms and technologies. Also, the visual effects may be computationally intensive and technically complex due to the large volume of information conveyed in the images. This problem may be further compounded by high resolution / high frame rate images.

[0004] Limitations and disadvantages of conventional and traditional approaches will become apparent to one of skill in the art, through comparison of described systems with some aspects of the present disclosure, as set forth in the remainder of the present application and with reference to the drawings.SUMMARY

[0005] An electronic device and method for determination of re-aged facial-image based on generative diffusion model is provided substantially as shown in, and / or described in connection with, at least one of the figures, as set forth more completely in the claims.

[0006] These and other features and advantages of the present disclosure may be appreciated from a review of the following detailed description of the present disclosure, along with the accompanying figures in which like reference numerals refer to like parts throughout.BRIEF DESCRIPTION OF THE DRAWINGS

[0007] FIG. 1 is a block diagram that illustrates an exemplary network environment for determination of re-aged facial-image based on generative diffusion model, in accordance with an embodiment of the disclosure.

[0008] FIG. 2 is a block diagram that illustrates an exemplary electronic device for determination of a re-aged facial-image based on a generative diffusion model, in accordance with an embodiment of the disclosure.

[0009] FIG. 3 is a processing pipeline diagram that illustrates determination of the reaged facial-image based on the generative diffusion model, in accordance with an embodiment of the disclosure.

[0010] FIG. 4 is a block diagram that illustrates an architecture of the generative diffusion model for determination of the re-aged facial image, in accordance with an embodiment of the disclosure.

[0011] FIG. 5 is an exemplary sequence diagram that illustrates a scenario of a media asset management (MAM) system including a control flow of determination of a re-aged facial image based on a generative diffusion model, in accordance with an embodiment of the disclosure.

[0012] FIG. 6A is a diagram that illustrates an exemplary generator architecture of a generative diffusion model for determination of a re-aged facial image, in accordance with an embodiment of the disclosure.

[0013] FIG. 6B is a diagram that illustrates an exemplary discriminator architecture of a generative diffusion model for determination of a re-aged facial image, in accordance with an embodiment of the disclosure.

[0014] FIG. 7 is a flowchart that illustrates operations of an exemplary method for merging re-aged images into coherent video sequences, in accordance with an embodiment of the disclosure.

[0015] FIG. 8 is a figure that illustrates an exemplary execution pipeline of video integration based on compilation of target images into video frames upon determination of re-aged facial image, in accordance with an embodiment of the disclosure.

[0016] FIG. 9 is a flowchart that illustrates determination of re-aged facial-image based on generative diffusion model, in accordance with an embodiment of the disclosure.DETAILED DESCRIPTION

[0017] The following described implementations may be found in a disclosed electronic device and a method for determination of re-aged facial-image based on generative diffusion model. Exemplary aspects of the disclosure may provide an electronic device that may receive a source image (for example, an image including a face of a user). The electronic device may also receive a textual prompt indicative of a set of facial attributes (for example, a target ethnicity of the user, a target gender of the user, and a description of target facial features of the user) of a target image of the user. The set of facial attributes may include, for example, the target age of the user in the target image. The electronic device may apply a generative diffusion model on the received source image, based on the received textual prompt. Further, the electronic device may predict a set of re-aging delta images based on the application of the generative diffusion model and determine the target image of the facial region of the user, based on the predicted set of re-aging delta images and the received source image. The determined target image may correspond to the set of facial attributes of the target image.

[0018] Typically, an appearance of a facial region of a user may be manually manipulated using digital effects to create aging and re-aging effects. An alternative approach to creation of the re-aged face may involve manually sculpting the face, which can be captured using performance capture or keyframe animation techniques. This reaged face may then be rendered from any desired viewpoint. For example, in case a character’s face needs to be re-aged, an artist may meticulously sculpt the facial features, adjust the wrinkles, skin texture, and overall appearance to reflect the desired age. This sculpted face can then be captured using performance capture techniques, where the character’s facial movements may be tracked and mapped onto the digital model. Alternatively, keyframe animation techniques may be employed, where the artist may manually animate the facial expressions and movements of the character’s face frame-by-frame. The character’s face can be seen from different angles, allowing an immersive and realistic portrayal of the aging process of the character’s face. However, manually re-aging faces in images or video frames can be a laborious and error-prone process that may consume a significant amount of time.

[0019] The disclosed electronic device may receive the source image including the facial region of the user. Further, the electronic device may receive the textual prompt indicative of the set of facial attributes of the target image of the facial region on the user. The set of facial attributes may include, for example, the target age of the user in the target image. The electronic device may apply the generative diffusion model on the received source image, based on the received textual prompt. The set of re-aging delta images may be predicted based on the application of the generative diffusion model. The electronic device may determine the target image of the facial region of the user, based on the predicted set of re-aging delta images and the received source image. The determined target image may correspond to the set of facial attributes of the target image. Thus, the versatility of the electronic device may allow the manipulation of input age maps, providing users with the ability to finely adjust re-aging effects across different facial regions with precision. Based on an integration of a face segmentation network (for example, by use of intermediate images), targeted re-aging may be performed, that may allow subtle shifts in wrinkle positions and the preservation of high-frequency details. Additionally, the utilization of a text prompt may further enhance control over the re-aging process, which may enable even more refined adjustments. The electronic device may adopt a focused approach based on prediction of the re-aging delta images instead of estimation of the current age. This may minimize computational overhead and also enhance efficiency. Based on a prioritization of this streamlined approach and an inherent temporal smoothness across image frames or video frames, the electronic device may ensure consistent and realisticre-aging results. Such realistic re-aging results can be further customized through text inputs, providing unparalleled creative freedom.

[0020] FIG. 1 is a block diagram that illustrates an exemplary network environment for determination of re-aged facial-image based on generative diffusion model, in accordance with an embodiment of the disclosure. With reference to FIG. 1 , there is shown a network environment 100. The network environment 100 includes an electronic device 102, a generative diffusion model 104, a server 106, a communication network 110, and an image-capture system 112. The electronic device 102 may communicate with the server 106 through one or more networks (such as, the communication network 110). The server 106 may be associated with a database 108. The electronic device 102 may include the generative diffusion model 104. The database 108 may include a source image 114, a textual prompt 118, and a target image 116.

[0021] The electronic device 102 may include suitable logic, circuitry, interfaces, and / or code that may be configured to receive the source image 114 including a facial region of a user. The electronic device 102 may be communicatively connected (wired / wireless) to the image-capture system 112. The network environment 100 may operate within a framework of a Media Asset Management (MAM) system. The MAM system may be designed to organize, store, and manage digital media assets such as images, videos, audio files, and documents. The MAM system may provide a centralized repository where these assets may be easily accessed, searched, and shared by users within an organization. The MAM systems may typically offer features such as metadata tagging, version control, rights management, workflow automation, and integration with other software tools. Such MAM systems may be commonly used by media and entertainment companies, marketing teams, and other organizations that deal with large volumes of digital media content. The MAM system may include for example, digital library system,workflow management system, collaboration platform, rights management system, analytics and reporting system, integration platform, and the like. The MAM system may be responsible for managing a wide range of content and supporting various workflows. In an example implementation of the MAM system, when a re-aging request is received through the Hypertext Transfer Protocol (HTTP) Representational State Transfer (REST) Application Programming Interface (API), the re-aging service may use customizable parameters to initiate job creation and thread management processes, which may increase a processing efficiency of the input video content.

[0022] The electronic device 102 may further receive the textual prompt 118 along with the source image 114. The textual prompt 118 may be indicative of a set of facial attributes of the target image 116 of the user. The set of attributes may include a target age of the user in the target image 116. The electronic device 102 may apply the generative diffusion model 104 on the received source image 114, based on the received textual prompt 118. Further, the electronic device 102 may predict a set of re-aging delta images based on the application of the generative diffusion model 104 and determine the target image of the facial region of the user, based on the predicted set of re-aging delta images and the received source image 114. The determined target image 116 may correspond to the set of facial attributes of the target image 116. For example, the electronic device 102 may be associated with a display device (such as, a display device 206A, shown in FIG. 2) that may display the target image 116. The electronic device 102 may be communicatively coupled to the image-capture system 112 to receive the source image 114, which may be captured by the image-capture system 112. In some embodiments, the source image 114 may be received from the database 108, via the server 106. Examples of the electronic device 102 may include, but may not be limited to, a desktop, a tablet computer, a television (TV), a laptop, a computing device, a smartphone, a cellular phone, a mobilephone, a machine learning computing device (enabled with or hosting, for example, a computing resource, a memory resource, and a networking resource), a consumer electronic (CE) device having a display.

[0023] The generative diffusion model 104 may be a neural network architecture. The generative diffusion model 104 may be a machine learning model that can compare and generate inferences based on received data input (for example, source images 114). The generative diffusion model 104 may work based on learning of context and meaning of words from large amounts of unlabeled text data, and then fine-tuning the model for specific tasks using labeled data. The generative diffusion model 104 may process both left and right context of each data input. The generative diffusion model 104 may include an embedding module, a stack of encoders, and an un-embedding module. The generative diffusion model 104 may be pre-trained on two tasks, such as, masked language modeling (MLM) and next-data-item prediction (for example, data items, such as, users and items). The MLM may randomly mask some tokens in the input data and predict the inference based on the context. The generative diffusion model 104 may include a set of nodes that may be associated with the source images 114. The nodes may represent embeddings associated with the source images 114. For example, a node associated with the source image 114 may be representative of an embedding associated with the user. Each node may be connected to a plurality of nodes that may be neighbors of the corresponding node in a graphical representation of the data. For example, the node associated with the source images may be connected to nodes associated with the source images that may be determined as neighbors of the source images.

[0024] In an embodiment, the generative diffusion model 104 may correspond to a neural network. The neural network may be a computational network or a system of artificial neurons, arranged in a plurality of layers, as nodes. The plurality of layers of theneural network may include an input layer, one or more hidden layers, and an output layer. Each layer of the plurality of layers may include one or more nodes (or artificial neurons, represented by circles, for example). Outputs of all nodes in the input layer may be coupled to at least one node of hidden layer(s). Similarly, inputs of each hidden layer may be coupled to outputs of at least one node in other layers of the neural network. Outputs of each hidden layer may be coupled to inputs of at least one node in other layers of the neural network. Node(s) in the final layer may receive inputs from at least one hidden layer to output a result. The number of layers and the number of nodes in each layer may be determined from hyper-parameters of the neural network. Such hyper-parameters may be set before, while training, or after training the neural network on a training dataset.

[0025] Each node of the neural network may correspond to a mathematical function (e.g., a sigmoid function or a rectified linear unit) with a set of parameters, tunable during training of the network. The set of parameters may include, for example, a weight parameter, a regularization parameter, and the like. Each node may use the mathematical function to compute an output based on one or more inputs from nodes in other layer(s) (e.g., previous layer(s)) of the neural network. All or some of the nodes of the neural network may correspond to same or a different same mathematical function.

[0026] In training of the neural network, one or more parameters of each node of the neural network may be updated based on whether an output of the final layer for a given input (from the training dataset) matches a correct result based on a loss function for the neural network. The above process may be repeated for same or a different input until a minima of loss function may be achieved, and a training error may be minimized. Several methods for training are known in art, for example, gradient descent, stochastic gradient descent, batch gradient descent, gradient boost, meta-heuristics, and the like.

[0027] The neural network may include electronic data, which may be implemented as, for example, a software component of an application executable on the electronic device 102. The neural network may rely on libraries, external scripts, or other logic / instructions for execution by a processing device, such as, a processor / circuitry of the electronic device 102. The neural network may include code and routines configured to enable the electronic device 102 to perform one or more operations for application of a generative diffusion model on the received source image, based on the received textual prompt. Additionally, or alternatively, the neural network may be implemented using hardware including a processor, a microprocessor (e.g., to perform or control performance of one or more operations), a field-programmable gate array (FPGA), or an application-specific integrated circuit (ASIC). Alternatively, in some embodiments, the neural network may be implemented using a combination of hardware and software.

[0028] The server 106 may include suitable logic, circuitry, interfaces, and / or code configured to receive the source image 114 (for example, image frames, moving pictures, videos, and the like) including the facial region of the user and the textual prompt 118 indicative of the set of facial attributes. The textual prompt may indicate the set of facial attributes of the target image 116 of the user. The server 106 may be configured to extract the re-aging delta images. Further, the server 106 may be configured to extract the target image 116 of the facial region of the user based on the predicted set of re-aging delta images and the received source image 114. In some embodiments, the server 106 may be configured to store the generative diffusion model 104. In an embodiment, the server 106 may be configured to predict the set of re-aging delta images based on the application of the generative diffusion model 104. The generative diffusion model 104 may correspond to Brownian diffusion model. The generative diffusion model 104 may further include alinear layer configured to enhance each feature of first latent image features and second latent text features.

[0029] The server 106 may be implemented as a cloud server and may execute operations through web applications, cloud applications, HTTP requests, repository operations, file transfer, and the like. Other example implementations of the server 106 may include, but are not limited to, a database server, a file server, a web server, a media server, an application server, a mainframe server, a machine learning server (enabled with or hosting, for example, a computing resource, a memory resource, and a networking resource), or a cloud computing server.

[0030] In at least one embodiment, the server 106 may be implemented as a plurality of distributed cloud-based resources by use of several technologies that are well known to those ordinarily skilled in the art. A person with ordinary skill in the art will understand that the scope of the disclosure may not be limited to the implementation of the server 106 and the electronic device 102, as two separate entities. In certain embodiments, the functionalities of the server 106 can be incorporated in its entirety or at least partially in the electronic device 102 without a departure from the scope of the disclosure. In certain embodiments, the server 106 may host the database 108. Alternatively, the server 106 may be separate from the database 108 and may be communicatively coupled to the database 108.

[0031] The database 108 may include suitable logic, circuitry, interfaces, and / or code configured to store information associated with source images. The database 108 may be associated with the server 106. In an example, the database 108 may include the source image 114 (such as, a set of image frames). The database 108 may further include information related facial attributes, for example, a target ethnicity of the user, a target gender of the user, and a description of target facial features of the user. The database108 may be stored or cached on one or more devices or servers, such as the server 106. The device storing a database may be configured to query the database 108 for certain information (such as, the source image, the textual prompt, etc.) based on reception of the source image 114 (including the facial region of the user) for the particular information from the electronic device 102. In response, the device storing the database 108 may be configured to retrieve, from the database, results (for example, records related to the queried information) based on the received query.

[0032] In some embodiments, the database 108 may be hosted on the servers stored at the same or different locations. The operations of the database may be executed using hardware including a processor, a microprocessor (e.g., to perform or control performance of one or more operations), a field-programmable gate array (FPGA), or an applicationspecific integrated circuit (ASIC). In some other instances, the database 108 may be implemented using software.

[0033] The communication network 110 may include a communication medium through which the electronic device 102 and the server 106 may communicate with each other. The communication network 110 may be a wired or wireless communication network. Examples of the communication network 110 may include, but are not limited to, Internet, a cloud network, Cellular or Wireless Mobile Network (such as, Long-Term Evolution and 5thGeneration (5G) New Radio (NR)), satellite communication system (using, for example, low earth orbit satellites), a Wireless Fidelity (Wi-Fi) network, a Personal Area Network (PAN), a Local Area Network (LAN), or a Metropolitan Area Network (MAN). Various devices in the network environment 100 may be configured to connect to the communication network 110, in accordance with various wired and wireless communication protocols. Examples of such wired and wireless communication protocols may include, but are not limited to, at least one of a Transmission Control Protocol andInternet Protocol (TCP / IP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), File Transfer Protocol (FTP), Zig Bee, EDGE, IEEE 802.11 , light fidelity(Li-Fi), 802.16, IEEE 802.11 s, IEEE 802.11 g, multi-hop communication, wireless access point (AP), device to device communication, cellular communication protocols, and Bluetooth (BT) communication protocols.

[0034] The image-capture system 112 may include suitable logic, circuitry, and interfaces that may be configured to capture an image / a plurality of images of a facial region of a user. The image-capture system 112 may be further configured to transmit the captured image or the captured plurality of images to the database 108 for storage. The image-capture system 112 may also transmit the captured image or the captured plurality of images to the electronic device 102 to generate a re-aged image. Examples of the image-capture system 112 may include, but are not limited to, an image sensor, a wide- angle camera, an action camera, a closed-circuit television (CCTV) camera, a camcorder, a camera with an integrated depth sensor, a cinematic camera, Digital Single-Lens Reflex (DSLR) camera, a Digital Single-Lens Mirrorless (DSLM) camera, a digital camera, camera phones, a time-of-flight camera (ToF camera), a night-vision camera, and / or other image capture devices.

[0035] In operation, the electronic device 102 may be configured to receive the source image 114 including the facial region of the user. The received source image 114 may include the set of image frames or video data received from the MAM system. The MAM system may transmit the source image 114 via Hypertext transfer Protocol (HTTP)-based Representational State Transfer (REST) Application Programming Interface (API). The reception of the source image 114 is described further, for example, in FIG. 3 (at 302).

[0036] The electronic device 102 may be configured to receive the textual prompt 118 indicative of the set of facial attributes of the target image 116 of the facial region of theuser. For example, the electronic device 102 may receive the textual prompt 118 as a user input from the user. An additional textual prompt may be provided to specify the desired de-aging type. Such additional textual prompt may provide auxiliary information to a generator of the generative diffusion model 104, which may incorporate the textual description of the target image 116. An example for the textual prompt 118 may include: “{A / An} {Age}-year-old {Ethnicity} {man / woman / person} with {description of facial features}”Example prompt may include:'A 20-year-old Caucasian man with trimmed beard, dark brown hair, no smile,"

[0037] The set of facial attributes may include the target age of the user in the target image 116. The set of facial attributes of the target image 116 may be associated with at least one of the target ethnicities of the user, the target gender of the user, the description of the target facial features of the user. The reception of the textual prompt is described further, for example, in FIG. 3 (at 304).

[0038] The electronic device 102 may be configured to apply the generative diffusion model 104 on the received source image 114, based on the received textual prompt 118. The generative diffusion model 104 may correspond to the Brownian bridge diffusion model. The received source image 114 may be segmented into a set of intermediate images associated with a set of first image channels (for example, a 5-channel tensor). The application of the generative diffusion model 104 may be further based on the set of intermediate images. The generative diffusion model 104 may include an image encoder configured to determine an encoded image based on the set of first intermediate images. The determined encoded image may correspond to first latent features of the received source image 114. Further, the generative diffusion model 104 may include a text encoder configured to determine a textual embedding based on the received textual prompt 118.The determined textual embedding may correspond to second latent features of the received textual prompt. The generative diffusion model 104 may further include a linear layer configured to enhance each feature of the first latent features and the second latent features. The generative diffusion model 104 may be configured to traverse each feature of the first latent features and the second latent features as at least one of a feedforward feature or feedback feature associated with the generative diffusion model 104. A set of second intermediate images may be determined based on the traversal of each feature of the first latent features and the second latent features. Also, the generative diffusion model 104 may further include a decoder model configured to determine a set of decoded images based on the determined set of second intermediate images. The determined set of decoded images may correspond to the predicted set of re-aging delta images. The application of the generative diffusion model on the received source image is described further, for example, in FIG. 3 (at 306).

[0039] The electronic device 102 may be configured to predict the set of re-aging delta images based on the application of the generative diffusion model 104. The predicted reaging delta images may be merged with the received source image 114. The determination of the target image 116 may be based on the merging of the received source image 114 and the predicted set of re-aging delta images. Based on the prediction of aging deltas rather than an estimation of current age, the generative diffusion model 104 may minimize computational overhead and enhance efficiency. The prediction of the set of re-aging delta images is described further, for example, in FIG 3 (at 308).

[0040] The electronic device 102 may be configured to determine the target image 116 of the facial region of the user, based on the predicted set of re-aging delta images and the received source image 114. The determined target image 116 may correspond to the set of facial attributes of the target image 116. The target image 116 of the facial regionmay include the re-aged image of the user based on a received input image (for example, the source image 114) and a textual prompt (e.g., the textual prompt 118) indicative of a set of facial attributes of a target image (e.g., the target image 116). The determination of the target image is described further, for example, in FIG 3 (at 310).

[0041] FIG. 2 is a block diagram that illustrates an exemplary electronic device for determination of a re-aged facial-image based on a generative diffusion model, in accordance with an embodiment of the disclosure. FIG. 2 is explained in conjunction with elements from FIG. 1. With reference to FIG. 2, there is shown a diagram 200 of the electronic device 102. The electronic device 102 may include a circuitry 202, a memory 204, the generative diffusion model 104, an input / output (I / O) device 206, and a network interface 208. In at least one embodiment, the I / O device 206 may also include a display device 206A. In at least one embodiment, the memory 204 may include the source image 114, the target image 116, and the textual prompt 118. The circuitry 202 may be communicatively coupled to the memory 204, the I / O device 206, the network interface 208, through wired or wireless communication of the electronic device 102.

[0042] The circuitry 202 may include suitable logic, circuitry, and interfaces that may be configured to execute program instructions associated with different operations to be executed by the electronic device 102. The operations may include the reception of the source image 114 and textual prompt 118, the prediction of the set of re-aging delta images, and the determination of the target images 116 of the facial region of the user. The operations may further include the application of the generative diffusion model 104 on the received source image 114. The operations may include merging the received source image 114 and the predicted set of re-aging delta images. The determination of the target image 116 of the facial region of the user may be further based on the merging of the received source image 114 and the predicted set of re-aging delta images. The circuitry202 may include one or more specialized processing units, which may be implemented as an integrated processor or a cluster of processors that perform the functions of the one or more specialized processing units, collectively. The circuitry 202 may be implemented based on a number of processor technologies known in the art. Examples of implementations of the circuitry 202 may be an x86-based processor, a Graphics Processing Unit (GPU), a Reduced Instruction Set Computing (RISC) processor, an Application-Specific Integrated Circuit (ASIC) processor, a Complex Instruction Set Computing (CISC) processor, a microcontroller, a central processing unit (CPU), and / or other computing circuits.

[0043] The memory 204 may include suitable logic, circuitry, interfaces, and / or code that may be configured to store the program instructions to be executed by the circuitry 202. The program instructions stored on the memory 204 may enable the circuitry 202 to execute operations of the circuitry 202 (and / or the electronic device 102). In at least one embodiment, the memory 204 may store the source image 114 including the facial region of the user, textual prompt 118 indicative of the set of facial attributes of the target image 116, the set of re-aging delta images, and the target image 116. The source image 114 may be captured by the image-capture system 112 and transferred to the memory 204. Alternatively, the image-capture system 112 may capture the source image 114 and transfer the captured source image 114 to the database 108 for storage. The source image 114 may be retrieved from the database 108 and transferred to the memory 204. Examples of implementation of the memory 204 may include, but are not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Electrically Erasable Programmable Read- Only Memory (EEPROM), Hard Disk Drive (HDD), a Solid-State Drive (SSD), a CPU cache, and / or a Secure Digital (SD) card.

[0044] The I / O device 206 may include suitable logic, circuitry, interfaces, and / or code that may be configured to receive an input and provide an output based on the received input. For example, the I / O device 206 may receive the source image 114 as a user input from the user. The user input may be indicative of information of the source image 114 (image data or video data including the face of the user) and the textual prompt 118 (indicative of a desired de-aging type, a textual description of the target face, and the like) by the user. The I / O device 206 may render the target image 116 on the display device 206A. Examples of the I / O device 206 may include, but are not limited to, a touch screen, a keyboard, a mouse, a joystick, a microphone, the display device 206A and a speaker. Examples of the I / O device 206 may further include braille I / O devices, such as, braille keyboards and braille readers.

[0045] The I / O device 206 may include the display device 206A. The display device 206A may include suitable logic, circuitry, and interfaces that may be configured to receive inputs from the circuitry 202 to render, on a display screen, the target image 116 (the reaged images or videos of the facial region of the user based on the set of facial attributes). In at least one embodiment, the display screen may be at least one of a resistive touch screen, a capacitive touch screen, or a thermal touch screen. The display device 206A or the display screen may be realized through several known technologies such as, but not limited to, at least one of a Liquid Crystal Display (LCD) display, a Light Emitting Diode (LED) display, a plasma display, or an Organic LED (OLED) display technology, or other display devices.

[0046] The network interface 208 may include suitable logic, circuitry, and interfaces that may be configured to facilitate communication between the circuitry 202 and the server 106, via the communication network 110. The network interface 208 may be implemented by use of various known technologies to support wired or wireless communication of theelectronic device 102 with the communication network 110. The network interface 208 may include, but is not limited to, an antenna, a radio frequency (RF) transceiver, one or more amplifiers, a tuner, one or more oscillators, a digital signal processor, a coder-decoder (CODEC) chipset, a subscriber identity module (SIM) card, or a local buffer circuitry.

[0047] The network interface 208 may be configured to communicate via wireless communication with networks, such as the Internet, an Intranet, or a wireless network, such as a cellular telephone network, a wireless local area network (LAN), a short-range network, and a metropolitan area network (MAN). The wireless communication may use one or more of a plurality of communication standards, protocols and technologies, such as Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), wideband code division multiple access (W-CDMA), Long Term Evolution (LTE), 5thGeneration (5G) New Radio (NR), code division multiple access (CDMA), time division multiple access (TDMA), Bluetooth, Wireless Fidelity (Wi-Fi) (such as IEEE 802.11a, IEEE 802.11 b, IEEE 802.11 g or IEEE 802.11 n), voice over Internet Protocol (VoIP), light fidelity (Li-Fi), Worldwide Interoperability for Microwave Access (Wi-MAX), a near field communication protocol, and a wireless pear-to-pear protocol.

[0048] The functions or operations executed by the electronic device 102, as described in FIG. 1 , may be performed by the circuitry 202. Operations executed by the circuitry 202 are described in detail, for example, in FIGs. 3, 4, 5, 6A, 6B, 7 and 8.

[0049] FIG. 3 is a processing pipeline diagram that illustrates determination of the reaged facial-image based on the generative diffusion model, in accordance with an embodiment of the disclosure. FIG. 3 is explained in conjunction with elements from FIG.1 and FIG. 2. With reference to FIG. 3, there is shown an exemplary execution pipeline 300 for determination of a re-aged facial-image based on the generative diffusion model 104. The execution pipeline 300 may include operations 302 to 310 executed by acomputing device, such as, the electronic device 102 of FIG. 1 or the circuitry 202 of FIG. 2.

[0050] At 302, an operation for source image reception may be executed. The circuitry 202 may be configured to receive the source image 114 including the facial region of the user. The source image 114 may correspond to the set of image frames or video frames and may be received from the MAM system via the HTTP-based REST API. The source image 114 may be received from the image-capture system 112. The source image 114 may include the facial region of the user. For example, an image-1 may include a user-1 , image-2 may include multiple users, and the like. The MAM system may be a comprehensive solution designed to handle, organize, and distribute digital media assets such as videos, images, and audio files. The MAM system may provide, for example, centralized repository, discovery and retrieval of the media files, metadata management, content ingestion, collaboration support, and the like. The MAM systems may provide the centralized location for media files including for example media frames, which may make the media files accessible and manageable across the MAM system. Centralization may ensure the accessibility of media files from the single location, making it easier for the users across the organization to find and manage the media files. The media files may be available on-site or in cloud-based storage (for example, the database 108). The MAM system may ensure the availability of the media file instantly through the API and the user interface. In an embodiment, upon receiving the request for facial re-aging for the user, the re-aging service within the MAM system may employ configurable parameters to instantiate job creation and thread management processes, which may thereby facilitate the processing of the input video content.

[0051] At 304, an operation for textual prompt reception may be executed. The circuitry 202 may be configured to receive the textual prompt 118 indicative of the set of facialattributes of the target image 116 of the facial region of the user. The set of facial attributes may include the target age of the user in the target image 116. Such additional textual prompt may provide auxiliary information to a generator of the generative diffusion model 104, which may incorporate the textual description of the target image 116. An example for the textual prompt 118 may include:“{A / An} {Age}-year-old {Ethnicity} {man / woman / person} with {description of facial features}”Example prompt may include:'A 20-year-old Caucasian man with trimmed beard, dark brown hair, no smile,"

[0052] In an embodiment, the set of facial attributes of the target image 116 may include the target ethnicity of the user, target gender of the user, the description of the facial features of the user. The ethnicity of the user may include, for example, Caucasian, Cajuns, Asian-American and the like. The ethnicity of the user may indicate cultural, historical, and social dimensions of the people. The target gender of the user may include, for example, male, female, and the like. The description of the facial features of the user may be, for example, a man with a beard, wearing spectacles, and the like. The target age may be the desired age of the user to which the user’s appearance may be aligned using the generative diffusion model 104.

[0053] At 306, an operation for generative diffusion model application may be executed. The circuitry 202 may be configured to apply the generative diffusion model 104 on the received source image 114, based on the received textual prompt 118. The generative diffusion model 104 may correspond to the Brownian bridge diffusion model. Upon extraction of a frame from the source image 114, the circuitry 202 may utilize image segmentation techniques to isolate regions of interest (ROIs) corresponding to facial features. Subsequently, the extracted frame may be transformed into a five-channel formatand fed into a re-aging model (e.g., the generative diffusion model 104) for analysis. The image segmentation may be performed on the source image 114 to obtain the first intermediate images. The first intermediate images may include the ROIs corresponding to the facial features. The segmentation of the received source image 114 may be associated with a set of first image channels (for example, a five-channel tensor). The set of first image channels may correspond to a set of second image channels, a third image channel, and a fourth image channel. The set of second image channels and the third image channel may be associated with the segmented source image 114. The fourth image channel may be associated with a segmentation of the determined target image 116. The set of second image channels may correspond to color information (for example, a Red-Blue-Green (RGB) color space) earmarked for re-aging associated with the segmented source image. The third image channel may correspond to a first age map (a face of a 45-year-old person) associated with the received source image 114. The fourth image channel may correspond to the second age map (a face of a 21 -year-old person) associated with the determined target image 116.

[0054] In an embodiment, the generative diffusion model 104 may include the image encoder. The image encoder may determine the encoded images based on the set of first intermediate images. The determined encoded images may correspond to first latent features of the received source image 114. The generative diffusion model 104 may further include a text encoder. The text encoder may determine a textual embedding based on the received textual prompt 118. The textual embedding may correspond to second latent feature of the received textual prompt 118. The generative diffusion model 104 may further include a linear layer configured to enhance each feature of the first latent features and the second latent features.

[0055] At 308, an operation for prediction of a set of re-aging delta images may be executed. The circuitry 202 may be configured to predict a set of re-aging delta images based on the application of the generative diffusion model 104. The generative diffusion model 104 may traverse each feature of the first latent features and the second latent features as a feedforward feature and / or a feedback feature associated with the generative diffusion model 104. The circuitry 202 may determine a set of second intermediate images based on the traversal of each feature of the first latent features and the second latent features. The generative diffusion model 104 may facilitate the generalization of age modification (for example, aging or re-aging) based on the age indicated in the textual prompt 118. The generalized age modification may be the second intermediate images. The facial variation may include, for example, a user with a beard, without a beard, with glasses, without glasses, and so on. The set of re-aging delta images may include a set of images with variations in the facial features. The re-aging delta images may be obtained from a re-aging node. The re-aging node may include a re-aging service bus that operates as HTTP server layered atop the internal REST API (IRA), serving as a conduit between microservices and process management modules. The prediction of the re-aging delta images is described further, for example, in FIG. 4, FIG. 5, and FIG.6A.

[0056] At 310, an operation for target image determination may be executed. The circuitry 202 may determine the target image 116 of the facial region of the user, based on the predicted set of re-aging delta images and the received source image 114. The determined target image 116 may correspond to the set of facial attributes of the target image 116. The generative diffusion model 104 may include a decoder model. The decoder model may determine a set of decoded images based on the determined set of second intermediate images. The determined set of decoded images may correspond to the predicted set of re-aging delta images. The predicted set of re-aging delta images maybe merged with the received source images 114. The determination of the target image 116 of the facial region of the user may be further based on the merging of the received source image 114 and the predicted set of re-aging delta images. Based on the prediction of the re-aging delta images and incorporation of the textual prompt 118, the generative diffusion model 104 may streamline computational processes and ensure natural-looking results. The circuitry 202 may maintain a subject's distinct identity, including unique facial characteristics and expressions in the set of delta images, and thus enhance the overall realism and authenticity of the transformation.

[0057] Typically, an appearance of a facial region of a user may be manually manipulated using digital effects to create aging and re-aging effects. An alternative approach to creation of the re-aged face may involve manually sculpting the face, which can be captured using performance capture or keyframe animation techniques. This reaged face may then be rendered from any desired viewpoint. For example, in case a character’s face needs to be re-aged, an artist may meticulously sculpt the facial features, adjust the wrinkles, skin texture, and overall appearance to reflect the desired age. This sculpted face can then be captured using performance capture techniques, where the character’s facial movements may be tracked and mapped onto the digital model. Alternatively, keyframe animation techniques may be employed, where the artist may manually animate the facial expressions and movements of the character’s face frame-by- frame. The character’s face can be seen from different angles, allowing an immersive and realistic portrayal of the aging process of the character’s face. However, manually re-aging faces in images or video frames can be a laborious and error-prone process that may consume a significant amount of time.

[0058] The disclosed electronic device 102 may receive the source image 114 including the facial region of the user. Further, the electronic device 102 may receive the textualprompt 118 indicative of the set of facial attributes of the target image of the facial region on the user. The set of facial attributes may include, for example, the target age of the user in the target image. The electronic device 102 may apply the generative diffusion model 104 on the received source image 114, based on the received textual prompt 118. The set of re-aging delta images may be predicted based on the application of the generative diffusion model 104. The electronic device 102 may determine the target image 116 of the facial region of the user, based on the predicted set of re-aging delta images and the received source image 114. The determined target image 116 may correspond to the set of facial attributes of the target image. Thus, the versatility of the electronic device 102 may allow the manipulation of input age maps, providing users with the ability to finely adjust re-aging effects across different facial regions with precision. Based on an integration of a face segmentation network (for example, by use of intermediate images), targeted re-aging may be performed, that may allow subtle shifts in wrinkle positions and the preservation of high-frequency details. Additionally, the utilization of a text prompt may further enhance control over the re-aging process, which may enable even more refined adjustments. The electronic device 102 may adopt a focused approach based on the prediction of re-aging delta images instead of estimation of the current age. This may minimize computational overhead and enhance efficiency. Based on a prioritization of this streamlined approach and an inherent temporal smoothness across image frames or video frames, the electronic device 102 may ensure consistent and realistic re-aging results. Such realistic re-aging results can be further customized through text inputs, providing unparalleled creative freedom.

[0059] FIG. 4 is a block diagram illustrating an architecture of the generative diffusion model for determination of the re-aged facial image, in accordance with an embodiment of the disclosure. FIG. 4 is explained in conjunction with elements from FIG. 1 , FIG. 2, andFIG. 3. With reference to FIG. 4, there is shown an exemplary architecture 400 of the generative diffusion model 104 for the determination of the re-aged facial image. The architecture 400 of the generative diffusion model 104 may be implemented on a computing device, such as, the electronic device 102 of FIG. 1 or the circuitry 202 of FIG. 2.

[0060] At 402, an operation of image segmentation may be executed. The circuitry 202 may receive an input image (for example, the source image 114) as a user input from the user, or the database 108. Referring to 404, the source image 114 may include the first image channels (for example, 5-channel source segmented image 404A). The first image channels may include the RGB image, age maps, and the output ages. The second image channel may include the RGB images. A 3-channel source segmented image 404A (for example, the second image channel) may be dedicated to the RGB images. A 1 -channel source segmented image 404B (For example, a third image channel) may be dedicated to the source segmented image. The source segmented image may include the age map of the source image 114. A 1 -channel target segmented image 404C (For example, a fourth image channel) may be dedicated to the target segmented image (for example, output ages). For example, a user may input a current age, for example ’86 years’, and may also input a target age, for example, ’25 years’ for re-aging of the facial region of the source image 114 to the age of ’25 years’.

[0061] In an embodiment, the first image channels may be connected to the generator 406. The generator 406 may include image encoder (EA) 408A, first latent features (LA) 410, second latent features (LT) 414, resulting latent features 412 (LB), a text encoder (ET) 416, a linear layer 418, and an image decoder (DB) 408B. The input for the generator 406 may be the first image channels (for example, the second image channel, the third image channel, and the fourth image channel). Further, the generator 406 may receive the textualprompt 118 as the input. The textual prompt 118 may be an additional text specifying a desired de-aging type. The textual prompts 118 may provide auxiliary information to the generator 406, which may incorporate a textual description of the target image 116. For example, the textual prompt 118 may include:“{A / An} {Age}-year-old {Ethnicity} {man / woman / person} with {description of facial features}”Example prompt may include: "A 20-year-old Caucasian man with trimmed beard, dark brown hair, no smile".

[0062] In an embodiment, the first image channels 404 (for example, the 5-channel image data) may be encoded based on the image encoder (EA) 408A. The first latent features (LA) 410 may be extracted from the image encoder (EA) 408A. In an embodiment, the input text received from the textual prompt 118 may be processed based on the text encoder (ET) 416. The second latent features (LT) (clip-based text embeddings) 414 may be extracted from the text encoder (ET) 416. The image encoder 408A (for example, a vision encoder-decoder model) may be used to initialize the image to any pretrained transformer-based vision model as the encoder (for example, a vision transformer (ViT), a bidirectional encoder representation from image transformers (BEiT), data-efficient image transformers (DeiT), and a Swin transformer) and any pretrained language model as the decoder (for example, Bidirectional Encoder Representations from Transformers (BERT) model and a DistilBERT model). The image encoder 408A may convert the first image channels 404 (for example, the 5-channel image data) of the source image 114 into the first latent features (LA) 410.

[0063] In an example, text-to-image (T2I) pipelines may include the text encoder (ET) 416 and the generative diffusion model 104. The text encoder (ET) 416 may convert the textual prompt 118 into the second latent features (LT) 412. The generative diffusion model104 may then utilize the second latent features (LT) 412 to generate a corresponding image. Various tools may be used to encode and decode text using different encoding methods. For example: a Base91 Encoder may encode data to Base91 , a Base64 Image Encoder may encode image data using a Multipurpose Internet Mail Extensions (MIME) Base64. The image encoding may be referred to as down-sampling. Within the latent space, the first latent features (for example, latent features of the source image 114) and clip-based text embeddings may traverse a Brownian bridge diffusion process 408 in both forward and backward directions. The Brownian bridge diffusion process 408 may facilitate a generalization of age modification, which may encompass both re-aging and de-aging, depending on the age indicated in the textual prompt 118. Examples of the textual prompt 118 may include facial variations like with a beard, without a beard, with glasses, without glasses, and so forth.

[0064] The linear layer 418 may perform a linear transformation on the input data (for example, the first latent features and the second latent features). The combination of the encoded image data and the encoded text data may be referred as the latent features (for example, embeddings). The latent features (for example, the image and the text latent features) may be concatenated or stacked to form a joint latent vector representation. Then, the linear layer 418 may be applied to the joint latent vector. In text-to-image generation models, the joint latent vector may serve as input to the generator 406 (e.g., a GAN or a diffusion model). The generator 406 may then produce an image that aligns with both the textual description and the image latent information. During training, weights of the linear layer 418 may be learned through backpropagation. Fine-tuning the combined latent space may allow the generative diffusion model 104 to capture meaningful interactions between the text and image modalities. In another example, a fusion of the (Contrastive Language-Image Pre-training) CLIP text embedding from the text encoder416 with latent diffused features may be introduced. The integration may facilitate a more control mechanism over the de-aged or re-aged output based on the textual prompt 118. The fusion may be passed through the linear layer 418 for more feature enhancement. The image decoder (DB) 408B may then produce predicted age maps. The predicted aging deltas may be merged with the source image 114 at input merger 422 to make a final deaged image 422A.

[0065] In an embodiment, the training process for the generative diffusion model 104 may involve use of synthetic data pairs and a composite loss function. The loss function may include several components, such as, but not limited to, a Frechet Inception Distance (FID), an L1 Loss, perceptual (LPIPS) Loss, and adversarial losses. The goal of the loss function is to ensure robustness and fidelity in re-aging outcomes. For example, a Textconditional Diffusion-based U-net model may be used to prioritize both identity preservation and spatial fidelity. The synthetic data pairs may consist of artificially generated combinations of text and corresponding images, which may be used during training to enhance the model’s performance. The FID may measure the similarity between the generated images and the source images 114 based on features extracted from a pretrained neural network (for example, the generative diffusion model 104). The L1 loss may measure pixel-wise difference between the generated images (for example, the output from the generator 406) and ground truth images. The perceptual loss may be a perceptual similarity metric that may work on high-level features (such as, texture and structure) rather than the pixel values. The combination of such loss terms may ensure that the generative diffusion model 104 optimizes for both perceptual quality and pixel fidelity. Based on the use of the synthetic data and the composite loss function, the generative diffusion model 104 may achieve robustness (i.e., an ability to handle diverse inputs) and fidelity (i.e., a faithfulness to the original image) in its re-aging predictions.

[0066] It should be noted that the architecture 400 of FIG. 4 is for exemplary purposes and should not be construed to limit the scope of the disclosure.

[0067] FIG. 5 is an exemplary sequence diagram that illustrates a scenario of a media asset management (MAM) system including a control flow of determination of a re-aged facial image based on a generative diffusion model, in accordance with an embodiment of the disclosure. FIG. 5 is explained in conjunction with FIG. 1 , FIG. 2, FIG. 3, and FIG. 4. With reference to FIG. 5, an exemplary scenario 500 of a MAM system including a control flow of determination of a re-aged facial image based on the generative diffusion model 104 is described. The scenario 500 includes a MAM client 502, a secure Application Programming Interface (API) gateway 504, an internal Representational State Transfer (REST) API 506, a re-aging node 508, and a video generation model 510. The video generation model 510 may be associated with various operations, such as, an operation for video frame extraction and processing 510A, an operation for face detection and segmentation 510B, an operation for ruse of a re-aging model 510C, and an operation for re-aging frame merger to video 510D.

[0068] The MAM system may include a process management module (not shown in FIG. 5). The process management module may serve as a central orchestration to handle tasks within the system, including workflow initiation from a MAM client 502. The MAM client 502 may access credentials, via the secure API gateway 504 and may facilitate communication with the internal REST API 506 for seamless workflow execution. Additionally, the MAM client 502 may ensure efficient resource allocation, error handling, and logging for smooth system operation.

[0069] The MAM client 502 may be a user interface through which users may interact with the MAM system. The MAM client 502 may initiate workflows like re-aging via a web interface, based on a user’s choice of parameters like file ID and target age. In the re-aging workflow, MAM client 502 may send media files and parameters to the secure API gateway 504 with access credentials in a request header.

[0070] The secure API gateway 504 may act as protective layer for the MAM system’s APIs and may manage authentication, authorization, and routing of requests to the appropriate services. In an example, upon request reception, credentials may be checked for re-aging service access within the secure API gateway 504. Upon verification, command parsing initiates the execution flow.

[0071] The internal REST API 506 within the system seamlessly may continue the flow from the secured API gateway. The internal REST API 506 may handle further processing of requests, authentication, and routing to various services within the MAM system, to ensure secure and efficient communication.

[0072] The re-aging node 508 may include a re-aging service bus that operates as a lightweight HyperText Transfer Protocol (HTTP) server layered atop the Internal REST API (IRA), that may serve as a conduit between microservices and process management modules. Upon job creation, an operational thread initializes to facilitate download of video files from cloud storage and trigger the re-aging system. Initial steps may involve frame extraction from videos, followed by face recognition and segmentation. The segmented face images undergo processing via the generative diffusion model 104, that may enable age regression based on target input specifications (for example, a target image). Subsequently, seamlessly integrated de-aged frames may be incorporated into a video stream to generate the resultant image output. In some embodiments, the output may be a video output.

[0073] The video generation model 510 may generate the video output based on the integration of the de-aged frames into the source images. The video generation model 510 may be associated with various operations, such as, the operation for video frameextraction and processing 51 OA, the operation for face detection and segmentation 51 OB, the operation for use of the re-aging model 51 OC, and the operation for re-aging frame merger to video 51 OD. The video frame extraction and processing may include extraction of frames from the video. The face detection and segmentation 51 OB may include face detection within the frames to find and identify facial regions of the user. The segmentation may involve isolation of the detected faces from the rest of the frame to focus on the facial features referred as the segmented face images. The segmented face images may undergo processing based on the re-aging model 51 OC to enable age regression based on target input specifications. Subsequently, the re-age frame merger to video 51 OD may seamlessly integrate the de-aged frames into the input image (for example, a source image) to generate the resultant video output. The video generation model 510 may transmit the re-aged video output and a re-aging status response to the re-aging node 508.

[0074] It should be noted that although the disclosure describes operations on images for re-aging, the disclosure may not be so limited. The disclosure may be also applicable to videos, moving pictures, other multi-media, and the like. It should be noted that the scenario 500 of FIG. 5 is for exemplary purposes and should not be construed to limit the scope of the disclosure.

[0075] FIG. 6A is a diagram that illustrates an exemplary generator architecture of a generative diffusion model for determination of a re-aged facial image, in accordance with an embodiment of the disclosure. FIG. 6A is explained in conjunction with elements from FIG. 1 , FIG. 2, FIG. 3, FIG. 4, and FIG. 5. With reference to FIG. 6A, there is shown an exemplary generator architecture 600A including the generator 406. The generator 406 may receive input image data (for example, the source image 114) including the first image channels 404 (for example, a 5-channel tensor) associated with an RGB image earmarkedfor re-aging, a third image channel for age maps associated with the input image data, and a fourth image channel for the output ages for each pixel. Further, the generator 406 may receive the textual prompt 118 as another input. The textual prompt 118 may be the additional text prompt that may specify the desired de-aging type. For example, the textual prompt 118 may provide the auxiliary information to the generator 406, which may incorporate the textual description of the input image of the facial region of the user. An example of a template of the textual prompt 118 may be:“{A / An} {Age}-year-old {Ethnicity} {man / woman / person} with {description of facial features}”Example prompt: 'A 20-year-old Caucasian man with trimmed beard, dark brown hair, no smile," is incorporated.

[0076] In an embodiment, the first image channels may be connected to the generator 406. The generator 406 may include the image encoder (EA) 408A, the first latent features (LA) 410, the second latent features (LT) 414, the resulting latent features 412 (LB), the text encoder (ET) 416, the linear layer 418, and the image decoder (DB) 408B. The first image channels 404 may be encoded via the image encoder 408A, while the textual prompt 118 may be processed by the text encoder 416. Latent features (e.g., the first latent features (LA) 410) of the input image may be extracted from the image encoder 408A and Clipbased text embeddings (e.g., the second latent features (LT) 414) may be extracted from the text encoder 416 based on the first latent features (LA) 410. Within the latent space, the image’s first latent features 410 and Clip-based text embeddings 414 may traverse the Brownian bridge diffusion process 408 in both forward and backward directions. The Brownian bridge diffusion process 408 may facilitate the generalization of age modification, which may encompass both re-aging and de-aging, based on the age indicated in the textual prompt 118. Examples include facial variations like with a beard,without a beard, with glasses, without glasses, and so on. Further, first latent features (LA) 410, second latent features (LT) 414, resulting latent features 412 (LB), text encoder (ET) 416 may be integrated. The integration may facilitate a more control over the de-aged or re-aged output based on the textual prompt 118. The fusion may be passed through linear layer 418 for more feature enhancement. The image decoder 408B then may produce predicted age maps. The predicted aging maps or deltas may be added to the input original image to make a final de-aged image. Training at a next layer such as, a convolution 2D (or conv 2D) 602 may involve synthetic data pairs and a composite loss function comprising FID, L1 , perceptual (LPIPS), and adversarial losses, which may ensure robustness and fidelity in re-aging outcomes. In an example, a model architecture such as, a Text-conditional Diffusion-based U-net may prioritize identity preservation and spatial fidelity. A tanH (hyperbolic tangent) 604 may be an activation function commonly used in an output layer of generator models for range matching, efficient gradient flow, and enhanced quality of the generated image. The generated image 608 based on the generator (G) 406 may be optimized using the loss calculation 610. The loss calculation 610 may be based on the target image 116. The loss calculation 610 may be performed to measure pixel-wise difference between the generated image 608 (for example, the output from the generator 406) and a ground truth image 610A. The perceptual loss may be a perceptual similarity metric that considers high-level features (such as texture and structure) rather than the pixel values. The combination of the loss terms may ensure that the generative diffusion model 104 optimizes (at 612A) for both perceptual quality and pixel fidelity.

[0077] The generator 406 may exhibit remarkable versatility during manipulation of input age maps. Users may gain precise control over re-aging effects across various facial regions. The integration of the face segmentation network may allow targeted re-aging,preservation of high-frequency details and accommodation of subtle shifts in wrinkle positions. Additionally, use of the textual prompt 118 may refine control over the de-aging process. Based on a focus on prediction of aging deltas rather than an estimation of a current age, the generator 406 may minimize computational overhead and enhance efficiency. Such streamlined approach, combined with inherent temporal smoothness over video frames, may ensure consistent and realistic re-aging results. The solution may represent a monumental advancement in the facial re-aging, which may provide unparalleled creative freedom and production-ready capabilities for real-world applications.

[0078] In the facial re-aging models, particularly in generative adversarial networks (GANs), a discriminator may be a neural network that may be trained to differentiate between the real and fake images. The discriminator applies the convolution layer (606a- 606n, as shown FIG. 6B) to generate a feature representation of the source image 114. The discriminator may use several layers with the down-sampling to process the features. The images may be classified as real image or fake image based on the GAN discriminator.

[0079] FIG. 6B is a diagram that illustrates an exemplary discriminator architecture of a generative diffusion model for determination of a re-aged facial image, in accordance with an embodiment of the disclosure. FIG. 6B is explained in conjunction with elements from FIG. 1 , FIG. 2, FIG. 3, FIG. 4, FIG. 5, and FIG. 6A. With reference to FIG. 6B, there is shown an exemplary discriminator architecture 600B including a discriminator (D) 606.

[0080] The discriminator (D) 606 in FIG 6B is a key component which may provide aging network with extra adversarial monitoring. Based on a careful selection of training dataset, the discriminator 606 may compare a re-aged RGB image 614 with an input target age map 616 and determine whether the appearance of the re-aged RGB image 614 isconsistent with the target age. With respect to image re-aging, a Text-conditional Diffusionbased ll-net may be used as the discriminator 606 to train various convolution layers of the discriminator, such as, CONV 606a - CONV 606n to discern between a real image and a fake image (for example, a re-aged image) 608A. Instances that are genuine and provide a standard of authenticity may be real examples taken from the artificial dataset with precise age labels. As 'fake' examples, on the other hand, samples produced by the re-aging network and genuine photos mixed with inaccurate age maps may be classified, as fake images for discriminator training. The discriminator 606 may follow an adversarial approach (such as, using a loss calculation or adversarial loss calculation 610 and optimization 612B) to add another degree of robustness to the MAM system's performance and improve its suitability for real-world applications by supporting the integrity of the reaging process and guaranteeing that the generated outputs closely match the target age.

[0081] FIG. 7 is a flowchart that illustrates operations of an exemplary method for merging re-aged images into coherent video sequences, in accordance with an embodiment of the disclosure. FIG. 7 is explained in conjunction with elements from FIG. 1 , FIG. 2, FIG. 3, FIG. 4, FIG. 5, FIG. 6A, and FIG. 6B. With reference to FIG. 7, there is shown an exemplary flowchart 700 for determination of merging re-aged images into coherent video sequences. The flowchart 700 may include operations 702 to 724 executed by a computing device, such as, the electronic device 102 of FIG. 1 or the circuitry 202 of FIG. 2. The operations may start at 702 and may proceed to 704.

[0082] At 704, a file ID may be fetched, and corresponding file may be downloaded from a cloud storage. The circuitry 202 may be configured to fetch the file ID of the file (e.g., an image or video) and download the file from a cloud storage, such as, the database 108 that may be hosted by the server 106. In some embodiment, the MAM system may include the cloud storage or data storage. The MAM system may retrieve a unique identifier (ID)associated with the file stored in the cloud. Once the file ID is obtained, the MAM system may download the file based on a transmission of the file from the cloud storage to a local system or another designated environment. The file ID may be crucial for location and differentiation of the file from other files stored in the same environment. Original input frames may be utilized as reference images for re-alignment of re-aged face images, and isolation of the re-aged face frames from the background to maintain background consistency.

[0083] At 706, it may be determined whether an input provided by a user is a video file. The circuitry 202 may be configured to determine whether the input provided by the user is a video file. In case, the input provided by the user is the video file, operation 708 may be executed. Otherwise, in case, the input provided by the user is other than a video file (for example, an image), control may pass to end.

[0084] At 708, video frames may be extracted from the video file. The circuitry 202 may be configured to extract video frames from the video file that may be received as the input from the user. The video frames may be the individual images that make up a video sequence. The frame extraction may involve capture of specific frames at regular intervals from the video file. The frames may serve as input for subsequent processing, such as age transformation.

[0085] At 710, the frames may be processed for face detection and face region of interest (Rol) segmentation. The circuitry 202 may be configured to process the frames for face detection and face Rol segmentation. The processed frames may include a 1 -channel source segmented image 710A, a 1 -channel target segmented image (obtained at 712), and a 3-channel source segmented image 710B. The face detection is a process for identification and location of human faces in digital images. The face detection may be an initial step in applications like facial recognition, emotion detection, or other forms ofhuman-computer interaction. The image may be divided into regions to simplify or change the representation of the image. The ROI segmentation may focus on the specific parts of the face, such as the eyes, nose, mouth, and the like.

[0086] At 712, the 1 -channel target segmented images may be obtained. The circuitry 202 may be configured to obtain the 1 -channel target segmented images based on the frame processing (at 710). The 1 -channel target segmented images may be associated with a segmentation of the determined target image 116. The fourth image channel corresponds to the second age map associated with the determined target image 116.

[0087] At 714, a 5-channel tensor with a source and a target segmented mask may be obtained. The circuitry 202 may be configured to obtain the 5-channel tensor with the source segmented mask and the target segmented mask. The 5-channel tensor with source and target segmented mask may include inputs, for example, the 3-channel source segmented image (at 710B), the 1 -channel source segmented image (at 710A), and the 1 -channel target segmented image (at 712). The tensor may be a multi-dimensional array used in machine learning to represent data. Here, the tensor may have five channels, which could mean different things depending on the context, such as color channels in an image plus additional data channels such as age maps including input age map and output age map.

[0088] At 716, the input may be received at the generator 406 of the re-aging model 510C. The circuitry 202 may be configured to receive the input at the generator 406 of the re-aging model 510C. The input may be the 5-channel tensor. The predicted aging deltas may be the output from the generator 406. The re-aging model 510C may include the reaging service bus that may operate as a lightweight HTTP server layered atop the internal REST API (IRA) and serve as a conduit between microservices and process management modules. Initial steps may involve frame extraction from videos, followed by facerecognition and segmentation. The segmented face images may undergo processing via the re-aging model 51 OC, to enable age regression based on target input specifications. Subsequently, seamlessly integrated de-aged frames may be incorporated into the video stream, for the generation of the resultant video output.

[0089] At 718, the source image 114 and the predicted aging deltas may be integrated. The circuitry 202 may be configured to integrate the source image 114 and the predicted aging deltas. The predicted aging deltas may be added to the input original image (e.g., the 3-channel source segmented image 71 OB) to generate a final de-aged image.

[0090] At 720, the final re-aged output may be merged with the video frames. The circuitry 202 may be configured to merge the final re-aged output with the video frames. Original input frames (e.g., the source image 114) and de-aged face frames may be used to generate the final de-aged frame. Initially, the original input frames may be utilized as a reference to realign the de-aged face images and isolate the de-aged face images from the background to maintain background consistency.

[0091] At 722, a seamless video integration operation may be executed. The circuitry 202 may be configured to execute the seamless video integration operation. Based on a utilization of cutting-edge video processing algorithms, the re-aged frames may be merged with prioritization of temporal coherence. Such an intricate process may involve blending facial features meticulously, seamlessly integrating de-aged output faces onto the original input frames or source images 114. The intricate process may ensure consistency in background details, harmonize visual elements, and minimize potential artifacts. To further enhance realism, color consistency may be achieved through color transfer and sharpening, alignment of a color profile of the de-aged face with that of the original source image 114 to facilitate a more natural transition. Additionally, temporal interpolation methods may be employed to ensure fluid transitions between consecutive frames, whichmay preserve the natural motion dynamics across the entire video sequence. Based on intricate synchronization of the re-aged frames with the original source image 114, a cohesive and visually consistent video output may be obtained such that an integrity of facial features and expressions may be maintained throughout the duration of the sequence. Such an approach to video generation may ensure a high-quality viewing experience, free from perceptible inconsistencies, and enhanced overall realism and fidelity of the re-aging process for various entertainment applications.

[0092] At 724, the generated re-aged video may be stored in the MAM system. The circuitry 202 may be configured to store the generated re-aged video in a storage device associated with the MAM system. For example, the generated re-aged video may be stored on the database 108 hosted by the server 106. Control may pass to end.

[0093] Although the flowchart 700 is illustrated as discrete operations, such as 704, 706, 708, 710, 712, 714, 716, 718, 720, 722, and 724, the disclosure is not so limited. Accordingly, in certain embodiments, such discrete operations may be further divided into additional operations, combined into fewer operations, or eliminated, depending on the implementation without detracting from the essence of the disclosed embodiments.

[0094] FIG. 8 is a figure that illustrates an exemplary execution pipeline of video integration based on compilation of target images into video frames upon determination of re-aged facial image, in accordance with an embodiment of the disclosure. FIG. 8 is explained in conjunction with elements from FIG. 1 , FIG. 2, FIG. 3, FIG. 4, FIG. 5, FIG. 6A, FIG. 6B, and FIG. 7. With reference to FIG. 8, there is shown an exemplary execution pipeline 800 including an output re-aged face frame(s) 802A (for example, the target image 116) and original input frame(s) 802B (for example, the source image 114). The execution pipeline 800 may further include a set of operations such as, an operation for realignment 804, an operation for seamless frame face blending, an operation of color transfer 808(from the original input image), an operation for sharpening 810, an operation for temporal interpolation 812, and an operation for video compilation 814. The set of operations may be executed by the electronic device 102 of FIG. 1 or the circuitry 202 or FIG. 2.

[0095] At 804, the operation for realignment may be executed. The circuitry 202 may be configured to realign the output re-aged face frame 802A and the original input frames 802B. The original input frames 802B may have multiple frames that may include differences in face alignment compared to the output re-aged face frame 802A. To blend the re-aged facial image into the original input frames 802B, the re-aged facial image may be realigned to match the original input frames 802B. The realignment may be performed based on input coordinates of the face of the user as reference to realign the re-aged facial images.

[0096] At 806, the operation for seamless frame face blending may be executed. The circuitry 202 may be configured to blend the re-aged facial images into the original input frames 802B. Blending may involve face mask with different settings, such as, an adjusted mask size and blur level of the re-aged mask during blending. Additionally, a histogrambased matching may be applied to match the re-aged mask and the original face based on the histograms. The face blending may be performed by using, for example, but not limited to, a Poisson blending.

[0097] At 808, the operation for color transfer (from the original input image 802B) may be executed. The circuitry 202 may be configured to transfer color from the original input frame 802B. The color transfer may be performed to achieve a more consistent appearance that may resemble the original input frame 802B. In some embodiments, color transfer and sharpening may be performed based on, but not limited to, a Reinhard color transfer and an iterative distribution transfer method. Such methods may aim to approximate the color of the original input frame 802B to the re-aged facial image. Theblending process may consider variations in skin tones, face shapes, and illumination conditions, particularly at the junctions between the original input frame 802B and the delimited region, as well as the target face.

[0098] At 810, the operation for sharpening may be executed. The circuitry 202 may be configured to sharpen the output re-aged face frame 802A. For face sharpening, a pretrained face super-resolution neural network may be used. The neural network may perform smoothening of the re-aged face frame 802A and remove moles and wrinkles.

[0099] At 812, the operation for temporal interpolation may be executed. The circuitry 202 may be configured to apply a temporal interpolation technique on the output re-aged face frame 802A to ensure smooth transitions between the re-aged face frames 802A, that may maintain a natural flow of motion and facial expressions throughout the video. The temporal interpolation technique may be, for example, but not limited to, a spline interpolation or optical flow-based methods. Such techniques may interpolate intermediate frames based on the motion and appearance of neighboring frames, which may result in fluid motion and consistent facial expressions.

[0100] At 814, the operation for video compilation may be executed. The circuitry 202 may be configured to compile a re-aged video based on the re-aged face frames 802A. The re-aged face frames 802A may be compiled to prepare the video. Once the frames are processed (from 804 to 812), the processed re-aged face frames may be compiled into a sequence of images. Each image may represent a frame of a final video. In some embodiment, audio may be added to the sequence of images. The sequence of images may be encoded using a method for example, but not limited to, an ffmpeg to convert to video format. The video encoding parameters may be set to maintain the quality of the original video, including resolution, frame rate, and compression settings.

[0101] FIG. 9 is a flowchart that illustrates determination of re-aged facial image based on generative diffusion model, in accordance with an embodiment of the disclosure. FIG. 9 is explained in conjunction with elements from FIG. 1 , FIG. 2, FIG. 3, FIG. 4, FIG. 5, FIG. 6A, FIG.6B, FIG. 7, and FIG. 8. With reference to FIG. 9, there is shown a flowchart 900 for determination of re-aged facial image based on generative diffusion model 104. The operations from 902 to 916 may be implemented by any computing system, such as, by the electronic device 102 of FIG. 1. The operations may start at 902 and may proceed to 904.

[0102] At 904, the source image 114 including the facial region of the user may be received. The circuitry 202 may be configured to receive the source image 114 (for example, image frames, video frames, and the like) including a facial region (head, eyes, ear, mouth, and the like) of the user. The source image 114 may be stored in the database 108. The database 108 may store the information (for example, source image) of the user. The reception of the source image is described further, for example, in FIG. 3 (at 302).

[0103] At 906, the textual prompt 118 indicative of the set of facial attributes of the target image 116 of the facial region of the user may be received, wherein the set of facial attributes may include the target age of the user in the target image 116. The circuitry 202 may be configured to receive the textual prompt 118 indicative of the set of facial attributes of the target image 116 of the facial region of the user. The textual prompt 118 may include for example, a target age of the user within the source image 114, ethnicity of the user, gender of the user, description of facial features of the user, and the like. The facial attributes of the target image 116 may include the target ethnicity of the user, the target gender of the user, the description of the target facial features of the user, and the like. The reception of the textual prompt is described further, for example, in FIG. 3 (at 304).

[0104] At 908, the generative diffusion model 104 may be applied on the received source image 114, based on the received textual prompt 118. The circuitry 202 may be configured to apply generative diffusion model 104 on the received source image 114, based on the received textual prompt 118. The generative diffusion model 104 may generalize the age modification. The generative diffusion model 104 may encompass both re-aging and de-aging, based on the age indicated in the textual prompt 118. The examples may include facial variations such as beard, without beard, with glasses, without glasses and so on. The application of the generative diffusion model is described further, for example, in FIG. 3 (at 306).

[0105] At 910, the set of re-aging delta images may be predicted based on the application of the generative diffusion model 104. The circuitry 202 may be configured to predict the set of re-aging delta images based on the application of the generative diffusion model 104. The generative diffusion model 104 may generate a target image with a generalized age of the source image 114 that may encompass both re-aging and de-aging, based on the age indicated in the textual prompt 118. The CLIP text embedding may be used to control the mechanism over the de-aged or re-aged output based on the textual prompt 118. Then the image decoder 408B may produce predicted age maps (for example, age deltas) that may be added to the source image 114 to generate the target image 116. The target image 116 can be the de-aged image or re-aged image. The prediction of the set of re-aged delta images is described further, for example, in FIG. 3 (at 308).

[0106] At 912, the target image 116 of the facial region of the user may be determined based on the predicted set of re-aging delta images and the received source image, wherein the determined target image 116 may correspond to the set of facial attributes of the target image 116. The circuitry 202 may be configured to determine the target image 116 of the facial region of the user, based on the predicted set of re-aging delta imagesand the received source image 114. The determined target image 116 corresponds to the set of facial attributes of the target image 116. The target image 116 may be stored in database 108. The determination of the target image is described further, for example, in FIG. 3 (at 310).

[0107] Although the flowchart 900 is illustrated as discrete operations, such as 904, 906, 908, 910, 912, 914, and 916, the disclosure is not so limited. Accordingly, in certain embodiments, such discrete operations may be further divided into additional operations, combined into fewer operations, or eliminated, depending on the implementation without detracting from the essence of the disclosed embodiments.

[0108] Various embodiments of the disclosure may provide a non-transitory computer- readable medium and / or storage medium having stored thereon, computer-executable instructions executable by a machine and / or a computer to operate an electronic device (such as the electronic device 102). The computer-executable instructions may cause the machine and / or computer to perform operations that include determination of re-aged facial image based on a generative diffusion model. The operations may include reception of a source image (e.g., the source image 114) including a facial region of a user. The operations may further include reception of a textual prompt (e.g., the textual prompt 118) indicative of a set of facial attributes of a target image of the facial region of the user. The set of facial attributes may include a target age of the user in the target image. The operation may further include application of a generative diffusion model (e.g., the generative diffusion model 104) on the received source image 114, based on the received textual prompt 118. The operations may further include prediction of a set of re-aging delta images based on the application of the generative diffusion model 104. The operations may further include determination of the target image of the facial region of the user, based on the predicted set of re-aging delta images and the received source image 114. Thedetermined target image 116 may correspond to the set of facial attributes of the target image.

[0109] Exemplary aspects of the disclosure may include an electronic device (such as, the electronic device 102 of FIG. 1 ) that may include circuitry (such as, the circuitry 202), that may be communicatively coupled to the electronic device (such as, the electronic device 102 of FIG. 1 ). The electronic device 102 may further include memory (such as, the memory 204 of FIG. 2). The circuitry 202 may be configured to receive the source image 114 including a facial region of a user. The circuitry 202 may be configured to receive the textual prompt 118 indicative of a set of facial attributes of a target image of the facial region of the user. The set of facial attributes may include a target age of the user in the target image. The circuitry 202 may be further configured to apply the generative diffusion model 104 on the received source image 114, based on the received textual prompt 118. The circuitry 202 may be configured to predict a set of re-aging delta images based on the application of the generative diffusion model 104. The circuitry 202 may be further configured to determine the target image 116 of the facial region of the user, based on the predicted set of re-aging delta images and the received source image 114. The determined target image 116 may correspond to the set of facial attributes of the target image 116.

[0110] In accordance with an embodiment, the received source image 114 may correspond to a set of media frames received from a Media Asset Management (MAM) system via a Hypertext Transfer Protocol (HTTP)-based Representational State Transfer (REST) Application Programming Interface (API).

[0111] In accordance with an embodiment, the set of facial attributes of the target image 116 may further include at least one of a target ethnicity of the user, a target gender of the user, and a description of target facial features of the user.

[0112] In accordance with an embodiment, the generative diffusion model 104 may correspond to a Brownian bridge diffusion model.

[0113] In accordance with an embodiment, the circuitry 202 may be further configured to segment the received source image 114 into a set of first intermediate images associated with a set of first image channels. The application of the generative diffusion model 104 may be further based on the set of first intermediate images.

[0114] In accordance with an embodiment, the set of first image channels may correspond to at least one of a set of second image channels associated with the segmented source image 114, a third image channel associated with the segmented source image 114, or a fourth image channel associated with a segmentation of the determined target image 116.

[0115] In accordance with an embodiment, the set of second image channels corresponds to color information associated with the segmented source image 114.

[0116] In accordance with an embodiment, the third image channel corresponds to a first age map associated with the received source image 114.

[0117] In accordance with an embodiment, the fourth image channel corresponds to a second age map associated with the determined target image 116.

[0118] In accordance with an embodiment, the generative diffusion model 104 may include an image encoder configured to determine an encoded image based on the set of first intermediate images. The determined encoded image may correspond to first latent features of the received source image 114.

[0119] In accordance with an embodiment, the generative diffusion model 104 may further include a text encoder configured to determine a textual embedding based on the received textual prompt 118. The determined textual embedding may correspond to second latent features of the received textual prompt 118.

[0120] In accordance with an embodiment, the generative diffusion model 104 may further include a linear layer configured to enhance each feature of the first latent features and the second latent features.

[0121] In accordance with an embodiment, the generative diffusion model 104 may further includes a diffusion model configured to traverse each feature of the first latent features and the second latent features as at least one of a feedforward feature or feedback feature associated with the diffusion model. The generative diffusion model 104 may further include determination of a set of second intermediate images based on the traversal of each feature of the first latent features and the second latent features.

[0122] In accordance with an embodiment, the generative diffusion model 104 may further include a decoder model configured to determine a set of decoded images based on the determined set of second intermediate images. The determined set of decoded images may correspond to the predicted set of re-aging delta images.

[0123] In accordance with an embodiment, the circuitry 202 may be further configured to merge the received source image 114 and the predicted set of re-aging delta images. The determination of the target image 116 of the facial region of the user may be further based on the merging of the received source image 114 and the predicted set of re-aging delta images.

[0124] In accordance with an embodiment, the circuitry 202 may be further configured to re-align the determined target image 116 with the received source image 114 to determine a re-aligned facial image and blend a face mask of the re-aligned facial image over the received source image 114 to determine a blended facial image. The circuitry 202 may be further configured to transfer color attributes from the received source image to the blended facial image and sharpen the blended facial image based on the transferred color attributes. The circuitry 202 may be further configured to temporally interpolate thesharpened blended facial image based on a set of neighboring frames associated with a video corresponding to the received source image 114. The video may include the sharpened blended facial image and the set of neighboring frames. The set of neighboring frames may be frames adjacent to the sharpened blended facial image in the video and generate the video based on the temporal interpolation.

[0125] The present disclosure may be realized in hardware, or a combination of hardware and software. The present disclosure may be realized in a centralized fashion, in at least one computer system, or in a distributed fashion, where different elements may be spread across several interconnected computer systems. A computer system or other apparatus adapted to carry out the methods described herein may be suited. A combination of hardware and software may be a general-purpose computer system with a computer program that, when loaded and executed, may control the computer system such that it carries out the methods described herein. The present disclosure may be realized in hardware that comprises a portion of an integrated circuit that also performs other functions.

[0126] The present disclosure may also be embedded in a computer program product, which comprises all the features that enable the implementation of the methods described herein, and which when loaded in a computer system is able to carry out these methods. Computer program, in the present context, means any expression, in any language, code or notation, of a set of instructions intended to cause a system with information processing capability to perform a particular function either directly, or after either or both of the following: a) conversion to another language, code or notation; b) reproduction in a different material form.

[0127] While the present disclosure is described with reference to certain embodiments, it will be understood by those skilled in the art that various changes may be made, andequivalents may be substituted without departure from the scope of the present disclosure. In addition, many modifications may be made to adapt a particular situation or material to the teachings of the present disclosure without departure from its scope. Therefore, it is intended that the present disclosure is not limited to the embodiment disclosed, but that the present disclosure will include all embodiments that fall within the scope of the appended claims.

Claims

CLAIMSWhat is claimed is:1 . An electronic device, comprising: circuitry configured to: receive a source image including a facial region of a user; receive a textual prompt indicative of a set of facial attributes of a target image of the facial region of the user, wherein the set of facial attributes includes a target age of the user in the target image; apply a generative diffusion model on the received source image, based on the received textual prompt; predict a set of re-aging delta images based on the application of the generative diffusion model; and determine the target image of the facial region of the user, based on the predicted set of re-aging delta images and the received source image, wherein the determined target image corresponds to the set of facial attributes of the target image.

2. The electronic device according to claim 1 , wherein the received source image corresponds to a set of media frames received from a Media Asset Management (MAM) system via a Hypertext Transfer Protocol (HTTP)-based Representational State Transfer (REST) Application Programming Interface (API).

3. The electronic device according to claim 1 , wherein the set of facial attributes of the target image further includes at least one of:a target ethnicity of the user, a target gender of the user, or a description of target facial features of the user.

4. The electronic device according to claim 1 , wherein the generative diffusion model corresponds to a Brownian bridge diffusion process.

5. The electronic device according to claim 1 , wherein the circuitry is further configured to: segment the received source image into a set of first intermediate images associated with a set of first image channels, wherein the application of the generative diffusion model is further based on the set of first intermediate images.

6. The electronic device according to claim 5, wherein the set of first image channels corresponds to at least one of: a set of second image channels associated with the segmented source image, a third image channel associated with the segmented source image, or a fourth image channel associated with a segmentation of the determined target image.

7. The electronic device according to claim 6, wherein the set of second image channels corresponds to color information associated with the segmented source image.

8. The electronic device according to claim 6, wherein the third image channel corresponds to a first age map associated with the received source image.

9. The electronic device according to claim 6, wherein the fourth image channel corresponds to a second age map associated with the determined target image.

10. The electronic device according to claim 5, wherein the generative diffusion model includes an image encoder configured to: determine an encoded image based on the set of first intermediate images, wherein the determined encoded image corresponds to first latent features of the received source image.11 . The electronic device according to claim 10, wherein the generative diffusion model further includes a text encoder configured to: determine a textual embedding based on the received textual prompt, wherein the determined textual embedding corresponds to second latent features of the received textual prompt.

12. The electronic device according to claim 11 , wherein the generative diffusion model further includes a linear layer configured to enhance each feature of the first latent features and the second latent features.

13. The electronic device according to claim 11 , wherein the generative diffusion model further includes a diffusion model configured to: traverse each feature of the first latent features and the second latent features as at least one of a feedforward feature or feedback feature associated with the diffusion model; and determine a set of second intermediate images based on the traversal of each feature of the first latent features and the second latent features.

14. The electronic device according to claim 13, wherein the generative diffusion model further includes a decoder model configured to: determine a set of decoded images based on the determined set of second intermediate images, wherein the determined set of decoded images corresponds to the predicted set of re-aging delta images.

15. The electronic device according to claim 1 , wherein the circuitry is further configured to: merge the received source image and the predicted set of re-aging delta images, wherein the determination of the target image of the facial region of the user is further based on the merging of the received source image and the predicted set of re-aging delta images.

16. The electronic device according to claim 1 , wherein the circuitry is further configured to:re-align the determined target image with the received source image to determine a re-aligned facial image; blend a face mask of the re-aligned facial image over the received source image to determine a blended facial image; transfer color attributes from the received source image to the blended facial image; sharpen the blended facial image based on the transferred color attributes; temporally interpolate the sharpened blended facial image based on a set of neighboring frames associated with a video corresponding to the received source image, wherein the video includes the sharpened blended facial image and the set of neighboring frames, and the set of neighboring frames are frames adjacent to the sharpened blended facial image in the video; and generate the video based on the temporal interpolation.

17. A method, comprising: in an electronic device: receiving source image including a facial region of a user; receiving a textual prompt indicative of a set of facial attributes of a target image of the facial region of the user, wherein the set of facial attributes includes a target age of the user in the target image; applying a generative diffusion model on the received source image, based on the received textual prompt;predicting a set of re-aging delta images based on the application of the generative diffusion model; and determining the target image of the facial region of the user, based on the predicted set of re-aging delta images and the received source image, wherein the determined target image corresponds to the set of facial attributes of the target image.

18. The method according to claim 17, wherein the set of facial attributes of the target image further includes at least one of: a target ethnicity of the user, a target gender of the user, or a description of target facial features of the user.

19. The method according to claim 17, further comprising: re-aligning the determined target image with the received source image to determine a re-aligned facial image; blending a face mask of the re-aligned facial image over the received source image to determine a blended facial image; transferring color attributes from the received source image to the blended facial image; sharpening the blended facial image based on the transferred color attributes; temporally interpolating the sharpened blended facial image based on a set of neighboring frames associated with a video corresponding to the received source image, whereinthe video includes the sharpened blended facial image and the set of neighboring frames, and the set of neighboring frames are frames adjacent to the sharpened blended facial image in the video; and generating the video based on the temporal interpolation.

20. A non-transitory computer-readable medium having stored thereon, computerexecutable instructions that when executed by an electronic device, causes the electronic device to execute operations, the operations comprising: receiving source image including a facial region of a user; receiving a textual prompt indicative of a set of facial attributes of a target image of the facial region of the user, wherein the set of facial attributes includes a target age of the user in the target image; applying a generative diffusion model on the received source image, based on the received textual prompt; predicting a set of re-aging delta images based on the application of the generative diffusion model; and determining the target image of the facial region of the user, based on the predicted set of re-aging delta images and the received source image, wherein the determined target image corresponds to the set of facial attributes of the target image.