Gated Network Generator for Digital Human Image Coordinate Adhesion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Generative adversarial networks, such as StyleGAN3, face challenges with image coordinate adhesion, leading to blurry facial details and reduced user experience in digital human video generation, particularly due to complex network requirements that hinder high automation applications.

Innovation Solution

A gated network-based generator is introduced, comprising an image input layer, feature encoding layer, and image output layer, utilizing gated convolutional and inverse gated convolutional networks to process and decode image sequences, along with an audio input layer for enhanced feature extraction and image-audio fusion, to resolve coordinate adhesion and improve detail clarity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If a general generator network architecture (convolutional structure, nonlinear structure, upsampling structure) is used, then the network can process images, but coordinate adhesion occurs causing blurry facial details

Engineering Contradiction:
Improvefacial detail clarityVSAvoidcoordinate adhesion problem
Core Design Contradiction:
Manufacturing precisionVSReliability

Solution Approach 1:

The patent changes the fundamental parameters of the generator network architecture by replacing standard convolutional layers with equivariant convolutional layers that incorporate rotation group operations. This parameter change in the mathematical foundation of the network ensures coordinate adhesion is avoided while maintaining facial detail clarity through the inherent equivariance properties of the transformed architecture.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If StyleGAN3 network is used to resolve coordinate adhesion, then coordinate adhesion is reduced, but the network becomes too complex and requires a lot of manual intervention

Engineering Contradiction:
Improvecoordinate adhesion resolutionVSAvoidnetwork complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the complex StyleGAN3 architecture into a simplified equivariant generator framework that processes images through rotation-equivariant convolutional layers. This segmentation approach maintains the essential functionality of resolving coordinate adhesion while eliminating unnecessary complexity and manual intervention requirements by focusing on the core equivariance mechanism.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent extracts and isolates the critical equivariance property from the complex StyleGAN3 network, creating a streamlined generator that implements only the necessary rotation-equivariant operations. This extraction removes redundant components and manual intervention steps while preserving the coordinate adhesion resolution capability.

Inventive Principle:
Principle #2Taking out (Extraction)

3Reliability

If StyleGAN3 network is used, then coordinate adhesion is addressed, but high automation requirements cannot be met

Engineering Contradiction:
Improvecoordinate adhesion resolutionVSAvoidautomation capability
Core Design Contradiction:
ReliabilityVSExtent of automation

Solution Approach 1:

The equivariant generator network is designed to automatically maintain coordinate consistency through its inherent mathematical properties, eliminating the need for manual intervention or post-processing adjustments. The rotation-equivariant convolutional layers self-correct coordinate adhesion issues during the generation process, enabling high automation in digital human video production workflows.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS12056903B2Generator, generator training method, and method for avoiding image coordinate adhesion
Publication Date: 2024.08.06 NANJING SILICON INTELLIGENCE TECH CO LTD
  • US12056903B2 patent drawing
  • US12056903B2 patent drawing
  • US12056903B2 patent drawing

AI summary

Disclosed are a gated network-based generator, a generator training method, and a method for avoiding image coordinate adhesion. The generator processes, by using an image input layer, a to-be-processed image as an image sequence and inputs it to a feature encoding layer. Multiple feature encoding layers encode the image sequence by using a gated convolutional network, to obtain an image code. Moreover, multiple image decoding layers decode the image code by using an inverse gated convolution unit, to obtain a target image sequence. Finally, an image output layer splices the target image sequence to obtain a target image. Therefore, a character feature in the obtained target image is more obvious, making details of a facial image of generated digital human more vivid, whereby solving a problem of image coordinate adhesion in a digital human image generated by an existing generator using a generative adversarial network, and improving user experience.