Face feature diversified synthesis self-adaptive digital human system for real-time interaction

By designing a system that includes data collection, feature diversity synthesis, privacy protection, intelligent interaction and adaptability modules, the existing digital human system technology is insufficient, low degree of diversity, privacy leakage and adaptability, and a more stable, personalized and secure real-time interactive experience is achieved.

CN120014124APending Publication Date: 2025-05-16CHINA UNICOM WO MUSIC & CULTURE CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411816280.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-11
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing synthetic adaptive digital human system for real-time interaction has problems such as insufficient technological maturity, difficulty in generating sufficiently diverse and personalized digital human images, privacy leakage or security risks, and lack of adaptability and scalability.

Method used

A system including data acquisition and preprocessing module, feature diversity synthesis module, privacy protection and data encryption module, intelligent interaction and processing module, and adaptability and scalability module are designed. The system uses a convolutional neural network algorithm for face recognition and feature extraction, and generates diverse face features using a Generative Adversarial Network (GANs) algorithm. At the same time, the data security is ensured through the privacy protection and data encryption module, and the adaptability and scalability module provides flexible interface and architectural support.

Benefits of technology

It solves the problems of lag, delay or error recognition that may occur in the system during real-time interaction, improves the diversity and personalization of the digital person's image, ensures the security of user data, and enhances the adaptability and scalability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014124A_ABST
    Figure CN120014124A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of digital human systems, in particular to a real-time interaction-oriented face feature diversification synthesis self-adaptive digital human system, which comprises a data acquisition and preprocessing module which is coupled with a feature diversification synthesis module. And the feature diversification synthesis module is coupled with a privacy protection and data encryption module. According to the real-time interaction-oriented face feature diversified synthesis adaptive digital human system, the problem that an existing system possibly faces insufficient technical maturity is solved, so that unstable conditions such as lagging, delay or misrecognition occur in the real-time interaction process; in practical application, due to limitation of an algorithm or a data set, sufficiently diversified and personalized digital human images are difficult to generate; privacy leakage or security risks may exist when an existing system processes the data. The problems that an existing system possibly lacks enough adaptability and expandability and is difficult to adapt to changes of different scenes and requirements are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of digital human systems, and in particular to a real-time interactive face feature diversified synthesis adaptive digital human system. Background Art

[0002] The adaptive digital human system with diversified synthesis of facial features for real-time interaction is a highly intelligent system that combines advanced computer graphics, deep learning, natural language processing and multimodal interaction technology. The system supports real-time facial feature capture and synthesis, and can dynamically adjust according to the user's real-time input (such as expression, action, etc.) to achieve a smooth interactive experience. Using deep learning technology, the system can synthesize a variety of different facial features, including but not limited to gender, age, race, expression, etc., to generate rich and diverse digital human images. The system has strong adaptive capabilities and can automatically adjust the digital human's expression and interaction strategy according to the user's interactive behavior and preferences, providing more personalized services, providing digital humans with realistic appearance and action performance, making digital humans visually close to real humans, and playing a key role in facial feature synthesis, speech recognition, natural language processing, etc., enabling digital humans to understand and generate complex language and behavior, enabling digital humans to understand and generate natural language, and have smooth conversations and exchanges with users, combining speech, image, Text and other multi-modal information to achieve a more natural and rich interactive experience. As a virtual idol, virtual anchor and other roles, it provides users with a novel and unique interactive experience. As a virtual teacher or lecturer, it provides personalized teaching plans to improve teaching efficiency and quality. As a brand spokesperson or customer service representative, it interacts with consumers to enhance brand image and customer satisfaction. It realizes three-dimensional communication across time and space on social platforms and opens up a new social model. With the development of AI large models, digital human systems will be able to handle more complex data and tasks and achieve more natural and intelligent interactions. Digital human systems will gradually penetrate into more fields, such as medical care, finance, and smart manufacturing, providing more intelligent and efficient services for all walks of life. The system will pay more attention to the personalized customization of user experience and provide users with more intimate and personalized services by analyzing users' historical behaviors and preference data. Although the adaptive digital human system with diversified synthesis of facial features for real-time interaction has broad development prospects, it still faces many challenges, such as technical maturity, privacy protection, and ethical issues. In the future, with the continuous advancement of technology and the increasing richness of application scenarios, the digital human system is expected to become an indispensable part of human life and achieve a more harmonious and symbiotic relationship with humans. In summary, the digital human system with diversified synthesis of facial features for real-time interaction is a highly intelligent system with the characteristics of real-time interaction, diversified synthesis of facial features and adaptive capabilities. It has broad application prospects in entertainment, education, commercial marketing and social networking. With the continuous advancement of technology and the expansion of application scenarios, the system will bring more convenient, intelligent and personalized service experience to humans;

[0003] However, the existing face feature diversified synthesis adaptive digital human system for real-time interaction still has the following problems in actual use:

[0004] Existing systems may face the problem of insufficient technical maturity, resulting in instability such as freezing, delays or misrecognition during real-time interaction;

[0005] In practical applications, it may be difficult to generate sufficiently diverse and personalized digital human images due to limitations of algorithms or data sets;

[0006] Existing systems may have privacy leaks or security risks when processing this data;

[0007] Existing systems may lack sufficient adaptability and scalability to adapt to different scenarios and changes in demand;

[0008] To solve the above problems, a facial feature diversified synthesis adaptive digital human system for real-time interaction is provided. Summary of the invention

[0009] The purpose of the present invention is to provide a real-time interactive face feature diversified synthesis adaptive digital human system to solve the problems raised in the above background technology. To achieve the above purpose, the present invention provides the following technical solutions: a real-time interactive face feature diversified synthesis adaptive digital human system, including a data acquisition and preprocessing module, the data acquisition and preprocessing module is coupled to a feature diversified synthesis module, the feature diversified synthesis module is coupled to a privacy protection and data encryption module, the privacy protection and data encryption module is coupled to an intelligent interaction and processing module, and the intelligent interaction and processing module is coupled to an adaptability and scalability module.

[0010] Preferably, the data acquisition and preprocessing module uses a convolutional neural network algorithm to perform face recognition and feature extraction.

[0011] Preferably, the convolutional neural network algorithm is as follows: the input layer: input a digital face image, the number of CNN model layers L and the types of all hidden layers; the convolution layer defines the size K of the convolution kernel, the dimension F of the convolution kernel matrix, the padding size P, and the stride S; the pooling layer defines the pooling area size k and the pooling standard; the fully connected layer defines the activation function and the number of neurons in each layer; the output layer: the output value of the CNN model is aL;

[0012] A. Filling the edge of the digital face image according to the padding size P of the input layer to obtain an input tensor al;

[0013] B. Initialize the parameters W and b of all the hidden layers;

[0014] C. for l = 2 to L-1:

[0015] (1) If the first layer is the convolution layer, the side output is

[0016] a L =σ(z l )=σ(a l-1 *W 1 +b 1 );

[0017] (2) If the first layer is the pooling layer, the side output is

[0018] a L =pool(a l-1 )(pool refers to the process of reducing the input tensor according to the pooling area size k and the pooling criterion);

[0019] (3) If the first layer is the fully connected layer, the output is

[0020] al=σ(zl)=σ(Wlal-1+bl);

[0021] (4) For the output layer L:

[0022] aL=σ(ZL)=σ(WLaL_1+bL);

[0023] Among them, the superscript represents the number of layers, W represents the convolution kernel, b represents the bias, and σ is the activation function ReLU.

[0024] Preferably, the feature diversity synthesis module uses a generative adversarial network (GANs) algorithm to generate diverse facial features.

[0025] Preferably, the generative adversarial network (GANs) algorithm includes a generator and a discriminator;

[0026] The generator and the discriminator are both composed of a convolutional neural network, the convolution kernel size of the input convolution layer of the convolutional neural network is 2×2, and the convolution kernel size of the remaining convolution layers is 1×3;

[0027] The generator is used to generate handwritten characters in the air with variable length, and its input is the mean vector of each class mixed with noise;

[0028] The discriminator is used to perform adversarial training with the generator, so as to make the generated variable-length aerial handwritten characters closer to real data;

[0029] The discriminator includes a global average pooling layer GAP, and the average pooling layer GAP is used to convert feature maps of different sizes generated after the discriminator performs a convolution operation into vectors of equal length.

[0030] Preferably, the generator includes a first generation network module and a second generation network module, both of which are residual structures. The generator also includes a third generation network module and a fourth generation network module, and the third network module and the fourth network module are used for feature extraction. The discriminator includes a first discriminant network module and a second discriminant network module, both of which are residual structures. The discriminator also includes a third discriminant network module, and the third discriminant network module is used for feature extraction. The first discriminant network module, the second discriminant network module and the third discriminant network module all include a Padding layer, and the Padding layer is used to align the length of the feature map after the input and convolution operation. The discriminator also includes a fully connected layer, and the fully connected layer is used to map the "distributed feature representation" learned by the generative adversarial network model to the sample labeling space. The discriminator also includes a binary classification layer, and the loss function of the discriminator is a binary cross entropy loss function BCE. The input of the discriminator is a real handwritten character sample in the air and a virtual handwritten character sample in the air generated by the generator.

[0031] Preferably, the data acquisition and preprocessing module includes a signal conditioning module, the signal conditioning module is coupled to a data cleaning and conversion module, the data cleaning and conversion module is coupled to a data integration module, and the data integration module is coupled to a data storage and management module.

[0032] Preferably, the signal conditioning module includes a signal amplifying unit, the signal amplifying unit is coupled to a signal filtering unit, the signal filtering unit is coupled to a signal isolation unit, the signal isolation unit is coupled to a signal conversion unit, the signal conversion unit is coupled to a signal linear unit, the signal linear unit is coupled to a signal protection unit, and the signal protection unit is coupled to a signal distribution unit.

[0033] Preferably, the signal amplifying unit comprises an input unit, the input unit is coupled to a gain unit, the gain unit is coupled to an output unit, the output unit is coupled to a bias circuit unit, and the bias circuit unit is coupled to a feedback network unit.

[0034] Preferably, the feedback network unit is composed of a resistor and a capacitor.

[0035] Compared with the prior art, the present invention has the following beneficial effects:

[0036] The data acquisition and preprocessing module transmits the processed data to the feature diversification synthesis module to generate diversified digital human images. The digital human images generated by the feature diversification synthesis module will be transmitted to the intelligent interaction and processing module for real-time interaction with users. The privacy protection and data encryption module will encrypt and protect the data transmission and storage process of the entire system to ensure the security of user data. The adaptability and scalability module will provide flexible interface and architecture support for other modules so that the system can easily adapt to changes in different scenarios and needs, solving the problem that the existing system may face insufficient technical maturity, resulting in instability such as freezes, delays or misrecognition during real-time interaction; in actual applications, it may be difficult to generate sufficiently diverse and personalized digital human images due to the limitations of algorithms or data sets; the existing system may have privacy leaks or security risks when processing these data; the existing system may lack sufficient adaptability and scalability and is difficult to adapt to changes in different scenarios and needs. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 is a system block diagram of the present invention;

[0038] Figure 2 This is a system block diagram of the data acquisition and preprocessing module of the present invention. DETAILED DESCRIPTION

[0039] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technical personnel in this field without creative work are within the scope of protection of the present invention.

[0040] See also Figure 1 to Figure 2 The present invention provides a technical solution: a real-time interactive facial feature diversified synthesis adaptive digital human system, including a data acquisition and preprocessing module, the data acquisition and preprocessing module is coupled with a feature diversified synthesis module, the feature diversified synthesis module is coupled with a privacy protection and data encryption module, the privacy protection and data encryption module is coupled with an intelligent interaction and processing module, and the intelligent interaction and processing module is coupled with an adaptability and scalability module.

[0041] In this embodiment, the data acquisition and preprocessing module uses a convolutional neural network algorithm to perform face recognition and feature extraction.

[0042] In this embodiment, the convolutional neural network algorithm is as follows: input layer: input a human face digital image, the number of CNN model layers L and the types of all hidden layers; the convolution base layer defines the size K of the convolution kernel, the dimension F of the convolution kernel matrix, the padding size P, and the stride S; the pooling layer defines the pooling area size k and the pooling standard; the fully connected layer defines the activation function and the number of neurons in each layer; output layer: the output value of the CNN model is aL;

[0043] A. Fill the edge of the face digital image according to the padding size P of the input layer to obtain the input tensor al;

[0044] B. Initialize the parameters W,b of all hidden layers;

[0045] C. for l = 2 to L-1:

[0046] (1) If the first layer is a convolution layer, the side output is

[0047] a L =σ(z l )=σ(a l-1 *W 1 +b 1 );

[0048] (2) If the first layer is a pooling layer, the side output is

[0049] a L =pool(a l-1 )(pool refers to the process of reducing the input tensor according to the pooling area size k and the pooling criterion);

[0050] (3) If the first layer is a fully connected layer, the output is

[0051] al=σ(zl)=σ(Wlal-1+bl);

[0052] (4) For the output layer L:

[0053] aL=σ(ZL)=σ(WLaL_1+bL);

[0054] Among them, the superscript represents the number of layers, W represents the convolution kernel, b represents the bias, and σ is the activation function ReLU.

[0055] In this embodiment, the feature diversity synthesis module uses a generative adversarial network (GANs) algorithm to generate diverse facial features.

[0056] In this embodiment, the generative adversarial network (GANs) algorithm includes a generator and a discriminator;

[0057] Both the generator and the discriminator are composed of convolutional neural networks. The convolution kernel size of the input convolution layer of the convolutional neural network is 2×2, and the convolution kernel size of the remaining convolution layers is 1×3;

[0058] The generator is used to generate handwritten characters in the air with variable length, and its input is the mean vector of each class mixed with noise;

[0059] The discriminator is used to conduct adversarial training with the generator to make the generated variable-length aerial handwritten characters closer to the real data;

[0060] The discriminator includes a global average pooling layer GAP, which is used to convert feature maps of different sizes generated by the discriminator after convolution operation into vectors of equal length.

[0061] In this embodiment, the generator includes a first generation network module and a second generation network module, both of which are residual structures. The generator also includes a third generation network module and a fourth generation network module, and the third network module and the fourth network module are used for feature extraction. The discriminator includes a first discriminant network module and a second discriminant network module, both of which are residual structures. The discriminator also includes a third discriminant network module, and the third discriminant network module is used for feature extraction. The first discriminant network module, the second discriminant network module, and the third discriminant network module all include a Padding layer, and the Padding layer is used to align the length of the feature map after the input and convolution operation. The discriminator also includes a fully connected layer, and the fully connected layer is used to map the "distributed feature representation" learned by the generative adversarial network model to the sample labeling space. The discriminator also includes a binary classification layer, and the loss function of the discriminator is a binary cross entropy loss function BCE. The input of the discriminator is a real handwritten character sample in the air and a virtual handwritten character sample in the air generated by the generator.

[0062] In this embodiment, the data acquisition and preprocessing module includes a signal conditioning module, which amplifies, filters, linearizes, and processes the electrical signal output by the sensor to ensure the accuracy and stability of the signal and meet the needs of subsequent data acquisition and processing. This module usually includes components such as amplifiers, filters, and analog-to-digital converters. The signal conditioning module is coupled to a data cleaning and conversion module to remove duplicate data, correct erroneous data, fill in missing values, etc., to ensure the integrity and consistency of the data. At the same time, the data is formatted and the units are unified as needed. The data cleaning and conversion module is coupled to a data integration module to merge and integrate data from different data sources to form a unified data set. This includes entity recognition, redundancy and correlation analysis, tuple duplication, detection and processing of data value conflicts, data conversion, data reduction, feature selection, and feature extraction. The data integration module is coupled to a data storage and management module to store the preprocessed data in an appropriate medium, such as a local hard disk, cloud storage, etc., for subsequent analysis and processing. At the same time, it is necessary to establish an effective data management mechanism to ensure the security and accessibility of the data.

[0063] In this embodiment, the signal conditioning module includes a signal amplification unit, which amplifies the weak signal output by the sensor to reach the amplitude range required by the subsequent processing circuit. The signal amplification unit is coupled to a signal filtering unit, which filters the signal, removes unnecessary frequency components, and retains useful signal information. The signal filtering unit is coupled to a signal isolation unit, which realizes isolated transmission of the signal, cuts off ground loop interference, and protects the subsequent processing circuit from external interference. The signal isolation unit is coupled to a signal conversion unit, which converts the signal from one form to another, such as conversion from analog signal to digital signal. The signal conversion unit is coupled to a signal linearity unit, which linearizes the nonlinear signal to make it more in line with the requirements of subsequent processing. The signal linearity unit is coupled to a signal protection unit, which performs overload protection, short circuit protection, etc. on the signal to prevent damage to the signal conditioning module and the subsequent processing circuit due to abnormal conditions. The signal protection unit is coupled to a signal distribution unit, which distributes the signal to multiple output channels for use by multiple subsequent processing circuits.

[0064] In this embodiment, the signal amplification unit includes an input unit, which receives an external input signal and performs preliminary amplification. The differential amplifier can amplify the voltage difference between the two input terminals, suppress the common mode signal, and improve the signal-to-noise ratio of the signal. The input unit is coupled with a gain unit to further amplify the signal amplified by the input stage and provide a higher voltage gain. The gain stage usually adopts a multi-stage common source amplifier cascade to achieve step-by-step amplification of the signal. The gain unit is coupled with an output unit to convert the signal amplified by the gain stage into a voltage or current signal suitable for driving the load. The output stage needs to have a lower output impedance and a higher output power capability to ensure that it can drive various types of loads. The output unit is coupled with a bias circuit unit to provide a stable operating point for the differential amplifier to ensure that the amplifier maintains stable performance under normal working conditions. The bias circuit unit is coupled with a feedback network unit to feed back part of the output signal to the input terminal to control the gain and frequency response characteristics of the amplifier. Negative feedback is one of the basic working principles of the operational amplifier, which can improve the stability and linearity of the amplifier.

[0065] In this embodiment, the feedback network unit is composed of a resistor and a capacitor.

[0066] The above shows and describes the basic principles, main features and advantages of the present invention. Technical personnel in this industry should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions are only preferred examples of the present invention and are not used to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, which fall within the scope of the present invention to be protected. The scope of protection of the present invention is defined by the attached claims and their equivalents.

Claims

1. A real-time interactive face feature diversified synthesis adaptive digital human system, characterized by: It includes a data acquisition and preprocessing module, which is coupled to a feature diversification synthesis module, which is coupled to a privacy protection and data encryption module, which is coupled to an intelligent interaction and processing module, and which is coupled to an adaptability and scalability module.

2. The face feature diversified synthesis adaptive digital human system for real-time interaction according to claim 1 is characterized by: The data acquisition and preprocessing module uses a convolutional neural network algorithm to perform face recognition and feature extraction.

3. The face feature diversified synthesis adaptive digital human system for real-time interaction according to claim 2 is characterized by: The convolutional neural network algorithm is as follows: the input layer: input a digital face image, the number of CNN model layers L and the types of all hidden layers; the convolution layer defines the size K of the convolution kernel, the dimension F of the convolution kernel matrix, the padding size P, and the stride S; the pooling layer defines the pooling area size k and the pooling standard; the fully connected layer defines the activation function and the number of neurons in each layer; the output layer: the output value of the CNN model is aL; A. Filling the edge of the digital face image according to the padding size P of the input layer to obtain an input tensor al; B. Initialize the parameters W and b of all the hidden layers; C. for l = 2 to L-1: (1) If the first layer is the convolution layer, the side output is a L =σ(z l )=σ(a l-1 *W 1 +b 1 ); (2) If the first layer is the pooling layer, the side output is a L =pool(a l-1 )(pool refers to the process of reducing the input tensor according to the pooling area size k and the pooling criterion); (3) If the first layer is the fully connected layer, the output is al=σ(zl)=σ(Wlal-1+bl); (4) For the output layer L: aL=σ(ZL)=σ(WLaL_1+bL); Among them, the superscript represents the number of layers, W represents the convolution kernel, b represents the bias, and σ is the activation function ReLU.

4. The real-time interactive face feature diversified synthesis adaptive digital human system according to claim 1, characterized in that: The feature diversity synthesis module uses a generative adversarial network (GANs) algorithm to generate diverse facial features.

5. The real-time interactive face feature diversified synthesis adaptive digital human system according to claim 5, characterized in that: The generative adversarial network (GANs) algorithm includes a generator and a discriminator; The generator and the discriminator are both composed of a convolutional neural network, the convolution kernel size of the input convolution layer of the convolutional neural network is 2×2, and the convolution kernel size of the remaining convolution layers is 1×3; The generator is used to generate handwritten characters in the air with variable length, and its input is the mean vector of each class mixed with noise; The discriminator is used to perform adversarial training with the generator, so as to make the generated variable-length aerial handwritten characters closer to real data; The discriminator includes a global average pooling layer GAP, and the average pooling layer GAP is used to convert feature maps of different sizes generated after the discriminator performs a convolution operation into vectors of equal length.

6. The face feature diversified synthesis adaptive digital human system for real-time interaction according to claim 5, characterized in that: The generator includes a first generation network module and a second generation network module, both of which are residual structures. The generator also includes a third generation network module and a fourth generation network module, and the third network module and the fourth network module are used for feature extraction. The discriminator includes a first discriminant network module and a second discriminant network module, both of which are residual structures. The discriminator also includes a third discriminant network module, and the third discriminant network module is used for feature extraction. The first discriminant network module, the second discriminant network module and the third discriminant network module all include a Padding layer, and the Padding layer is used to align the length of the feature map after the input and convolution operation. The discriminator also includes a fully connected layer, and the fully connected layer is used to map the "distributed feature representation" learned by the generative adversarial network model to the sample labeling space. The discriminator also includes a binary classification layer, and the loss function of the discriminator is a binary cross entropy loss function BCE. The input of the discriminator is a real air handwriting character sample and a virtual air handwriting character sample generated by the generator.

7. The real-time interactive face feature diversified synthesis adaptive digital human system according to claim 1, characterized in that: The data acquisition and preprocessing module includes a signal conditioning module, the signal conditioning module is coupled to a data cleaning and conversion module, the data cleaning and conversion module is coupled to a data integration module, and the data integration module is coupled to a data storage and management module.

8. The face feature diversified synthesis adaptive digital human system for real-time interaction according to claim 7, characterized in that: The signal conditioning module includes a signal amplifying unit, the signal amplifying unit is coupled to a signal filtering unit, the signal filtering unit is coupled to a signal isolation unit, the signal isolation unit is coupled to a signal conversion unit, the signal conversion unit is coupled to a signal linear unit, the signal linear unit is coupled to a signal protection unit, and the signal protection unit is coupled to a signal distribution unit.

9. The face feature diversified synthesis adaptive digital human system for real-time interaction according to claim 8, characterized in that: The signal amplifying unit comprises an input unit, the input unit is coupled to a gain unit, the gain unit is coupled to an output unit, the output unit is coupled to a bias circuit unit, and the bias circuit unit is coupled to a feedback network unit.

10. The face feature diversified synthesis adaptive digital human system for real-time interaction according to claim 9, characterized in that: The feedback network unit is composed of a resistor and a capacitor.