Method and system for generating digital human video
By obtaining user behavior data and sentiment analysis, combining adversarial networks and fuzzy logic control algorithms, personalized digital person images and dialogue content are generated, and digital person videos are generated through distributed rendering, the shortcomings of existing digital person generation methods in high-quality generation and dynamic performance are solved, and a higher user experience and interaction depth are achieved.
Patent Information
- Application Number
- CN202411679883.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-22
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-11-22
AI Technical Summary
The existing digital life generation methods require a large amount of labeled data and computing resources when generating high-quality digital life. The dynamic expression and emotional expression are insufficient, and it is difficult to adapt to different scenarios and individual differences.
By obtaining user behavior data, extracting user preferences and interest tags, generating user feature vectors, and inputting them into the adversarial network to generate initial digital human image. Combining emotion analysis and fuzzy logic control algorithms, digital human expressions, tone and body language are dynamically adjusted to generate personalized digital human images and dialogue content, and digital human videos are generated through distributed rendering.
The generation of personalized digital person images and dialogue content has been realized, which has improved the realism and emotional connection between digital persons and users, and enhanced the immersion and experience of users.
Smart Images

Figure CN119180895B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of digital human technology, and in particular to a method and system for generating digital human videos. Background Art
[0002] With the rapid development of digital technology, the application of digital humans (also known as virtual humans or digital twins) in multiple fields has gradually become popular, including entertainment, education, medical care and customer service. Digital humans can simulate the appearance, behavior and emotions of real humans, making the interaction between people and computers more natural and efficient. This trend not only provides companies with innovative business models, but also brings a new experience to users. Especially in the context of the increasing maturity of virtual reality and augmented reality technologies, the use of digital humans has become more and more widespread, and its potential in digital marketing, social media and online education has been valued by more and more organizations.
[0003] Although there are many methods for generating digital humans, such as deep learning-based generative adversarial networks (GANs), 3D modeling technology, and expression capture technology, these methods still face many challenges. First, existing technologies often require a large amount of labeled data and computing resources when generating high-quality digital humans, which limits their application in small businesses or individual developers. Secondly, many existing methods lack realism in dynamic performance and emotional expression, and it is difficult to effectively simulate complex human emotions and micro-expressions. In addition, the existing systems lack flexibility in adapting to different scenarios and individual differences, making it difficult for digital humans to meet specific user needs. These shortcomings limit the application effect and user experience of digital humans, and further technological innovation and system improvement are urgently needed. Summary of the invention
[0004] In view of the deficiencies of the prior art, the present invention provides a method and system for generating digital human videos, which solve the problems of the above-mentioned background technology.
[0005] To achieve the above objectives, the present invention is implemented through the following technical solutions: a method for generating digital human videos, comprising the following steps: S1. obtaining user behavior data, preprocessing the user behavior data, extracting user preferences, interest tags and interaction records, and generating user feature vectors through a deep learning model; S2. inputting the user feature vectors into an adversarial network to generate an initial digital human image, obtaining user emotional input data, extracting user emotional state characteristics through an emotional analysis model, setting the emotional response of the digital human, analyzing the user's real-time interaction data, and obtaining a user interaction matching index; S3. based on the user interaction matching index, establishing a feedback mechanism through a fuzzy logic control algorithm, adjusting the digital human's expression details, tone and body language, and generating a personalized digital human image; S4. based on the user feature vector and the interaction matching index, generating personalized dialogue content through a natural language generation model, and generating digital human videos in combination with distributed rendering.
[0006] Furthermore, the specific process of preprocessing the user behavior data and extracting user preferences, interest tags and interaction records is as follows: normalize the user's historical browsing records, click behaviors and dwell time, analyze the user behavior data through clustering algorithms, extract user preference characteristics and interest tags, and record the user's behavior frequency in different interaction scenarios.
[0007] Furthermore, the specific process of generating user feature vectors through the deep learning model is as follows: the preprocessed user behavior data is input into the deep learning model, a multi-layer perceptron is used to extract the user's preference features and interest tags, the output of each layer is nonlinearly transformed through the activation function, and finally the user feature vector is generated through the fully connected layer.
[0008] Furthermore, the specific process of extracting the characteristics of the user's emotional state through the sentiment analysis model and setting the emotional response of the digital human is as follows: input the user's emotional input data into the sentiment analysis model, analyze the user's text, voice or facial expression data through the long short-term memory network, and extract the user's emotional characteristics; classify the emotional state of the digital human based on the extracted emotional characteristics, and set the corresponding emotional response rules to match the user's emotional state.
[0009] Furthermore, the specific process of analyzing the real-time user interaction data and obtaining the user interaction matching index is as follows: obtaining the real-time user interaction data, including click behavior, browsing time, and feedback response, to form a multidimensional interaction data set, and normalizing different types of interaction data through multidimensional data aggregation; extracting features from the normalized data through principal component analysis, generating main feature vectors, and constructing a weighted interaction matching model. Based on the historical interaction data, the importance of each interaction behavior is analyzed, and different interaction behaviors are given variable weights in combination with collaborative filtering algorithms and dynamic weighting mechanisms. The weight values are updated in real time to adapt to the user's current behavior pattern; a comprehensive scoring formula is established through multi-factor linear regression analysis, the extracted main features and dynamic weights are combined, and the user interaction matching index is calculated to quantify the interaction quality and consistency between the user and the digital human.
[0010] Furthermore, based on the user interaction matching index, a feedback mechanism is established through a fuzzy logic control algorithm to adjust the digital human's expression details, tone and body language. The specific process is as follows: according to the user interaction matching index, the input variables of the fuzzy logic controller are set, the fuzzy sets and rule bases are defined, and the fuzzy rules are constructed. Through the fuzzy reasoning mechanism, the input variables are mapped to the fuzzy sets, and fuzzy operations are performed to generate fuzzy outputs to adjust the digital human's expression details, tone and body language.
[0011] Furthermore, the specific process of generating digital human videos in combination with distributed rendering is as follows: construct a digital human model and scene environment, divide the rendering task into multiple subtasks through a distributed rendering system, and assign them to different computing nodes for parallel processing; during the rendering process, dynamically adjust the digital human's facial expression details, tone and body language in combination with the user's interactive matching index; each computing node independently executes the rendering task according to its own processing power and current load, and sends the rendering results back to the central processing unit after completion; the central processing unit synthesizes the received rendering results to generate a digital human video.
[0012] The system for generating digital human videos comprises the following modules: a data acquisition module, an image generation module, an emotion adjustment module and a generation module; the data acquisition module is used to acquire user behavior data, pre-process the user behavior data, extract user preferences, interest tags and interaction records, and generate user feature vectors through a deep learning model; the image generation module is used to input the user feature vector into an adversarial network, generate an initial digital human image, acquire user emotion input data, extract user emotional state characteristics through an emotion analysis model, set the emotional response of the digital human, analyze the user's real-time interaction data, and acquire the user interaction matching index; the emotion adjustment module is used to establish a feedback mechanism based on the user interaction matching index through a fuzzy logic control algorithm, adjust the digital human's expression details, tone and body language, and generate a personalized digital human image; the generation module is used to generate personalized dialogue content through a natural language generation model based on the user feature vector and the interaction matching index, and generate digital human videos in combination with distributed rendering.
[0013] The present invention has the following beneficial effects:
[0014] (1) This method for generating digital human videos can accurately extract user preferences and interest tags by preprocessing user behavior data, thereby laying the foundation for subsequent personalized services. The generated user feature vector can fully reflect the user's behavior pattern and provide an important basis for the construction of the digital human image. By inputting the user feature vector into the adversarial network, the initial digital human image can be generated according to the user's personalized needs, and the user's emotional state can be obtained by combining sentiment analysis. This process ensures that the emotional response of the digital human is consistent with the user's emotions, which helps to enhance the realism and emotional connection of the interaction and increase the user's immersion.
[0015] (2) The system for generating digital human videos adjusts the digital human's expression, tone of voice and body language according to the user interaction matching index through a fuzzy logic control algorithm to achieve the generation of personalized digital human images. This mechanism improves the digital human's sensitivity to user emotions and interactions, enabling it to adapt to user needs more naturally, thereby enhancing user experience and satisfaction. By combining user feature vectors, interaction matching indexes and emotional states to generate personalized conversation content, and generating digital human videos through distributed rendering, it is possible to ensure that each interaction is unique and relevant. This dynamic generation capability not only improves the interactive quality of digital human videos, but also increases the depth of interaction between users and digital humans, ultimately achieving higher user engagement.
[0016] Of course, any product implementing the present invention does not necessarily need to achieve all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1The present invention is a flow chart of a method for generating digital human video.
[0018] Figure 2 The figure is a flow chart of the system for generating digital human video according to the present invention. DETAILED DESCRIPTION
[0019] The embodiments of the present application solve the problems of existing digital humans in terms of insufficient realism in dynamic performance and emotional expression, dependence on large amounts of annotated data and computing resources, and insufficient flexibility in adapting to different scenarios and individual differences through a method and system for generating digital human videos, thereby improving the application effect and user experience of digital humans.
[0020] The overall idea of the solution in the embodiments of this application is as follows:
[0021] Obtain user behavior data, preprocess the user behavior data, extract user preferences, interest tags and interaction records, and generate user feature vectors through deep learning models.
[0022] Input the user feature vector into the adversarial network to generate the initial digital human image, obtain the user's emotional input data, extract the user's emotional state characteristics through the sentiment analysis model, set the digital human's emotional response, analyze the user's real-time interaction data, and obtain the user interaction matching index.
[0023] Based on the user interaction matching index, a feedback mechanism is established through the fuzzy logic control algorithm to adjust the digital human's facial expressions, tone and body language to generate a personalized digital human image.
[0024] Based on user feature vectors and interaction matching index, personalized conversation content is generated through the natural language generation model, and digital human videos are generated in combination with distributed rendering.
[0025] See also Figure 1 , an embodiment of the present invention provides a technical solution: a method for generating a digital human video, comprising the following steps: S1. obtaining user behavior data, preprocessing the user behavior data, extracting user preferences, interest tags and interaction records, and generating a user feature vector through a deep learning model; S2. inputting the user feature vector into an adversarial network to generate an initial digital human image, obtaining user emotional input data, extracting user emotional state characteristics through an emotional analysis model, setting the emotional response of the digital human, analyzing the user's real-time interaction data, and obtaining a user interaction matching index; S3. based on the user interaction matching index, establishing a feedback mechanism through a fuzzy logic control algorithm, adjusting the digital human's expression details, tone and body language, and generating a personalized digital human image; S4. based on the user feature vector and the interaction matching index, generating personalized dialogue content through a natural language generation model, and generating a digital human video in combination with distributed rendering.
[0026] In this embodiment, in step S1, the user's behavior data on the application or platform is collected. By monitoring the user's clicks, browsing time and feedback, a preliminary understanding of the user's interests and preferences can be obtained. User behavior data preprocessing: in the acquired data, irrelevant information is removed and the data is cleaned to ensure the accuracy of subsequent analysis. Extracting user preferences, interest tags and interaction records: extracting useful information from the preprocessed data, such as the user's interest tags (such as hobbies, shopping preferences, etc.) and interaction history (such as past conversations, feedback, etc.). User feature vector: using a deep learning model (such as a convolutional neural network or a recurrent neural network) to convert the extracted user information into a high-dimensional vector. This feature vector contains various characteristics of the user, which is helpful for the generation of subsequent personalized services. In step S2, this step inputs the user feature vector into the deep learning model to generate a digital human image that matches the user's needs. Adversarial network: refers to the generative adversarial network (GAN), which is a deep learning model composed of a generator and a discriminator. The generator is responsible for generating new digital human images, and the discriminator judges the authenticity of these images. Through this adversarial mechanism, the quality of the generated digital human images will continue to improve. User emotional input data: At this stage, the system also needs to obtain the user's emotional data (such as anger, happiness, sadness, etc.) in order to set the corresponding emotional response for the digital human. Emotional analysis model: Use natural language processing technology to analyze the user's input (such as text or voice) and extract the characteristics of the user's emotional state. These emotional characteristics will be used to adjust the performance of the digital human. User interaction matching index: A quantitative indicator generated based on real-time interaction data, which represents the quality of interaction between the user and the digital human. In step S3, this step uses the user interaction matching index to adjust the performance of the digital human to improve the user experience. Fuzzy logic control algorithm: Use fuzzy logic (a method of dealing with uncertainty) to achieve intelligent adjustment of the performance of the digital human. The system will dynamically adjust the digital human's expression, tone and body language based on the user's feedback information to make it more natural and consistent with the user's emotional state. Feedback mechanism: By continuously collecting user feedback information, the performance of the digital human is adjusted in real time to enhance the realism of the interaction. In step S4, the system will generate personalized conversation content based on the user's feature vector, interaction matching index and emotional state. Natural language generation model: This model converts the user's characteristics and emotions into natural language text, allowing the digital human to have a more personalized conversation. Distributed rendering: When generating digital human videos, distributed computing resources are used to quickly process and render video content to improve efficiency and quality. Distributed rendering can distribute complex rendering tasks to multiple computers for processing, thereby speeding up video generation. Deep learning model: A machine learning method based on artificial neural networks that automatically learns and extracts data features through a multi-layer network structure.Generative Adversarial Network (GAN): A model consisting of two neural networks, the generator and the discriminator, which improves the quality of generated data through adversarial training and is suitable for a variety of generation tasks such as images and audio. Fuzzy Logic: A method of dealing with uncertainty, which uses fuzzy sets and rule definitions to make inferences and decisions, enabling the system to operate effectively in complex and uncertain environments. Natural Language Generation (NLG): A technology that automatically generates natural language text through computer programs, which is widely used in fields such as chatbots and virtual assistants. Distributed Rendering: A technology that distributes rendering tasks to multiple computers for parallel processing, usually used to improve computing efficiency and handle complex scenes.
[0027] Specifically, the specific process of preprocessing user behavior data and extracting user preferences, interest tags and interaction records is as follows: normalize the user's historical browsing records, click behaviors and dwell time, analyze the user behavior data through clustering algorithms, extract user preference characteristics and interest tags, and record the user's behavior frequency in different interaction scenarios.
[0028] In this implementation scheme, the user's historical behavior data is collected, including: Historical browsing history: a list of web pages or content that the user has visited, usually including the visit time and dwell time of each page. Click behavior: the links or buttons that the user clicks on the interface, these behaviors can reflect the user's interests and needs. Dwelling time: the time the user stays on a specific content or page, indicating the attractiveness of the content to the user. The collected user behavior data is normalized to ensure the comparability between different features. Cluster analysis uses a clustering algorithm to analyze user behavior data and extract user preference features and interest tags from it. The main purpose of the clustering algorithm is to classify similar behaviors to discover potential user preferences. K-means clustering: users are divided into K different categories by calculating the mean of user behavior features. The algorithm iteratively optimizes the center point of each category and finally forms a set of user features. After the cluster analysis is completed, according to the cluster to which each user belongs, their preference features and interest tags are extracted. For example: Preference features: based on the user's behavior performance in a specific category, it is concluded that the user is more inclined to browse which types of content (such as entertainment, education, technology, etc.). Interest tags: Assign corresponding interest tags to each user (such as game lovers, fashion followers), which will be used to generate personalized services in the future. Count the frequency of user behaviors in different interaction scenarios, such as the number of times a user participates in a certain type of activity, the preferred interaction method, etc. This helps to understand the user's performance in different situations and provide a more accurate user portrait for subsequent models.
[0029] Specifically, the specific process of generating user feature vectors through deep learning models is as follows: the preprocessed user behavior data is input into the deep learning model, a multi-layer perceptron is used to extract user preference features and interest tags, the output of each layer is nonlinearly transformed through an activation function, and finally a user feature vector is generated through a fully connected layer.
[0030] In this embodiment, the preprocessed user behavior data (including user preference features and interest tags) is used as input and passed to the deep learning model. At this point, the data has been standardized and processed to ensure its quality and availability. A multi-layer perceptron (MLP) structure is used for feature extraction. MLP is a typical feedforward neural network consisting of multiple neurons (nodes), usually including an input layer, multiple hidden layers, and an output layer. MLP can capture complex nonlinear relationships through a multi-layer network structure. Input layer: The feature vector that receives the user behavior data is usually a one-dimensional array, and each element represents a certain behavior or preference of the user. Hidden layer: The neurons in the middle layer are nonlinearly transformed through the activation function, so that the network can learn complex feature relationships. Multiple hidden layers can be set to enhance the expressive power of the model. In each hidden layer, the activation function is applied to perform a nonlinear transformation on the output, and the output is 0 for negative values and remains unchanged for positive values. Through the nonlinear transformation of the activation function, the model can capture complex patterns and features in the data, thereby improving the learning ability. After passing through multiple hidden layers, the output results will be sent to the fully connected layer, which is connected to all nodes in the previous layer. This layer is responsible for integrating the features extracted by the previous layers and generating the final user feature vector. The fully connected layer usually uses linear transformation and combines an activation function to enhance nonlinear capabilities. After being processed by the fully connected layer, the output generated by the model is the user feature vector, which is usually a high-dimensional array that can effectively represent the user's preferences and behavioral characteristics. The user feature vector provides an important foundation for the subsequent digital human generation and personalized services.
[0031] Specifically, the specific process of extracting the characteristics of the user's emotional state through the sentiment analysis model and setting the emotional response of the digital human is as follows: input the user's emotional input data into the sentiment analysis model, analyze the user's text, voice or facial expression data through the long short-term memory network, and extract the user's emotional characteristics; classify the emotional state of the digital human based on the extracted emotional characteristics, and set the corresponding emotional response rules to match the user's emotional state.
[0032] In this implementation, the user's emotional input data comes from multiple channels, including: Text data: text information entered by the user when talking to the digital person. Voice data: what the user says when interacting with the digital person through a voice assistant. Facial expression data: real-time facial expressions of the user are captured by the camera to reflect their emotional state. These data will be used as input to the sentiment analysis model to help the system understand the user's current emotions. Construction of sentiment analysis model: Sentiment analysis models are usually based on long short-term memory networks (LSTM), which are recurrent neural networks suitable for processing time series data. The specific steps include: Text analysis: For text input, LSTM encodes each word and identifies the emotional tendency in the sentence. For example, if the user enters "I am very happy", the model can detect the emotional word "happy" and evaluate its emotional intensity. Voice analysis: For voice input, the model analyzes features such as pitch, speaking speed, and intonation to identify emotional states. For example, a high pitch and fast speaking speed may indicate that the user is excited or happy, while a low pitch and slow speaking speed may indicate frustration or sadness. Facial expression analysis: Through computer vision technology, the captured facial image is input into the LSTM model to analyze facial features (such as smile, frown) to identify emotional states. For example, a smile may be classified as "happy", while a frown may be identified as "confused" or "angry". Feature extraction: After being processed by the LSTM model, the system will extract the user's emotional features. These features are usually a vector containing information about the type of emotion (such as happiness, sadness, anger, etc.) and the intensity of the emotion (such as mild, moderate, severe). Emotional state classification: Using the extracted emotional features, the model classifies the user's emotional state. The specific classification steps can be: Setting categories: defining preset emotional categories, such as "happy", "sad", "angry", "surprised", etc. Applying classification algorithms: using machine learning classifiers (such as support vector machines or decision trees) to classify emotional states based on the extracted features. This process can be supervised by labeled training data, so that the model can accurately classify under different emotional states. Setting emotional response rules: based on the classification results, emotional response rules are set for the digital human. The rules include: Emotional response: how the digital human should react when the user shows a certain emotion. For example: If the user shows happiness, the digital human may smile and respond with a pleasant tone. If the user shows sadness, the digital human may adopt a gentle tone and provide comforting words. Dynamic adjustment: During the interaction process, the digital human will monitor the changes in the user's emotions in real time and dynamically adjust its response. For example, if the user changes from happy to angry, the digital human will quickly adapt and adopt a calm and soothing tone. Matching the user's emotional state: The digital human will match the user's emotional state based on the classified emotional state and the set response rules.This emotional matching makes the digital human's performance more natural and humane, which can enhance the user's interactive experience and strengthen emotional resonance.
[0033] Specifically, the specific process of analyzing user real-time interaction data and obtaining the user interaction matching index is as follows: obtain user real-time interaction data, including click behavior, browsing time, and feedback response, to form a multidimensional interaction data set, and normalize different types of interaction data through multidimensional data aggregation; extract features from the normalized data through principal component analysis, generate main feature vectors, and construct a weighted interaction matching model. Based on historical interaction data, the importance of each interaction behavior is analyzed, and combined with collaborative filtering algorithms and dynamic weighting mechanisms, different interaction behaviors are given variable weights, and the weight values are updated in real time to adapt to the user's current behavior pattern; a comprehensive scoring formula is established through multi-factor linear regression analysis, the extracted main features and dynamic weights are combined, and the user interaction matching index is calculated to quantify the interaction quality and consistency between the user and the digital human.
[0034] In this implementation scheme, real-time user interaction data is obtained: real-time interaction data includes: click behavior: the frequency and duration of users clicking on the information or buttons provided by the digital human. Browsing time: the total time the user stays when interacting with the digital human. Feedback response: the feedback given by the user during the interaction process (such as ratings, comments, etc.). These data will be recorded and aggregated into a multidimensional interaction data set. Through multidimensional data aggregation, different types of interaction data are normalized to eliminate the impact of dimensions. The normalization formula can be expressed as: in: is the original data value. and are the minimum and maximum values of the feature, respectively. The normalized values ensure that all features are on the same scale for easy comparison. Principal component analysis is performed on the normalized data to extract the main features. PCA is implemented by calculating the data covariance matrix and extracting its eigenvalues and eigenvectors. The main steps are as follows: Calculate the covariance matrix , ;in is the number of samples, is the sample mean. Compute the eigenvalues and eigenvectors: ;in is the characteristic value, is the eigenvector. Select the largest The eigenvectors corresponding to the eigenvalues form the main eigenvector . Build a weighted interaction matching model: Analyze the importance of each interaction behavior based on historical interaction data, and combine collaborative filtering algorithms and dynamic weighting mechanisms to assign variable weights to each interaction behavior. ;in: For the Interaction behavior in time The weight of . is the weight of the previous time step. Score the importance of the behavior at the current time step (such as click-through rate, browsing time, etc.). and is the coefficient for weight update, and A comprehensive scoring formula is established using multi-factor linear regression analysis, and the user interaction matching index is calculated by combining the extracted main features and dynamic weights. The specific calculation formula is as follows: ;in: It is the user interaction matching index. For the Interaction behavior in time The weight of . For the The main characteristic values of the interactive behavior (such as normalized click behavior, browsing time, etc.). is the total number of interactive behaviors. Through the above formula, the system will calculate the user's interactive matching index ,This index quantifies the quality and consistency of the interaction between the user and the ,digital human. The higher value indicates that the interaction between the user and ,the digital human is smoother and more personalized.
[0035] Specifically, based on the user interaction matching index, a feedback mechanism is established through the fuzzy logic control algorithm to adjust the digital human's expression details, tone and body language. The specific process is as follows: set the input variables of the fuzzy logic controller according to the user interaction matching index, define the fuzzy set and rule base, construct fuzzy rules, map the input variables to the fuzzy set through the fuzzy reasoning mechanism, perform fuzzy operations, generate fuzzy outputs, and adjust the digital human's expression details, tone and body language.
[0036] In this implementation, the input variable of the fuzzy logic controller is set: the user interaction matching index (MDI) is used as the key input variable of the fuzzy logic controller to represent the quality and consistency of the interaction between the user and the digital human. In general, the range of MDI can be divided into several levels, such as "low", "medium" and "high". These levels will be used to define the fuzzy set. ;in, , , and is the boundary parameter that determines the fuzzy set, is the value of the user interaction matching index, Indicates that the value is in the fuzzy set According to the requirements of fuzzy logic control, a rule base is constructed to map input variables to output variables. These rules are generally defined in the form of "if...then...", as shown below: If the interaction matching index is "low", then the expression details are "angry", the tone is "deep", and the body language is "stiff". If the interaction matching index is "medium", then the expression details are "neutral", the tone is "normal", and the body language is "natural". If the interaction matching index is "high", then the expression details are "smiling", the tone is "brisk", and the body language is "relaxed". Constructing fuzzy rules Use the defined rule base to establish fuzzy rules. Each rule associates the input variables and the corresponding output variables to describe the output response under different input conditions. Define fuzzy sets For each input variable, such as the user interaction matching index, define the corresponding fuzzy set. The elements in the fuzzy set can be represented by the membership function. The fuzzy reasoning mechanism is the process of mapping the input variables (such as the user interaction matching index) to the fuzzy set and obtaining the fuzzy output through fuzzy operation. The reasoning steps are as follows: The membership of the input variables is calculated to obtain the membership value of each fuzzy set. According to the membership value, the fuzzy rules are applied to determine the membership of the corresponding output variables. Through the aggregation operation, the outputs of all rules are synthesized to form a fuzzy output. Fuzzy operation and output generation, through fuzzy operation, the outputs of each fuzzy set are synthesized to obtain the fuzzy output. Then, the fuzzy output is converted into a clear control signal using defuzzification technology. Common defuzzification methods include: ;in, is the final output value, is the specific value of the fuzzy output, is the corresponding membership degree. Through the above process, the generated control signal will be used to adjust the digital human's expression details, tone and body language. For example, if the final output points to a "smiling" expression, a "brisk" tone and a "relaxed" body language, the digital human will make corresponding adjustments based on these instructions. This will enhance the interactive experience with the user.
[0037] Specifically, the specific process of generating digital human videos in combination with distributed rendering is as follows: construct a digital human model and scene environment, divide the rendering task into multiple subtasks through a distributed rendering system, and assign them to different computing nodes for parallel processing; during the rendering process, dynamically adjust the digital human's facial expression details, tone and body language in combination with the user's interactive matching index; each computing node independently executes the rendering task based on its own processing power and current load, and sends the rendering results back to the central processing unit after completion; the central processing unit synthesizes the received rendering results to generate a digital human video.
[0038] In this implementation scheme, constructing a digital human model and scene environment: before generating a digital human video, it is first necessary to construct a three-dimensional model of the digital human and its scene environment. This includes: Digital human model: designing the appearance features of the digital human, such as facial expressions, body structure, clothing, etc. This is usually achieved using three-dimensional modeling software and is personalized according to the user's preferences and emotional state. Scene environment: constructing the background environment for the digital human's activities, including indoor and outdoor scenes, props and light source settings, etc. The complexity of the scene will affect the amount of computation required for rendering. Construction of a distributed rendering system: in order to improve rendering efficiency, a distributed rendering system is constructed, which divides the rendering task into multiple subtasks and assigns them to different computing nodes (such as multiple servers or computers) for parallel processing. The main steps are as follows: Task division: according to the complexity of the digital human model and the scene, the rendering task is subdivided into multiple small tasks, for example, rendering details from different perspectives and at different levels. Task scheduling: the divided subtasks are assigned to different computing nodes through a load balancing algorithm, which can be local computers, cloud servers, etc. Dynamically adjust the performance of the digital human: During the rendering process, it is necessary to dynamically adjust the expression details, tone and body language of the digital human in combination with the user's interaction matching index (MDI). According to different values of MDI, the following adjustments can be made during the rendering process: Expression details: When MDI is high, more delicate expression details can be rendered, such as smiles, eye contact, etc.; when MDI is low, simple expressions may be selected. Tone: The speech synthesis module will adjust the tone according to MDI, making the digital human more vivid and natural when the interaction matching degree is high. Body language: Through the animation control system, the body language of the digital human will also be adjusted in real time to adapt to the user's emotional state and behavior. Parallel processing of rendering tasks: Each computing node independently executes the rendering tasks assigned to it according to its own processing power and current load. This process includes: Data processing: Each node receives specific rendering task data and performs real-time image rendering, special effects application, etc. Rendering output: After completing the task, the computing node sends the rendering results (such as image frames) back to the central processing unit in the form of data packets. The data transmission here can use high-speed network protocols to ensure timely interaction of rendering data. Synthesis of the central processing unit: The central processing unit is responsible for receiving the rendering results from each computing node and synthesizing them. This process mainly includes: Data reception: receiving the rendering results from each computing node, which may be image frames, depth information, etc. Image synthesis: synthesizing the received images, usually including: Frame synthesis: synthesizing multiple frames of images into animations in time sequence. Post-processing: applying light and shadow effects, filters, special effects, etc. to enhance the overall visual effect of the video. Video generation: The final digital human video will be encoded and output in a specific format for easy storage and playback.
[0039] The system for generating digital human videos comprises the following modules: a data acquisition module, an image generation module, an emotion adjustment module and a generation module; the data acquisition module is used to acquire user behavior data, pre-process the user behavior data, extract user preferences, interest tags and interaction records, and generate user feature vectors through a deep learning model; the image generation module is used to input the user feature vector into an adversarial network, generate an initial digital human image, acquire user emotion input data, extract user emotional state characteristics through an emotion analysis model, set the emotional response of the digital human, analyze the user's real-time interaction data, and acquire the user interaction matching index; the emotion adjustment module is used to establish a feedback mechanism based on the user interaction matching index through a fuzzy logic control algorithm, adjust the digital human's expression details, tone and body language, and generate a personalized digital human image; the generation module is used to generate personalized dialogue content through a natural language generation model based on the user feature vector and the interaction matching index, and generate digital human videos in combination with distributed rendering.
[0040] In this implementation scheme, the data acquisition module: Function: Acquire user behavior data: This module collects user behavior data through various channels, including but not limited to user browsing history, click behavior, dwell time, social media interaction, etc. These data are the basis for constructing user features. Historical behavior data preprocessing: Clean and normalize the collected user behavior data, remove noise data, and ensure data quality. Extract user preferences and interest tags: Analyze user behavior patterns through clustering algorithms and other technologies to extract user preference features and interest tags, which can be used to describe users' personalized needs and interests. Generate user feature vectors: Input the processed data into the deep learning model, use structures such as multi-layer perceptron (MLP) to abstract features, and generate user feature vectors. This vector is the basis for subsequent modules to generate personalized digital people. Image generation module: Generate initial digital human image: Input the user feature vector into the adversarial network (such as GAN) to generate a digital human image that meets the user's preferences. This image can dynamically adjust the appearance according to the user's characteristics. Obtain user emotional input data: This module also needs to obtain the user's emotional input data, which may come from multimodal information such as the user's voice, text or facial expressions. Emotional state feature extraction: Use the emotion analysis model (such as long short-term memory network LSTM) to analyze the user's emotional input and extract the user's current emotional state. Set emotional response: Set the emotional response of the digital human according to the extracted emotional features so that it can make corresponding emotional expressions when interacting with the user. Obtain the user interaction matching index: Analyze the user's real-time interaction data, including the user's feedback, click behavior, etc., and calculate the user interaction matching index (MDI) to evaluate the interaction quality between the digital human and the user. Emotion adjustment module: Based on the interaction matching index: This module uses MDI as input to dynamically adjust the performance of the digital human according to the user's interaction. Fuzzy logic control algorithm: Through the fuzzy logic controller, a feedback mechanism is established to map the MDI to the fuzzy set to determine the expression details, tone and body language that the digital human needs to adjust. For example, if the MDI is higher, the expression of the digital human can be richer and more natural. Personalized digital human image generation: The adjusted expression and tone will generate a digital human image that is more in line with the user's emotional state, enhancing the naturalness and attractiveness of the interaction. Generation module: Generate personalized conversation content: Based on user feature vectors and interaction matching index, use natural language generation models (such as the GPT series or other conversation generation algorithms) to generate personalized conversation content, so that digital human communication is more in line with user needs. Combined with distributed rendering: Through distributed rendering technology, the generated digital human image and conversation content are combined to generate high-quality digital human videos. This process can effectively improve the rendering speed and video quality, making the final generated video more smooth and realistic.
[0041] In summary, this application has at least the following effects:
[0042] The method and system for generating digital human videos obtain user behavior data, pre-process the user behavior data, generate user feature vectors, and improve the accuracy of user personalized identification. The user feature vector is input into the adversarial network to generate the initial digital human image, analyze the user's real-time interaction data, obtain the user interaction matching index, realize the dynamic matching of the digital human image and the user's emotion, and enhance the interactive experience. A feedback mechanism is established through the fuzzy logic control algorithm to adjust the digital human's expression details, tone and body language to generate a personalized digital human image, generate personalized dialogue content through the natural language generation model, and generate digital human videos in combination with distributed rendering, achieving high-quality and personalized digital human video output, thereby improving user satisfaction and participation.
[0043] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0044] The present invention is described with reference to flowcharts and / or block diagrams of systems, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0045] These computer program instructions may also be stored in a computer readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0046] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0047] Although the preferred embodiments of the present invention have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0048] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.
Claims
1. A method for generating a digital human video, characterized in that: The following steps are involved: S1. Obtain user behavior data, pre-process the user behavior data, extract user preferences, interest tags and interaction records, and generate user feature vectors through deep learning models; S2. Input the user feature vector into the adversarial network to generate the initial digital human image, obtain the user's emotional input data, extract the user's emotional state characteristics through the sentiment analysis model, set the digital human's emotional response, analyze the user's real-time interaction data, and obtain the user interaction matching index; S3. Based on the user interaction matching index, a feedback mechanism is established through the fuzzy logic control algorithm to adjust the digital human's facial expressions, tone and body language to generate a personalized digital human image; S4. Based on the user feature vector and interaction matching index, the natural language generation model is used to generate personalized conversation content, and the digital human video is generated in combination with distributed rendering; The specific process of establishing a feedback mechanism based on the user interaction matching index through a fuzzy logic control algorithm to adjust the digital human's facial expression details, tone of voice and body language is as follows: According to the user interaction matching index, the input variables of the fuzzy logic controller are set, the fuzzy sets and rule base are defined, and the fuzzy rules are constructed. Through the fuzzy reasoning mechanism, the input variables are mapped to fuzzy sets, and fuzzy operations are performed to generate fuzzy outputs and adjust the digital human's facial expression details, tone and body language.
2. The method for generating a digital human video according to claim 1, characterized in that: The specific process of preprocessing user behavior data and extracting user preferences, interest tags, and interaction records is as follows: The user's historical browsing records, click behaviors and dwell time are normalized, the user behavior data is analyzed through clustering algorithms, the user's preference characteristics and interest tags are extracted, and the user's behavior frequency in different interaction scenarios is recorded.
3. The method for generating a digital human video according to claim 2, characterized in that: The specific process of generating user feature vectors through deep learning models is as follows: The preprocessed user behavior data is input into the deep learning model, and the user's preference features and interest tags are extracted using a multi-layer perceptron. The output of each layer is nonlinearly transformed through an activation function, and finally a user feature vector is generated through a fully connected layer.
4. The method for generating a digital human video according to claim 3, characterized in that: The specific process of extracting the user's emotional state characteristics through the sentiment analysis model and setting the emotional response of the digital human is as follows: Input the user's emotional input data into the sentiment analysis model, analyze the user's text, voice or facial expression data through the long short-term memory network of the sentiment analysis model, and extract the user's emotional characteristics; The emotional state of the digital human is classified based on the extracted emotional features, and corresponding emotional response rules are set to match the user's emotional state.
5. The method for generating a digital human video according to claim 4, characterized in that: The specific process of analyzing real-time user interaction data and obtaining the user interaction matching index is as follows: Obtain real-time user interaction data, including click behavior, browsing time, and feedback response, to form a multi-dimensional interaction data set, and normalize different types of interaction data through multi-dimensional data aggregation; The normalized data is subjected to feature extraction through principal component analysis to generate the main feature vectors and build a weighted interaction matching model. The importance of each interaction behavior is analyzed based on historical interaction data. The collaborative filtering algorithm and dynamic weighting mechanism are combined to assign variable weights to different interaction behaviors and update the weight values in real time to adapt to the user's current behavior pattern. A comprehensive scoring formula is established through multi-factor linear regression analysis, and the extracted main features and dynamic weights are combined to calculate the user interaction matching index to quantify the interaction quality and consistency between users and digital humans.
6. The method for generating a digital human video according to claim 5, characterized in that: The specific process of generating digital human video in combination with distributed rendering is as follows: Build digital human models and scene environments, divide rendering tasks into multiple subtasks through a distributed rendering system, and assign them to different computing nodes for parallel processing; During the rendering process, the digital human’s expression details, tone of voice, and body language are dynamically adjusted based on the user’s interactive matching index; Each computing node independently performs rendering tasks based on its own processing capabilities and current load, and sends the rendering results back to the central processing unit after completion; The central processing unit synthesizes the received rendering results to generate a digital human video.
7. A system for generating digital human videos, applying the method for generating digital human videos according to any one of claims 2 to 6, characterized in that: It includes the following modules: data acquisition module, image generation module, emotion adjustment module, generation module; The data acquisition module is used to acquire user behavior data, pre-process the user behavior data, extract user preferences, interest tags and interaction records, and generate user feature vectors through a deep learning model; The image generation module is used to input the user feature vector into the adversarial network, generate the initial digital human image, obtain the user's emotional input data, extract the user's emotional state characteristics through the emotional analysis model, set the digital human's emotional response, analyze the user's real-time interaction data, and obtain the user interaction matching index; The emotion adjustment module is used to establish a feedback mechanism based on the user interaction matching index through a fuzzy logic control algorithm to adjust the digital human's expression details, tone and body language to generate a personalized digital human image; The generation module is used to generate personalized dialogue content through a natural language generation model based on user feature vectors and interaction matching index, and generate digital human video in combination with distributed rendering; The specific process of establishing a feedback mechanism based on the user interaction matching index through a fuzzy logic control algorithm to adjust the digital human's facial expression details, tone of voice and body language is as follows: According to the user interaction matching index, the input variables of the fuzzy logic controller are set, the fuzzy sets and rule base are defined, and the fuzzy rules are constructed. Through the fuzzy reasoning mechanism, the input variables are mapped to fuzzy sets, and fuzzy operations are performed to generate fuzzy outputs and adjust the digital human's facial expression details, tone and body language.
Citation Information
Patent Citations
Voice interaction method based on emotion engine technology, intelligent terminal and storage medium
CN111368609A
Emotion symbiosis communication system
CN118379403A