Digital human interaction video stream determination method and device, equipment, medium and product

By extracting and smoothing the emotional features of the target object's video stream, generating emotional frame images and inserting them into the video stream, the problem of inaccurate emotion recognition in the existing technology is solved, and the accuracy and real-time performance of digital human interaction are improved.

CN120636388APending Publication Date: 2025-09-12CHINA MOBILE SHANGHAI ICT CO LTD +2
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510766327.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing emotion recognition technology has low accuracy in identifying human emotions and is unable to accurately identify rapid changes in emotions, resulting in inaccurate determination of the facial emotions of digital humans in digital human interaction video streams.

Method used

By collecting the video stream of the target object, the current emotional features in the current frame image are determined, and feature smoothing is performed to generate an emotional feature vector. Based on the vector, an emotional frame image is generated and inserted into the interactive video stream. The implicit key point extraction model and the multimodal large model are used for feature extraction and processing.

Benefits of technology

The accuracy of determining the interactive video stream of digital humans is improved, the stable state representation of emotions is achieved, the streaming transition is made smoother, and the accuracy and real-time performance of digital human interaction are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120636388A_ABST
    Figure CN120636388A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a digital human interaction video stream determination method and device, equipment, a medium and a product, and the method comprises the steps: determining a current emotion feature of a target object in a current frame image based on a collected video stream of the target object interacting with a target digital human; the video stream comprises a current frame image; performing feature smoothing processing on the current emotion feature to obtain an emotion feature vector; generating an emotion frame image when the target digital person interacts with the target object based on the emotion feature vector; and inserting the emotion frame image into an interactive video stream of the target digital human to obtain a target video stream. According to the invention, the accuracy of determining the digital human interaction video stream can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a method, apparatus, device, storage medium, and computer program product for determining a digital human interactive video stream. Background Art

[0002] Multimodal virtual humans are virtual characters that interact naturally and efficiently with humans by integrating and analyzing multiple sensory inputs (such as vision and hearing). By integrating multiple sources of information, including voice, facial expressions, and body language, to form a unified interactive interface, multimodal virtual humans can more comprehensively understand the user's intentions and emotions. Serving as the gateway to human-computer interaction in virtual spaces, multimodal virtual humans provide customers with more personalized and humane services.

[0003] Existing emotion recognition technology still has limitations due to the complexity and diversity of human emotional expression. Different emotions can share similar physiological and behavioral characteristics, potentially leading to confusion during system recognition. Furthermore, the characteristics of human emotions vary from person to person, resulting in low recognition accuracy. Furthermore, emotion recognition methods based on isolated image frames cannot accurately identify the rapid changes in human emotional characteristics, leading to cumulative errors and inaccurate determination of facial emotions in interactive video streams. This reduces the accuracy of the determination of digital human facial emotions in interactive video streams. Summary of the Invention

[0004] To solve the above technical problems, the embodiments of the present application hope to provide a method, device, equipment, storage medium and computer program product for determining a digital human interactive video stream, which can improve the accuracy of determining a digital human interactive video stream.

[0005] The technical solution of this application is achieved as follows:

[0006] The present invention provides a method for determining a digital human interactive video stream, the method comprising:

[0007] Determining, based on a collected video stream of a target object interacting with a target digital human, a current emotional feature of the target object in a current frame image; the video stream includes the current frame image;

[0008] Performing feature smoothing processing on the current emotion feature to obtain an emotion feature vector;

[0009] generating an emotion frame image when the target digital human interacts with the target object based on the emotion feature vector;

[0010] The emotion frame image is inserted into the interactive video stream of the target digital human to obtain a target video stream.

[0011] An embodiment of the present application provides a device for determining a digital human interactive video stream, the device comprising:

[0012] a determination unit, configured to determine, based on a collected video stream of a target object interacting with a target digital human, a current emotional feature of the target object in a current frame image; the video stream includes the current frame image;

[0013] a processing unit, configured to perform feature smoothing processing on the current emotion feature to obtain an emotion feature vector;

[0014] A generating unit, configured to generate an emotion frame image when the target digital human interacts with the target object based on the emotion feature vector;

[0015] The inserting unit is used to insert the emotion frame image into the interactive video stream of the target digital human to obtain a target video stream.

[0016] An embodiment of the present application provides a device for determining a digital human interactive video stream, the device comprising:

[0017] A memory, a processor and a communication bus, wherein the memory communicates with the processor via the communication bus, and the memory stores a program for determining a digital human interactive video stream that is executable by the processor. When the program for determining a digital human interactive video stream is executed, the above-mentioned method for determining a digital human interactive video stream is executed by the processor.

[0018] An embodiment of the present application provides a storage medium having a computer program stored thereon, which is applied to a core network, and is characterized in that when the computer program is executed by a processor, the above-mentioned method for determining a digital human interactive video stream is implemented.

[0019] An embodiment of the present application also provides a computer program product, including a computer program, which can be executed by a processor in the core network to complete the steps of the aforementioned method for determining the digital human interactive video stream.

[0020] The present invention provides a method, apparatus, device, storage medium, and computer program product for determining a digital human interactive video stream. The method comprises: determining the current emotional features of the target object in a current frame image based on a collected video stream of the target object interacting with the target digital human; the video stream includes the current frame image; performing feature smoothing on the current emotional features to obtain an emotional feature vector; generating an emotional frame image of the target digital human interacting with the target object based on the emotional feature vector; and inserting the emotional frame image into the interactive video stream of the target digital human to obtain a target video stream. Using the above method implementation scheme, after determining the current emotional features of the target object in the current frame image based on the collected video stream of the target object interacting with the target digital human, feature smoothing can be performed on the current emotional features to obtain an emotional feature vector. By performing feature smoothing, the emotional feature vector represents the user's stable state, making the churn push stream transition smoother, and improving the accuracy of determining the target object's emotion. Based on the accurate target object's emotional information, the facial emotion of the target digital human in the emotional frame image can be accurately determined, thereby improving the accuracy of determining the digital human interactive video stream. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 A flow chart of a method for determining a digital human interactive video stream provided in an embodiment of the present application;

[0022] Figure 2 A schematic diagram of an exemplary method for determining a digital human interactive video stream provided in an embodiment of the present application;

[0023] Figure 3 A schematic diagram of the structure of a device for determining a digital human interactive video stream provided in an embodiment of the present application;

[0024] Figure 4 A schematic diagram of the composition structure of a digital human interactive video stream determination device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0025] The following will be combined with the accompanying drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. It should be understood that the specific embodiments described here are only used to explain the present application and are not used to limit the present application.

[0026] The embodiment of the present application provides a method for determining a digital human interactive video stream, and the method is applied to a digital human interactive video stream determining device. Figure 1 A flow chart of a method for determining a digital human interactive video stream provided in an embodiment of the present application is shown as follows: Figure 1 As shown, the method for determining the digital human interactive video stream may include:

[0027] S101 : Determine the current emotional features of the target object in the current frame image based on the collected video stream of the target object interacting with the target digital human.

[0028] A method for determining a digital human interactive video stream provided by an embodiment of the present application is applicable to a scenario in which a target digital human and a target object interact with each other and a target video stream of the target digital human is determined.

[0029] In the embodiments of the present application, the apparatus for determining a digital human interactive video stream can be implemented in various forms. For example, the apparatus for determining a digital human interactive video stream described in the present application can include a core network, a network element, a server, or other devices, and the specific embodiments of the present application are not limited thereto.

[0030] In the embodiment of the present application, a camera may be used to capture image data of a target object that is engaging in a human-computer dialogue with a target digital human, thereby obtaining a video stream.

[0031] It should be noted that the camera may be a high-definition camera with a resolution greater than or equal to 1920x1080.

[0032] In the embodiment of the present application, the target object may be a human being interacting with the target digital human. It should be noted that the number of target objects may be one or more, and the specific number of target objects may be determined according to actual conditions, and the embodiment of the present application does not limit this.

[0033] Exemplarily, the number of target objects may be one.

[0034] In the embodiment of the present application, during the process of human-computer interaction between the target digital human and the target object, the video stream of the target object interacting with the target digital human can be continuously collected.

[0035] It should be noted that the video stream may be facial image video data of a user having a human-computer dialogue with a digital human.

[0036] In an embodiment of the present application, the video stream includes multiple frames of continuous frame images, and the multiple frames of continuous frame images include the current frame image.

[0037] In an embodiment of the present application, a digital human interactive video stream determination device determines the current emotional characteristics of the target object in the current frame image based on the collected video stream of the target object interacting with the target digital human, including: decoding the video stream to obtain multiple frames of continuous frame images of the target object; obtaining the current frame image from the multiple frames of continuous frame images; preprocessing the current frame image to obtain a preprocessed target image; extracting the current emotional information of the target object from the preprocessed target image; and determining the current emotional characteristics of the target object based on the current emotional information.

[0038] In an embodiment of the present application, the process of decoding a video stream to obtain multiple continuous frame images of a target object includes: sampling the video stream according to a preset sampling frequency to obtain multiple continuous frame images of the target object; the video stream can also be decoded in other ways to obtain multiple continuous frame images of the target object; the specific process of decoding the video stream to obtain multiple continuous frame images of the target object can be determined according to actual conditions, and the embodiment of the present application does not limit this.

[0039] It should be noted that by sampling the video stream according to a preset sampling frequency, continuous image frames at a fixed frame rate can be obtained.

[0040] It should also be noted that the preset sampling frequency can be the sampling frequency configured in the digital human interactive video stream determination device, or it can be information transmitted to the digital human interactive video stream determination device by other devices, or it can be information obtained by the digital human interactive video stream determination device through other methods. The specific method of obtaining the preset sampling frequency can be determined according to actual conditions, and the embodiments of this application do not limit this.

[0041] In an embodiment of the present application, an image frame closest to the current moment can be obtained from multiple consecutive frame images to obtain the current image frame.

[0042] In an embodiment of the present application, preprocessing may include a processing method for adjusting image clarity, a method for adjusting image brightness, or other methods. The specific preprocessing method can be determined based on actual conditions, and the embodiment of the present application does not limit this.

[0043] In an embodiment of the present application, the process of extracting the current emotional information of the target object from the preprocessed target image includes: inputting the preprocessed target image into an implicit key point extraction model (such as a deep neural network) to obtain the current emotional information of the target object; the current emotional information of the target object can also be extracted from the preprocessed target image by other means; the specific implementation method can be determined according to actual conditions, and the embodiment of the present application does not limit this.

[0044] It should be noted that the current emotional information can be represented by implicit key points. Implicit key points are not equivalent to the commonly used facial feature points of humans, but are key points used to determine the current emotional information in the latent space obtained by training the feature extraction model (i.e., implicit key point extraction model) on the digital human emotion and behavior extraction task.

[0045] In the embodiment of the present application, the current emotion information corresponds to the current emotion feature one to one, that is, one piece of current emotion information corresponds to one current emotion feature.

[0046] In the embodiment of the present application, the current emotion information includes: key point information of the target object's face for controlling the target object's face to make different expressions (the key point information can be represented by a vector X exp and X1) and the transformation vector (R a and T s ).

[0047] In the embodiment of the present application, the current emotion information (X exp , X1, R a and T s ), determine the current emotional characteristics of the target object (X E ):

[0048] X E =(X exp ⊕X1)·R a ⊕T s (1)

[0049] Specifically, The facial emotion information of the target object can be represented by a vector in the latent space (i.e., latent vector). The implicit key point extraction model is used to represent the key point features of the target object’s facial emotion information. Through experiments, we get X exp It is appropriate to use 47-dimensional representation, X1 is appropriate to use 32-dimensional vector representation, facial emotion information (X exp ) and key point features (X1) are both space vectors. is an affine matrix, which represents the vector affine transformation in 3D space. is the spatial offset. ⊕ is the sum of vectors by position. X obtained after spatial transformation E is the current emotional characteristics of the target object.

[0050] In an embodiment of the present application, the digital human interactive video stream determination device preprocesses the current frame image to obtain a preprocessed target image, including: extracting the image of the target part of the target object from the current frame image to obtain the target image; and positioning and correcting the target image to obtain the preprocessed target image.

[0051] In an embodiment of the present application, a face recognition model can be used to extract, locate, and correct the facial image of the person in the figure (i.e., the face recognition model is used to extract the image of the target part of the target object from the current frame image to obtain the target image; and the target image is located and corrected to obtain a pre-processed target image).

[0052] In the embodiment of the present application, the target part may be the face or the entire head of the target object.

[0053] In an embodiment of the present application, the process of positioning and correcting the target image to obtain a preprocessed target image includes: positioning and correcting the target part in the target image to obtain the preprocessed target image.

[0054] For example, if the face in the target image is tilted, the face in the target object needs to be located and rectified to obtain a pre-processed target image. If the face in the target image is in profile, the face in the target object needs to be located and rectified to obtain a frontal face, thereby obtaining a pre-processed target image.

[0055] This application addresses the problem of low recognition accuracy when constructing a multimodal digital human (i.e., a target digital human) because the characteristic expressions of human emotions vary from person to person and different emotions may have similar physiological and behavioral characteristics. A method for modeling user emotions and behavioral states using latent space representations of key points on the human face is proposed. Specifically, a feature extraction model (i.e., an implicit key point extraction model) constructed using a deep neural network is used to extract the current emotional information of the target object interacting with the target digital human (including key point information and transformation vectors on the target object's face). The current emotional features are obtained by training the implicit key point extraction model on the digital human emotion and behavior extraction task. The implicit key point extraction model of the latent space representation is obtained through multiple experiments. In the latent space, the implicit key points of the facial features are extracted by the implicit key point extraction model to obtain the current emotional features.

[0056] In an embodiment of the present application, the implicit key point extraction model can be used to determine the motion changes of implicit key points in the user image. Based on the implicit key point extraction model, the facial feature changes associated with the user's emotions and behaviors can be accurately determined.

[0057] In the embodiments of the present application, existing digital human technology, in order to achieve precise control of digital humans, requires the creation of 3D model assets of digital human movements through motion capture technology, and the use of lines or stick figures to represent the movements of people or other objects. While investing huge manpower, the movements of digital humans are determined by the production level of the modeler, and it is impossible to accurately create precise facial expressions similar to those of the human body. Based on the use of an implicit key point extraction model, a feature extraction network trained with a large amount of data can simultaneously solve the two key problems of precise control and precise modeling.

[0058] S102: Perform feature smoothing processing on the current emotional feature to obtain an emotional feature vector.

[0059] In an embodiment of the present application, the digital human interactive video stream determination device determines the current emotional features of the target object in the current frame image based on the collected video stream of the target object interacting with the target digital human, and then performs feature smoothing processing on the current emotional features to obtain an emotional feature vector.

[0060] In an embodiment of the present application, the digital human interactive video stream determination device performs feature smoothing processing on the current emotional features to obtain an emotional feature vector, including: obtaining historical emotional features corresponding to a preset frame image before the current frame image; using the historical emotional features to perform differential processing on the current emotional features to obtain an emotional feature vector.

[0061] In the embodiment of the present application, the preset frame image may be n frames of image. For example, the preset frame image may be 5 frames of image, or the preset frame image may be 3 frames of image, or the preset frame image may be another number of frames. The specific number of images in the preset frame image may be determined according to actual conditions, and the embodiment of the present application does not limit this.

[0062] In an embodiment of the present application, the process of using historical emotional features to perform differential processing on current emotional features to obtain an emotional feature vector includes obtaining a weight coefficient; determining the offset between the historical emotional features and the current emotional features; and performing differential processing on the current emotional features according to the weight coefficient and the offset to obtain an emotional feature vector.

[0063] It should be noted that the weight coefficient can be a value configured in the digital human interactive video stream determination device, or a value transmitted to the digital human interactive video stream determination device by other devices, or a value determined by the digital human interactive video stream determination device through other methods. The specific way in which the digital human interactive video stream determination device obtains the weight coefficient can be determined according to actual conditions, and the embodiments of the present application do not limit this.

[0064] In the embodiment of the present application, the number of historical emotion features may be n, i.e., n historical emotion features may be determined corresponding to n image frames prior to the current image frame. Specifically, the method for determining n historical emotion features based on n image frames is the same as the method for determining the current emotion feature based on the current image frame. For details, please refer to the method for determining the current emotion feature based on the current image frame.

[0065] In the embodiment of the present application, a fusion smoothing algorithm can be used to perform differential processing on the current emotional features to obtain a stable feature representation (i.e., according to the weight coefficient (α) and the offset between the historical emotional features and the current emotional features (including X △E,t-1 、X △E,t-2 ,…,X △E,t-n ) for the current emotional feature (i.e., the X on the right side of the equation (2) E,t ) is processed by differential processing to obtain the emotional feature vector (i.e., the X on the left side of the equal sign in formula (2) E,t )), the fusion smoothing algorithm formula is shown in formula (2):

[0066] X E,t =αX E,t +α(1-α)X △E,t-1 +α(1-α) 2 X △E,t-2 +...+α(1-α) n X △E,t-n (2)

[0067] Among them, X △E,t-n Represents the emotional characteristics of n periods before the current moment t (i.e., historical emotional characteristics) relative to the current emotional characteristics X E,t α is the weight coefficient. Experiments have shown that n should be less than 5 and α should be in the range of 0.15–0.3.

[0068] This application uses differential processing to process facial emotion information (i.e. X exp ) and key point features (i.e., X1) are used to model the current emotion features, and the increment (i.e., offset) between the historical emotion features and the current emotion features is used to smoothly fuse the current emotion features, so that the transition between the current emotion features and the historical emotion features is smooth and accurate, thereby obtaining a stable emotion feature vector of the user, making the streaming transition smoother.

[0069] In an embodiment of the present application, the process of using historical emotional features to perform differential processing on current emotional features to obtain an emotional feature vector may also include determining the feature mean of historical emotional features and current emotional features to obtain an emotional feature vector; the current emotional features may also be differentially processed using other methods and historical emotional features to obtain an emotional feature vector; the specific implementation method may be determined based on actual conditions, and the embodiment of the present application does not limit this.

[0070] S103: Generate an emotion frame image when the target digital human interacts with the target object based on the emotion feature vector.

[0071] In an embodiment of the present application, the digital human interactive video stream determination device performs feature smoothing processing on the current emotional features, and after obtaining the emotional feature vector, generates an emotional frame image when the target digital human interacts with the target object based on the emotional feature vector.

[0072] In an embodiment of the present application, a digital human interactive video stream determination device generates an emotional frame image based on an emotional feature vector when a target digital human interacts with a target object, comprising: obtaining audio information collected when collecting a video stream; obtaining initial state information and noise information of the target digital human when there is no interactive flow between the target digital human and the target object; adding noise information to the initial state information to obtain target state information; and using a multimodal large model to process the audio information, target state information, and emotional feature vector to obtain an emotional frame image.

[0073] In the embodiment of the present application, when the digital human performs human-computer interaction with the target object, a camera can be used to capture the video stream of the target object, and a microphone can be used to capture the audio information of the target object.

[0074] In an embodiment of the present application, when it is determined that the digital human and the target object are performing human-computer interaction, it is determined that there is an interactive flow state between the target digital human and the target object; when it is not determined that the digital human and the target object are performing human-computer interaction, it is determined that there is no interactive flow state between the target digital human and the target object.

[0075] In an embodiment of the present application, the initial state information can be the preset emotional information of the face of the target digital person, or it can be the state flow of the target digital person at the current moment. The specific initial state information can be determined according to actual conditions, and the embodiment of the present application does not limit this.

[0076] It should be noted that the initial state information can be obtained from the target digital human, or through other methods. The specific method of obtaining the initial state information can be determined according to actual conditions, and the embodiments of the present application do not limit this.

[0077] In the embodiment of the present application, the noise information can be information transmitted from other devices to the digital human interactive video stream determination device, or it can be information configured in the digital human interactive video stream determination device, or it can be information obtained by the digital human interactive video stream determination device through other means. The specific way in which the digital human interactive video stream determination device obtains the noise information can be determined according to actual conditions, and the embodiment of the present application does not limit this.

[0078] In the embodiment of the present application, the manner of adding noise information to the initial state information can be determined according to actual conditions, and the embodiment of the present application does not limit this.

[0079] In the embodiment of the present application, the multimodal large model can be information configured in the digital human interactive video stream determination device, or it can be information transmitted to the digital human interactive video stream determination device by other devices, or it can be information obtained by the digital human interactive video stream determination device through other methods. The specific way in which the digital human interactive video stream determination device obtains the multimodal large model can be determined according to actual conditions, and the embodiment of the present application does not limit this.

[0080] In the embodiments of the present application, the target state information is the digital human state stream of the target digital human. The digital human state stream is a video stream that continuously outputs the digital human's facial expressions. Because the audio stream is driven quickly enough, it is generally synchronized and output after the video stream is generated and then synchronized to the video encoder for processing. For the digital human state stream, the system maintains the digital human's initial state information E0. In the state without interactive stream output, the system adds a noise variable (i.e., noise information) to the initial state information and directly outputs it to the video encoder for synthesized video output. By maintaining this state and combining the interpolation calculation of the latent space state equation, the digital human can complete streaming operations such as posture switching in silent state and real-time interruption.

[0081] In order to achieve the streaming output of digital humans, this application uses a method of adding state increments (noise information) to the basic state (i.e., initial state information) to determine the emotional state of the digital human. The state increments are modeled using a large multimodal model to improve the speed of streaming reasoning of the digital human.

[0082] In an embodiment of the present application, a digital human interactive video stream determination device uses a multimodal large model to process audio information, target state information, and emotional feature vectors. Before obtaining an emotional frame image, the device obtains target state information corresponding to the target part of the target digital human when there is an interactive flow between the target digital human and the target object.

[0083] In an embodiment of the present application, a digital human interactive video stream determination device uses a multimodal large model to process audio information, target state information and emotional feature vectors to obtain an emotional frame image, which includes: extracting an audio feature vector from the audio information; filtering the audio feature vector based on the time of the video stream to obtain a target audio vector; extracting the digital human emotional state information of the target digital human from the target state information to obtain a digital human emotional state vector; inputting the target audio vector, the digital human emotional state vector and the emotional feature vector into the multimodal large model to obtain output information; and determining the emotional frame image based on the initial state information and the output information.

[0084] In the embodiment of the present application, an automatic speech recognition (ASR) model or an audio voiceprint feature extraction model can be used to extract an audio feature vector from the audio information, and the video image data of the same time range (the image data of the current frame image) can be matched by the cosine similarity algorithm to obtain the target audio vector X. A,t (i.e., the audio feature vector is screened based on the time of the video stream to obtain the target audio vector), ensuring the synchronization relationship between the audio data and the image data.

[0085] In an embodiment of the present application, if the target state information is video data, the current target frame image of the target digital human at the current moment can be obtained from the target state information; and the digital human's emotional state vector can be determined from the current target image frame. The method for determining the digital human's emotional state vector from the current target image frame is the same as the method for determining the emotional feature vector from the current frame image. That is, the digital human's emotional state vector can be determined from the current target image frame in the same manner as the emotional feature vector is determined from the current frame image.

[0086] It should also be noted that the target state information can also be other forms of data, such as text data or binary data. The specific target state information can be determined according to actual conditions, and the embodiments of the present application do not limit this.

[0087] In the embodiment of the present application, when a user initiates a conversation, the interactive state of the digital human will be triggered. The video stream of the target digital human can be obtained, and the hidden state of the target digital human can be extracted from the video stream of the target digital human to obtain the digital human emotional state vector H of the target digital human. E,t .

[0088] In the embodiment of the present application, the target audio vector (X A,t ), digital human emotional state vector (H E,t ) and the emotional feature vector (X E,t) to obtain splicing information, and input the splicing information into the multimodal large model to obtain output information.

[0089] In the embodiment of the present application, the multimodal large model accepts input as (X E,t ,X A,t ,H E,t ) The model maintains a fixed-length context through window attention.

[0090] In an embodiment of the present application, the process of determining an emotion frame image based on initial state information and output information includes: determining a deviation between the output information and the initial state information, and determining the emotion frame image based on the deviation and the output information.

[0091] In the embodiment of the present application, the decoding process of the multimodal large model is divided into two steps:

[0092] Through the model decoding process, the next step H is generated E,t+1 predictions.

[0093] Through the additional decoding generator, the architecture is the Transformer Decoder architecture, through E0,H E,t+1 Input, get Estimates, It can be superimposed with E0 to obtain the continuous output of the target digital human’s emotional frame image. And the digital human’s emotional state vector H E,t Control the coherent output of digital humans.

[0094] S104: Insert the emotion frame image into the interactive video stream of the target digital human to obtain the target video stream.

[0095] In an embodiment of the present application, after the digital human interactive video stream determination device generates an emotional frame image when the target digital human interacts with the target object based on the emotional feature vector, the emotional frame image is inserted into the interactive video stream of the target digital human to obtain the target video stream.

[0096] It should be noted that the target video stream includes the current emotional state information of the target digital human.

[0097] In an embodiment of the present application, the generated emotion frame image of the target digital human is inserted into the interactive video stream of the target digital human, and the synthesized video stream (ie, the target video stream) is obtained through a video encoder for display.

[0098] In the embodiments of the present application, the accuracy and real-time performance of virtual digital human interaction are significantly improved by innovatively adopting a multimodal large model and latent space state representation. By collecting the user's current emotional information, the latent state space in the current frame image can be modeled, and the generation speed can be greatly improved through streaming output. This technical improvement not only solves the problems of inaccurate emotion recognition and slow generation speed in existing methods, but also achieves accurate modeling of user emotions and behaviors through differential processing and state smoothing technology. In addition, the present application also achieves continuous and smooth output of the digital human state by maintaining the current state stream of the virtual digital human and using a multimodal large model to model the state increments. The technical solution in this application provides a more natural and efficient human-computer interaction experience for virtual digital humans, and at the same time provides strong technical support for the application of digital human technology in content creation, education and training, entertainment, and marketing.

[0099] In the embodiments of the present application, the streaming synthesis method can also be changed to use pre-rendering technology to generate key frames or animation sequences of digital humans in advance, and then select the corresponding pre-rendered content for playback based on user input during real-time interaction. This can reduce the computational burden of real-time rendering and improve response speed.

[0100] It should be noted that based on state incremental modeling, the digital human gains control of the process state and decouples the synthesis of model interaction and digital human video stream. This application is different from the digital human technology based on single-frame image rendering. The digital human streaming interaction created based on this application has higher controllability.

[0101] For example, Figure 2As shown in the figure: the video stream (camera input) of the target object interacting with the target digital human is decoded to obtain multiple frames of continuous frame images of the target object; the current frame image is obtained from the multiple frames of continuous frame images; the image of the target part of the target object is extracted from the current frame image to obtain the target image; the target image is positioned and corrected to obtain the preprocessed target image; the current emotional information of the target object is extracted from the preprocessed target image (features are extracted into latent space using a neural network); the current emotional features of the target object are determined based on the current emotional information (key point modeling in latent space); the historical emotional features corresponding to the preset frame images before the current frame image are obtained; the current emotional features are differentially processed using the historical emotional features to obtain the emotional feature vector (differential model feature fusion, emotional and behavioral information input); the audio information collected when the video stream is collected is obtained (audio input); and the target digital human and the target object are positioned between the target digital human and the target object. In the absence of an interactive flow state, the initial state information and noise information of the target digital human are obtained; noise information is added to the initial state information to obtain the target state information (virtual digital human state flow); audio feature vectors are extracted from the audio information; audio feature vectors are filtered based on the time of the video stream to obtain the target audio vector (speech and voiceprint recognition module, text information input); the latent state information of the target digital human is extracted from the target state information to obtain the digital human emotional state vector (latent cause space key point modeling); the target audio vector, the digital human emotional state vector and the emotional feature vector are input into the multimodal large model to obtain output information; based on the initial state information and output information, the emotional frame image is determined (the multimodal large model integrates the input of multiple modalities, and the generator generates the digital human state flow according to the increment); the emotional frame image is inserted into the interactive video flow of the target digital human to obtain the target video flow (the state flow output is inserted into the current flow).

[0102] It can be understood that after determining the current emotional features of the target object in the current frame image based on the collected video stream of the target object interacting with the target digital human, the current emotional features can be smoothed to obtain an emotional feature vector. By performing feature smoothing, the emotional feature vector is made to represent the user's stable state, making the churn push transition smoother, thereby improving the accuracy of determining the target object's emotions. Based on the accurate emotional information of the target object, the facial emotions of the target digital human in the emotional frame image can be accurately determined, thereby improving the accuracy of determining the digital human interaction video stream.

[0103] Based on the same inventive concept as the above-mentioned method for determining a digital human interactive video stream, an embodiment of the present application provides a digital human interactive video stream determining device 1, corresponding to a digital human interactive video stream determining method; Figure 3 A schematic diagram of the structure of a digital human interactive video stream determination device provided in an embodiment of the present application Figure 1, the digital human interactive video stream determining device 1 may include:

[0104] A determination unit 11 is configured to determine a current emotional feature of a target object in a current frame image based on a collected video stream of the target object interacting with the target digital human; the video stream includes the current frame image;

[0105] A processing unit 12 is configured to perform feature smoothing processing on the current emotional feature to obtain an emotional feature vector;

[0106] A generating unit 13 is configured to generate an emotion frame image when the target digital human interacts with the target object based on the emotion feature vector;

[0107] The inserting unit 14 is configured to insert the emotion frame image into the interactive video stream of the target digital human to obtain a target video stream.

[0108] In some embodiments of the present application, the apparatus further includes an acquisition unit;

[0109] The acquisition unit is configured to acquire historical emotion features corresponding to a preset frame image before the current frame image;

[0110] The processing unit 12 is configured to perform differential processing on the current emotion feature using the historical emotion feature to obtain the emotion feature vector.

[0111] In some embodiments of the present application,

[0112] The acquisition unit is configured to acquire the audio information collected when the video stream is collected; and when there is no interactive flow between the target digital human and the target object, to acquire the initial state information and noise information of the target digital human;

[0113] The adding unit is configured to add the noise information to the initial state information to obtain target state information;

[0114] The processing unit 12 is configured to process the audio information, the target state information, and the emotion feature vector using a multimodal large model to obtain the emotion frame image.

[0115] In some embodiments of the present application, the apparatus further comprises an extraction unit, a screening unit, and an input unit;

[0116] The extraction unit is configured to extract an audio feature vector from the audio information; extract the digital human emotional state information of the target digital human from the target state information to obtain the digital human emotional state vector;

[0117] The screening unit is configured to screen the audio feature vector based on the time of the video stream to obtain a target audio vector;

[0118] The input unit is used to input the target audio vector, the digital human emotional state vector and the emotional feature vector into the multimodal large model to obtain output information;

[0119] The determining unit 11 is configured to determine the emotion frame image based on the initial state information and the output information.

[0120] In some embodiments of the present application, the acquisition unit is configured to acquire target state information corresponding to a target part of the target digital human when there is an interactive flow between the target digital human and the target object.

[0121] In some embodiments of the present application, the apparatus further comprises a decoding unit;

[0122] The decoding unit is used to decode the video stream to obtain multiple consecutive frame images of the target object;

[0123] The acquiring unit is configured to acquire the current frame image from the multiple consecutive frame images;

[0124] The processing unit 12 is configured to preprocess the current frame image to obtain a preprocessed target image;

[0125] The extraction unit is used to extract the current emotion information of the target object from the preprocessed target image;

[0126] The determining unit 11 is configured to determine the current emotional characteristics of the target object according to the current emotional information.

[0127] In some embodiments of the present application, the extraction unit is configured to extract an image of a target portion of the target object from the current frame image to obtain a target image;

[0128] The processing unit 12 is used to perform positioning and correction processing on the target image to obtain the pre-processed target image.

[0129] It should be noted that, in actual applications, the above-mentioned determination unit 11, processing unit 12, generation unit 13 and insertion unit 14 can be implemented by a processor 15 on the digital human interactive video stream determination device, specifically a CPU (Central Processing Unit), MPU (Microprocessor Unit), DSP (Digital Signal Processing) or Field Programmable Gate Array (FPGA); the above-mentioned data storage can be implemented by a memory 16 on the digital human interactive video stream determination device.

[0130] The embodiment of the present application also provides a device for determining a digital human interactive video stream, such as Figure 4 As shown, the digital human interactive video stream determination device includes: a processor 15, a memory 16 and a communication bus 17. The memory 16 communicates with the processor 15 through the communication bus 17. The memory 16 stores a program executable by the processor 15. When the program is executed, the digital human interactive video stream determination method described above is executed by the processor 15.

[0131] In practical applications, the memory 16 may be a volatile memory, such as a random-access memory (RAM); or a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD) or a solid-state drive (SSD); or a combination of the above types of memory, and provide instructions and data to the processor 15.

[0132] An embodiment of the present application provides a computer-readable storage medium having a computer program thereon, which, when executed by the processor 15 , implements the method for determining a digital human interactive video stream as described above.

[0133] Illustratively, an embodiment of the present application further provides a computer program product, including a computer program, which can be executed by the processor 15 in the digital human interactive video stream determination device to complete the steps of the aforementioned digital human interactive video stream determination method.

[0134] It can be understood that after determining the current emotional features of the target object in the current frame image based on the collected video stream of the target object interacting with the target digital human, the current emotional features can be smoothed to obtain an emotional feature vector. By performing feature smoothing, the emotional feature vector is made to represent the user's stable state, making the churn push transition smoother, thereby improving the accuracy of determining the target object's emotions. Based on the accurate emotional information of the target object, the facial emotions of the target digital human in the emotional frame image can be accurately determined, thereby improving the accuracy of determining the digital human interaction video stream.

[0135] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of hardware embodiments, software embodiments, or embodiments combining software and hardware. Furthermore, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage and optical storage, etc.) containing computer-usable program code.

[0136] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0137] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0138] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1A step that specifies a function in one or more boxes.

[0139] The above description is merely a preferred embodiment of the present application and is not intended to limit the scope of protection of the present application.

Claims

1. A method for determining a digital human interactive video stream, characterized in that: The method comprises: Determining, based on a collected video stream of a target object interacting with a target digital human, a current emotional feature of the target object in a current frame image; the video stream includes the current frame image; Performing feature smoothing processing on the current emotion feature to obtain an emotion feature vector; generating an emotion frame image when the target digital human interacts with the target object based on the emotion feature vector; The emotion frame image is inserted into the interactive video stream of the target digital human to obtain a target video stream.

2. The method according to claim 1, characterized in that The step of performing feature smoothing on the current emotion feature to obtain an emotion feature vector includes: Acquire historical emotion features corresponding to a preset frame image before the current frame image; The historical emotion feature is used to perform differential processing on the current emotion feature to obtain the emotion feature vector.

3. The method according to claim 1, characterized in that The generating of the emotion frame image when the target digital human interacts with the target object based on the emotion feature vector includes: Acquiring audio information collected when collecting the video stream; When there is no interactive flow between the target digital human and the target object, obtaining initial state information and noise information of the target digital human; Adding the noise information to the initial state information to obtain target state information; The audio information, the target state information and the emotion feature vector are processed using a multimodal large model to obtain the emotion frame image.

4. The method according to claim 3, characterized in that The method of processing the audio information, the target state information, and the emotion feature vector using a multimodal large model to obtain the emotion frame image includes: extracting an audio feature vector from the audio information; Filtering the audio feature vector based on the time of the video stream to obtain a target audio vector; Extracting the digital human emotional state information of the target digital human from the target state information to obtain the digital human emotional state vector; Inputting the target audio vector, the digital human emotional state vector and the emotional feature vector into a multimodal large model to obtain output information; The emotion frame image is determined based on the initial state information and the output information.

5. The method according to claim 3, characterized in that Before the method processes the audio information, the target state information, and the emotion feature vector using the multimodal large model to obtain the emotion frame image, the method further includes: When there is an interactive flow between the target digital human and the target object, target state information corresponding to the target part of the target digital human is obtained.

6. The method according to claim 1, characterized in that The determining of the current emotional features of the target object in the current frame image based on the collected video stream of the target object interacting with the target digital human includes: Decoding the video stream to obtain multiple consecutive frame images of the target object; Acquire the current frame image from the multiple consecutive frame images; Preprocessing the current frame image to obtain a preprocessed target image; extracting current emotional information of the target object from the preprocessed target image; Determine the current emotional characteristics of the target object according to the current emotional information.

7. The method according to claim 6, characterized in that The preprocessing of the current frame image to obtain a preprocessed target image includes: Extracting an image of a target part of the target object from the current frame image to obtain a target image; Positioning and correction processing are performed on the target image to obtain the pre-processed target image.

8. A device for determining a digital human interactive video stream, characterized in that: The device comprises: a determination unit, configured to determine, based on a collected video stream of a target object interacting with a target digital human, a current emotional feature of the target object in a current frame image; the video stream includes the current frame image; a processing unit, configured to perform feature smoothing processing on the current emotion feature to obtain an emotion feature vector; A generating unit, configured to generate an emotion frame image when the target digital human interacts with the target object based on the emotion feature vector; The inserting unit is used to insert the emotion frame image into the interactive video stream of the target digital human to obtain a target video stream.

9. A device for determining a digital human interactive video stream, characterized in that: The device comprises: A memory, a processor and a communication bus, wherein the memory communicates with the processor via the communication bus, the memory stores a program determined by a digital human interactive video stream executable by the processor, and when the program determined by the digital human interactive video stream is executed, the method according to any one of claims 1 to 7 is executed by the processor.

10. A storage medium having a computer program stored thereon, applied to a device for determining a digital human interactive video stream, characterized in that: When the computer program is executed by a processor in the core network, the method according to any one of claims 1 to 7 is implemented.

11. A computer program product comprising a computer program, characterized in that The computer program implements the method according to any one of claims 1 to 7 when executed by a processor.

Citation Information

Patent Citations

  • Digital human driving method and system, electronic equipment and readable storage medium

    CN117218251A

  • Digital human posture adjusting method and device combined with image recognition

    CN119473209A

  • Digital human driving method and system for museum interaction in combination with large model technology

    CN119828896A

  • Emotional evolution method and terminal for virtual avatar in educational metaverse

    US20250014470A1