Video communication method and device, computer readable storage medium and electronic equipment
By extracting voiceprint and rhythmic features and matching them with a local virtual avatar model, an animation sequence is generated, which solves the problem of video interruption caused by network instability in video communication, achieves high-quality, low-latency video completion, and improves user experience and resource utilization efficiency.
Patent Information
- Application Number
- CN202511122371.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-11
- Publication Date
- 2025-11-18
AI Technical Summary
Existing video communication technologies struggle to achieve high-fidelity, low-latency dynamic video completion when the network is unstable. Edge processing is limited by computing power, and cloud solutions cannot stably transmit models when the network is abnormal, resulting in video interruptions and poor user experience.
By extracting the speaker's voiceprint and rhythmic features, matching them with a local virtual avatar model, and combining them with facial animation templates to generate animation sequences, the virtual avatar animation is rendered in real time, taking over the communication video stream, and switching back to the real-time video stream when the network recovers. The local caching model is used to achieve continuous completion of video content.
To avoid screen stuttering and interruptions when the network is unstable, provide a high-quality, low-latency, and personalized video communication experience, enhance the anthropomorphism and realism of interaction, and optimize the utilization of terminal resources.
Smart Images

Figure CN120980267A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of video processing, and in particular, to a video communication method, a video communication device, a computer readable storage medium, and an electronic device. BACKGROUND
[0002] In the process of real-time video communication, the instability of network environment often leads to packet loss, delay, freezing or even interruption of video stream, which seriously affects the continuity of conversation and user experience. To address this problem, the industry uses dynamic video completion technology to generate alternative pictures when communication quality decreases to maintain visual coherence.
[0003] However, the existing solutions generally have the problem of unsatisfactory completion effect under network communication abnormality: on the one hand, end-side processing is limited by the computing power of terminal devices, making it difficult to drive high-fidelity digital human models, and the generated animations often exhibit misaligned mouth shapes and stiff expressions, especially on mid- and low-end devices; on the other hand, while cloud-based solutions have strong computing power support, they cannot stably transmit model or animation data when the network is abnormal, resulting in the inability to timely deliver the completion content, which exacerbates the sense of picture interruption and loses the meaning of completion.
[0004] Therefore, how to break through the dual limitations of performance and network under poor communication quality, and realize voice-driven, low-latency, high-fidelity dynamic video completion has become a focus problem for relevant technical personnel.
[0005] In view of this, there is an urgent need in the art to develop a new video communication method and device.
[0006] It should be noted that the information disclosed in the above background section is only used to strengthen the understanding of the background of the present disclosure. SUMMARY
[0007] The purpose of the present disclosure is to provide a video communication method, a video communication device, a computer readable storage medium, and an electronic device, thereby at least partially overcoming the technical problem of how to break through the dual limitations of performance and network under poor communication quality, and realize voice-driven, low-latency, high-fidelity dynamic video completion due to the limitations of related technologies.
[0008] Other characteristics and advantages of the present disclosure will become apparent from the following detailed description, or will be learned by practice of the present disclosure.
[0009] According to a first aspect of the present disclosure, a video communication method is provided, comprising:
[0010] extracting a voiceprint feature corresponding to the speech voice of the current speaker;
[0011] determine whether the communication quality meets a preset quality standard based on the voiceprint feature.
[0012] extract a rhythm and prosody feature corresponding to the speech, and generate a face animation sequence of the target virtual image model based on the voiceprint feature, the rhythm and prosody feature, and a pre-stored face animation template.
[0013] render a virtual image animation by real-time rendering the face animation sequence, and replace a real-time communication video stream based on the virtual image animation.
[0014] In an example embodiment of the present disclosure, whether the communication quality meets the preset quality standard is determined by:
[0015] monitoring a network transmission indicator of a communication link, the network transmission indicator including at least one of a packet loss rate, a transmission delay, a jitter, and a bandwidth utilization rate;
[0016] determining whether the communication quality meets the preset quality standard based on a comparison result between the network transmission indicator and a preset indicator threshold.
[0017] In an example embodiment of the present disclosure, the matching and loading of the target virtual image model corresponding to the current speaker from a preset virtual image model library based on the voiceprint feature includes:
[0018] calculating a hash value corresponding to the voiceprint feature;
[0019] matching and loading the target virtual image model corresponding to the current speaker according to a correspondence between the hash value and a file index of a virtual image model pre-set.
[0020] In an example embodiment of the present disclosure, the extraction of the rhythm and prosody feature corresponding to the speech includes:
[0021] dividing the speech into a plurality of speech segments with a unit time length;
[0022] extracting a speech energy, a stress position, and a speech speed change feature corresponding to each of the speech segments based on a Fourier transform;
[0023] wherein the speech energy and a lip shape control parameter of the target virtual image model have a preset mapping relationship.
[0024] In an example embodiment of the present disclosure, after the rhythm and prosody feature corresponding to the speech is extracted, the method further includes:
[0025] extracting text sentiment features and acoustic features corresponding to the spoken voice; the text sentiment features have a preset mapping relationship with micro-expression control parameters of the target virtual image model;
[0026] performing multi-modal feature fusion on the text sentiment features, the acoustic features, the voiceprint features, and the rhythm and prosody features to obtain fused features;
[0027] generating a facial animation sequence of the target virtual image model according to the fused features and the pre-stored facial animation template.
[0028] In an example embodiment of the present disclosure, the generating of the facial animation sequence of the target virtual image model according to the fused features and the pre-stored facial animation template comprises:
[0029] collecting eye movement parameters of the current speaker;
[0030] generating a facial animation sequence of the target virtual image model according to the fused features, the eye movement parameters, and the pre-stored facial animation template.
[0031] In an example embodiment of the present disclosure, after the facial animation sequence of the target virtual image model is generated, the method further comprises:
[0032] injecting a body action sequence into the target virtual image model; the body action sequence comprises at least one of head action, hand action, or action of other parts of the body.
[0033] In an example embodiment of the present disclosure, the replacing of the real-time communication video stream with the virtual image animation comprises:
[0034] performing Gaussian blur processing on the last valid video frame of the real-time communication video stream to obtain a first processed picture;
[0035] performing Gaussian blur processing on the initial frame of the virtual image animation to obtain a second processed picture;
[0036] generating a first transition picture according to the first processed picture and the second processed picture;
[0037] replacing the real-time communication video stream with the first transition picture and the virtual image animation.
[0038] In an example embodiment of the present disclosure, the method further comprises:
[0039] when it is detected that the communication quality meets the preset quality standard, replacing the virtual image animation with the recovered real-time communication video stream.
[0040] In the example embodiment of the present disclosure, the resuming of the real-time communication video stream replaces the animation of the virtual image, including:
[0041] Gaussian blur processing is performed on the last valid frame of the animation of the virtual image to obtain a third processed picture;
[0042] Gaussian blur processing is performed on the initial frame of the resumed real-time communication video stream to obtain a fourth processed picture;
[0043] A second transition picture is generated according to the third processed picture and the fourth processed picture;
[0044] The second transition picture and the resumed real-time communication video stream replace the animation of the virtual image.
[0045] In the example embodiment of the present disclosure, after extracting the voiceprint feature corresponding to the speech of the current speaker, the method further includes:
[0046] When it is detected that the communication quality meets the preset quality standard, it is determined whether the speech is from a new speaker;
[0047] In the case where it is determined that the speech is from the new speaker, video data of the new speaker is collected and uploaded to the cloud;
[0048] According to a virtual image model generated by the cloud based on the video data, the preset virtual image model library is updated.
[0049] In the example embodiment of the present disclosure, the determination of whether the speech is from a new speaker includes:
[0050] The voiceprint feature is encoded into a speaker representation vector;
[0051] The speaker representation vector is matched with a reference speaker representation vector stored in advance, and it is determined whether the speech is from a new speaker according to the similarity matching result.
[0052] In the example embodiment of the present disclosure, the determination of whether the speech is from a new speaker according to the similarity matching result includes:
[0053] If the similarity matching result does not meet a preset similarity threshold, it is determined that the speech is from a new speaker;
[0054] If the similarity matching result meets the preset similarity threshold, it is determined that the speech is not from the new speaker.
[0055] In the example embodiment of the present disclosure, the video data includes continuous video data of a preset time length;
[0056] uploading the video data to a cloud, comprising:
[0057] performing motion blur detection and brightness normalization on the video data to obtain a first processed video;
[0058] selecting a key frame image from the first processed video;
[0059] labeling image information for the key frame image, and uploading the processed key frame image to the cloud; wherein the image information comprises a current timestamp and / or a device identifier.
[0060] In an example embodiment of the present disclosure, the cloud generates the virtual avatar model based on the following manner:
[0061] extracting facial parameters of the new speaker based on the key frame image, and reconstructing a face mesh of the new speaker based on the facial parameters; the facial parameters comprise a facial geometry basis, facial deformation control points, and expression parameters;
[0062] performing face neural modeling on the face mesh to generate an initial virtual avatar model;
[0063] constructing a speech-driven model according to acoustic characteristics of the new speaker in combination with the expression parameters; the speech-driven model is used to define a mapping relationship between acoustic characteristics and facial animation parameters; the facial animation parameters comprise mouth opening degree and mouth corner offset;
[0064] fusing the initial virtual avatar model and the speech-driven model to generate the virtual avatar model.
[0065] In an example embodiment of the present disclosure, the virtual avatar model generated by the cloud based on the video data updates the preset virtual avatar model library, comprising:
[0066] receiving a model file of the encrypted virtual avatar model sent by the cloud;
[0067] storing the decrypted model file into the preset virtual avatar model library, and recording an index identifier corresponding to the model file;
[0068] calculating a hash value corresponding to the voiceprint feature, and storing a mapping relationship between the hash value and the index identifier.
[0069] According to a second aspect of the present disclosure, a video communication device is provided, comprising:
[0070] a first feature extraction module configured to extract a voiceprint feature corresponding to a speech of a current speaker;
[0071] a model loading module, configured to match and load a target virtual image model corresponding to the current speaker from a preset virtual image model library based on the voiceprint feature when it is detected that the communication quality does not meet the preset quality standard;
[0072] a second feature extraction module, configured to extract rhythm and prosody features corresponding to the speech, and generate a face animation sequence of the target virtual image model based on the voiceprint feature, the rhythm and prosody features, and a pre-stored face animation template;
[0073] an animation generation module, configured to render a virtual image animation in real time based on the face animation sequence, and replace a real-time communication video stream based on the virtual image animation.
[0074] According to a third aspect of the present disclosure, a computer readable storage medium is provided, which stores a computer program. The computer program is executed by a processor to implement the video communication method of the first aspect.
[0075] According to a fourth aspect of the present disclosure, an electronic device is provided, which includes a processor and a memory for storing executable instructions of the processor. The processor is configured to execute the video communication method of the first aspect by executing the executable instructions.
[0076] According to the above technical solutions, the video communication method, the video communication device, the computer readable storage medium and the electronic device of the exemplary embodiments of the present disclosure have at least the following advantages and positive effects:
[0077] In the technical solution provided in some embodiments of the present disclosure, by extracting the voiceprint features corresponding to the speech of the current speaker, when it is detected that the communication quality does not meet the preset quality standard, the target virtual image model corresponding to the current speaker is matched and loaded from the preset virtual image model library based on the voiceprint features, the rhythm and prosody features corresponding to the speech are extracted, the face animation sequence of the target virtual image model is generated based on the voiceprint features, the rhythm and prosody features, and the pre-stored face animation template, the virtual image animation is generated by real-time rendering of the face animation sequence, and the real-time communication video stream is replaced based on the virtual image animation. On the one hand, in the weak network scene such as insufficient network bandwidth, high packet loss rate or video stream interruption, the picture freezing, freezing or black screen problems in the traditional scheme can be avoided, the video content can be continuously completed by using the locally cached virtual image model, and the smoothness and stability of the communication process can be ensured. On the other hand, by fusing the voiceprint identity information, the speech rhythm features and the personalized face animation template, the virtual image is driven to realize the natural cooperation of the mouth opening and closing, the micro-expression change and the head movement, the personification degree and the interactive authenticity of the digital person expression are significantly improved, the dependence on the cloud is reduced, the resource utilization efficiency of the terminal side is optimized, and high-quality, low-delay and personalized video communication experience is provided for users.
[0078] It should be understood that the foregoing general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF DRAWINGS
[0079] The drawings incorporated into the specification and forming a part of the specification, show embodiments consistent with the present disclosure, and together with the specification, serve to explain the principles of the present disclosure. Obviously, the drawings in the following description are only some embodiments of the present disclosure, and other drawings can be obtained by those skilled in the art without creative labor.
[0080] Figure 1 A flowchart of a video communication method in an embodiment of the present disclosure is shown;
[0081] Figure 2 A flowchart of how to extract the rhythm and prosody features corresponding to the speech in an embodiment of the present disclosure is shown;
[0082] Figure 3 A flowchart of another way to generate the face animation sequence of the target virtual image model in an embodiment of the present disclosure is shown;
[0083] Figure 4 A flowchart of how to update the local virtual image model library in an embodiment of the present disclosure is shown;
[0084] Figure 5A flowchart illustrating how to determine whether the speech is from a new speaker in the embodiment of the present disclosure is shown.
[0085] Figure 6 A flowchart illustrating how to upload the video data to the cloud in the embodiment of the present disclosure is shown.
[0086] Figure 7 A flowchart illustrating how the cloud generates the virtual avatar model in the embodiment of the present disclosure is shown.
[0087] Figure 8 A flowchart illustrating how to update the preset virtual avatar model library according to the virtual avatar model in the embodiment of the present disclosure is shown.
[0088] Figure 9 A flowchart illustrating the overall process of the video communication method in the embodiment of the present disclosure is shown.
[0089] Figure 10 A structural schematic diagram of a video communication device in the exemplary embodiment of the present disclosure is shown.
[0090] Figure 11 A structural schematic diagram of an electronic device in the exemplary embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0091] Example implementations will now be described more fully with reference to the accompanying drawings. Example implementations may, however, be implemented in many different forms and should not be construed as limited to the implementations set forth herein; rather, these implementations are provided so that this disclosure will be thorough and complete, and will fully convey the scope of example implementations to those skilled in the art. The described features, structures, or characteristics can be combined in one or more implementations. In the following description, numerous specific details are provided to give a thorough understanding of implementations of the disclosure. One skilled in the relevant art will recognize, however, that the implementations of the disclosure can be practiced without one or more of the specific details, or with other methods, components, materials, and so forth. In other instances, well-known structures have not been described in detail to avoid obscuring aspects of the disclosure.
[0092] The terms "one", "a", "an", and "the" as used herein mean "at least one" or "one or more" unless expressly specified otherwise. The term "includes" means "includes but is not limited to" and the term "including" means "including, but not limited to". The term "based on" means "based, at least in part, on" and not "based solely on" unless specifically specified.
[0093] In addition, the accompanying drawings are only schematic and are non-drawing to scale unless otherwise specified, and specific dimensions are intended to be exemplary only. The same or similar designations may be used to represent like or similar functionalities throughout the accompanying drawings and specification.
[0094] In the field of digital human generation and dynamic video completion, the existing technology mainly adopts a single end-side processing or cloud centralized processing architecture, which has significant technical limitations.
[0095] On the one hand, the end-side generation scheme is limited by the computing power and storage resources of the terminal device, making it difficult to achieve high-fidelity 3D digital human modeling and real-time rendering on low-end devices, resulting in insufficient realism of the generated virtual image and stiff animation performance, which cannot meet the needs of high-quality video communication. On the other hand, although the pure cloud scheme has strong computing power support, it can complete complex modeling tasks in the cloud, but when the network communication quality is poor or packet loss or delay occurs, the generated virtual image model cannot be timely delivered to the terminal, resulting in video completion link interruption, picture freezing, freezing or even black screen, which seriously affects user experience.
[0096] In addition, the existing digital human system generally has the problem of unreasonable resource utilization. Most schemes use periodic collection of user video clips to update or initialize the model, lack intelligent judgment of the triggering conditions, and perform collection and upload operations regardless of whether the speaker is a new user or the network state allows, resulting in waste of bandwidth, power and computing resources.
[0097] Furthermore, the current method generally ignores the timing dynamic characteristics of the speech signal in the animation driving link of the virtual image, fails to fully exploit the mapping relationship between the rhythm and intonation information (such as speech rate, stress, pause) in the speech and the facial movements (such as mouth shape, blinking, nodding), resulting in the generated digital human animation being out of sync with the actual speech, lacking natural expression and emotional expression, and having low interactive realism.
[0098] In view of the above problems, although some related research has explored, there is still no effective solution. For example, some research proposes an interdisciplinary audio-video visualization framework for analyzing the temporal dynamics of expressive behavior, but it focuses on data visualization and academic analysis and does not involve the technical implementation of digital human modeling and real-time video completion; another research proposes a video quality evaluation method for dynamic digital humans, evaluating the rendering quality through geometric and spatiotemporal features, but still does not solve the core engineering problems such as end-cloud collaborative modeling, on-demand triggering, and speech-driven synchronization.
[0099] In summary, the existing technology faces four key technical bottlenecks:
[0100] ① The end-side computing power is insufficient to support high-quality digital human local generation;
[0101] ②Cloud dependence is too strong, and video completion continuity cannot be guaranteed when the network is abnormal;
[0102] ③Lack of on-demand triggering mechanism, resulting in waste of data collection and transmission resources;
[0103] ④Neglecting voice-driven modeling, animation and voice rhythm are out of sync, and expression is unnatural.
[0104] To solve the above problems, the present disclosure proposes a voice-driven digital human generation and dynamic video completion method based on end-to-cloud collaboration. This scheme realizes efficient allocation of resources through a collaborative architecture of "cloud generation + terminal presentation": when the network is good, high-performance computing power of the cloud is used to complete high-fidelity 3D modeling (such as based on FLAME and NeRF technology), and a 75MB-level model that can be downloaded is generated through lightweight processing; when the network is abnormal, the terminal matches the target model based on the virtual image model library cached locally, combined with voiceprint recognition, to realize seamless completion.
[0105] In an embodiment of the present disclosure, a video communication method is first provided, which at least partially overcomes the defects in the related art of how to break through the dual limitations of performance and network in the case of poor communication quality, and realize voice-driven, low-latency, high-fidelity dynamic video completion.
[0106] Figure 1 A flowchart of a video communication method in an embodiment of the present disclosure is shown, and the execution subject of the video communication method can be a terminal device.
[0107] Reference Figure 1 According to an embodiment of the present disclosure, the video communication method includes the following steps:
[0108] Step S110, when it is detected that the communication quality does not meet the preset quality standard, extracting the voiceprint features corresponding to the speaking voice of the current speaker;
[0109] Step S120, matching and loading the target virtual image model corresponding to the current speaker from the preset virtual image model library based on the voiceprint features;
[0110] Step S130, extracting the rhythm and prosody features corresponding to the speaking voice, and generating the face animation sequence of the target virtual image model based on the voiceprint features, the rhythm and prosody features, and the pre-stored face animation template;
[0111] Step S140, rendering the virtual image animation by real-time rendering the face animation sequence, and replacing the real-time communication video stream based on the virtual image animation.
[0112] In Figure 1In the technical solution provided by the illustrated embodiment, by extracting the voiceprint features corresponding to the speech of the current speaker, when it is detected that the communication quality does not meet the preset quality standard, the target virtual image model corresponding to the current speaker is matched and loaded from the preset virtual image model library based on the voiceprint features, the rhythm and prosody features corresponding to the speech are extracted, the face animation sequence of the target virtual image model is generated based on the voiceprint features, the rhythm and prosody features, and the pre-stored face animation template, the virtual image animation is generated by real-time rendering of the face animation sequence, and the real-time communication video stream is replaced based on the virtual image animation. On the one hand, in the weak network scene of insufficient network bandwidth, high packet loss rate or video stream interruption, the picture freezing, freezing or black screen problems in the traditional scheme can be avoided, the video content can be continuously completed by using the locally cached virtual image model, and the smoothness and stability of the communication process can be ensured. On the other hand, by fusing the voiceprint identity information, the speech rhythm features and the personalized face animation template, the virtual image is driven to realize the natural cooperation of the mouth opening and closing, the micro-expression change and the head movement, the personification degree and the interactive reality of the digital person expression are significantly improved, the dependence on the cloud is reduced, the resource utilization efficiency of the terminal side is optimized, and high-quality, low-delay and personalized video communication experience is provided for users.
[0113] The specific implementation process of each step in the following Figure 1 will be described in detail:
[0114] In step S110, the voiceprint features corresponding to the speech of the current speaker are extracted.
[0115] In this step, the voiceprint features corresponding to the speech of the current speaker can be extracted. The voiceprint features can be used to represent the unique acoustic characteristics of each person, so as to identify the identity of the current speaker.
[0116] After that, the network transmission indicators of the communication link can be monitored in real time to determine whether the communication quality meets the preset quality standard. The network transmission indicators can include at least one of the packet loss rate, the transmission delay, the jitter and the bandwidth utilization rate. The monitored network transmission indicators are compared with the preset indicator threshold. If the packet loss rate exceeds 15%, or the transmission delay is greater than 2 seconds, the jitter is greater than 20 milliseconds, or the bandwidth utilization rate is less than the minimum threshold for ensuring smooth transmission of the video for more than 1 second, it can be determined that the communication quality does not meet the preset quality standard. Otherwise, it can be determined that the communication quality meets the preset quality standard.
[0117] In step S120, when it is detected that the communication quality does not meet the preset quality standard, the target virtual image model corresponding to the current speaker is matched and loaded from the preset virtual image model library based on the voiceprint features.
[0118] In this step, the terminal locally stores a preset virtual image model library, the virtual image model library stores personalized virtual image model files of a plurality of registered users, each model file corresponds to an identity of a user, and a mapping relationship between a voiceprint feature extracted when the user registers and a virtual image model file index is pre-established. To improve model retrieval efficiency, the first voiceprint feature currently extracted can be hashed to generate a corresponding hash value; the terminal looks up the pre-constructed "hash value-model file index" mapping table according to the hash value, quickly locates and loads the target virtual image model matched therewith.
[0119] By locally caching the model library and using the hash index mechanism, low-delay matching and fast loading of the virtual image model can be realized in the case of communication exception and inability to access the cloud, thereby guaranteeing the real-time and continuity of video completion.
[0120] In step S130, the rhythm and prosody features corresponding to the speech are extracted, and a face animation sequence of the target virtual image model is generated based on the voiceprint features, the rhythm and prosody features, and a pre-stored face animation template.
[0121] In an optional embodiment, the rhythm and prosody features corresponding to the speech can be extracted, and a face animation sequence of the target virtual image model is generated based on the voiceprint features, the rhythm and prosody features, and a pre-stored face animation template.
[0122] Specifically, referring to Figure 2 , Figure 2 A flowchart showing how to extract the rhythm and prosody features (which can include speech energy, stress position, and speech rate change features) corresponding to the speech in the embodiments of the present disclosure is shown, which includes steps S201-S202:
[0123] In step S201, the speech is divided into a plurality of speech segments with a unit time length.
[0124] In this step, the speech can be divided into a plurality of speech segments with a unit time length (for example, 20 milliseconds, which can be set by the user according to the actual situation, and the present disclosure does not make special limitations thereon).
[0125] In step S202, the speech energy, stress position, and speech rate change features corresponding to each speech segment are extracted based on Fourier transform.
[0126] In this step, the speech energy, stress position and speech rate variation features corresponding to each speech segment can be extracted based on short-time Fourier transform. Among them, the speech energy has a preset mapping relationship with the mouth shape control parameter of the target virtual image model, which is used to drive the opening and closing amplitude of the virtual mouth shape, and realize the dynamic synchronization of speech and mouth shape animation. The stress position refers to the time position of the syllable or word that is pronounced more loudly and prominently in a piece of speech. These syllables usually have higher energy (loudness), higher fundamental frequency (pitch) or longer duration, which are used to express emphasis, semantic emphasis or emotional tendency. The speech rate variation feature describes the variation of the pronunciation speed per unit time, that is, whether the speaking speed is faster or slower, which is not simply "how many words are spoken per minute on average", but focuses on the fluctuation of local speech rate, for example, a sentence is spoken quickly, and the next sentence is slowed down.
[0127] Subsequently, the face animation sequence of the target virtual image model can be generated based on the voiceprint feature, rhythm and prosody feature, and a pre-stored face animation template. Among them, the above-mentioned face animation template is used to describe the basic action parameters including mouth opening (amplitude 0-80%), blinking (frequency 0.3-0.7 Hz), nodding (angle ±8).
[0128] Specifically, the above-mentioned face animation template can be a pre-constructed face action primitive library, which contains basic action units and dynamic parameters bound to the current speaker identity, and specifically covers atomic-level expression actions such as mouth opening (amplitude 0-80%), blinking (frequency 0.3-0.7 Hz), nodding (angle ±8°). Each action unit is defined with time envelope (such as start and end time, duration), spatial deformation parameter (such as mouth opening and closing degree, eyelid closing degree) and intensity coefficient, supporting parameterized adjustment and multi-action combination playback, to ensure natural and coherent animation performance.
[0129] In the animation generation process, the rhythm and prosody feature can be used as the core driving signal and mapped to the corresponding face control channel: the speech energy can establish a preset mapping relationship with the mouth opening and closing degree to realize the dynamic change of the lip shape with the speech intensity; the stress position can trigger micro-expressions such as blinking or eyebrow raising to enhance the rhythm of semantic expression; the speech rate variation can be used to adjust the action playback rate, and when the speech rate is fast, the action is compact and coherent, and when the speech rate is slow or paused, natural micro-expressions are inserted to improve the realism. At the same time, based on the first voiceprint feature, the user's exclusive action style parameters (such as habitual nodding frequency, expression amplitude preference) can be called to make the generated animation not only realize "sound and picture synchronization", but also have "personalized" personification expression. Through the cooperative mechanism of speech driving and personalized template, the present disclosure can efficiently generate high-quality face animation locally without relying on cloud rendering, which significantly improves the video completion effect and interaction experience in a weak network environment.
[0130] In an alternative embodiment, after extracting the rhythm and prosody features corresponding to the speech, the disclosure further provides another scheme for generating a facial animation sequence of the target virtual image model. Referring to Figure 3 , Figure 3 A flowchart for generating a facial animation sequence of the target virtual image model in another embodiment of the disclosure is shown, comprising steps S301-S303:
[0131] In step S301, text sentiment features and acoustic features corresponding to the speech are extracted; the text sentiment features have a preset mapping relationship with micro-expression control parameters of the target virtual image model.
[0132] In this step, the speech can first be converted into corresponding text by an automatic speech recognition (ASR) module; then, a terminal-side lightweight natural language processing model (such as a fine-tuned BERT-mini) is used to analyze the semantics of the text, extract text sentiment features, and generate a sentiment vector representing sentiment polarity (positive, neutral, negative) and intensity. Meanwhile, acoustic features are extracted from the original speech signal, including mel-frequency cepstral coefficients (MFCC), fundamental frequency (F0), and energy envelope, etc., to characterize the emotional prosody information of the speech.
[0133] Among them, a preset mapping relationship is established between the above-mentioned text sentiment features and the micro-expression control parameters of the target virtual image model, for example: when "positive" emotion is detected, the AU12 (mouth corner up) action intensity is enhanced; when "negative" emotion is detected, the AU4 (eyebrow up) or AU1 (forehead up) sad-related action is activated, providing semantic basis for subsequent micro-expression driving.
[0134] In step S302, the text sentiment features, acoustic features, voiceprint features, and rhythm and prosody features are fused to obtain fused features.
[0135] In this step, the feature vectors from different modalities can be time-aligned and normalized, and then input into a lightweight gated recurrent unit (GRU) network to realize dynamic interaction and fusion of cross-modal time-series features. Among them, the text sentiment features provide high-level semantic sentiment tendency, the acoustic features reflect the emotional prosody dynamics of the speech, the voiceprint features retain the speaker identity information, and the rhythm and prosody features depict the stress and speed changes of the speech. The GRU learns the weight distribution of each modality feature adaptively through the gating mechanism, and outputs a high-dimensional fused feature vector that integrates the joint representation of identity, semantics, speech rhythm, and emotional state, effectively improving the richness and accuracy of the driving signal, and providing support for generating more expressive facial animations.
[0136] In step S303, a facial animation sequence of the target virtual avatar model is generated according to the fusion features in combination with a pre-stored facial animation template.
[0137] In this step, in an optional embodiment, the fusion features can be decoded into parameter instructions of multiple facial control channels, including mouth shape control, head movement, micro-expression intensity, etc., through a multi-branch decoding network, and animation sequence synthesis is performed in combination with a pre-stored personalized facial animation template. Specifically, the speech rhythm information (such as speech speed, stress) in the fusion features is used to drive the time alignment and amplitude change of the mouth opening and closing, to achieve accurate synchronization of lip shape and speech; the emotional semantic part is used to dynamically adjust the activation intensity and duration of micro-expression actions: for example, when a positive emotion is detected, the control coefficient of action unit AU12 (mouth corner up) is increased by 15%-20%, to enhance the expressiveness of a smile; when a negative emotion is identified, the intensity coefficient of AU4 (eyebrow up) or AU1 (forehead up) is increased by 10%-15%, to express worry or sadness. All action units can be parameterized combined based on time envelope and spatial deformation parameters, to generate natural and coherent facial animation sequences.
[0138] In another optional embodiment, to further enhance the eye contact ability and interactive realism of the virtual avatar, the current speaker's eye movement parameters can be collected. The parameters are extracted in real time through the terminal front camera combined with a lightweight eye tracking algorithm (such as based on the pupil-iris reflection method or a deep learning gaze estimation model), to obtain the two-dimensional coordinate trajectory of the user's gaze point. Subsequently, the eye movement parameters and the aforementioned fusion features are jointly used as driving signals, which are input into a facial animation synthesis module to be mapped into rotation angle control instructions of the digital person's eyeball, to achieve natural gaze movement within a horizontal direction of ±15° and a vertical direction of ±10°. By fusing visual attention information, the virtual avatar has the ability to "follow the gaze", accurately simulates the eye movement behavior in real interpersonal interaction in a conversation, and significantly enhances the immersion and emotional resonance. Finally, in combination with the action primitives defined in the facial animation template, a complete facial animation sequence including mouth shape, expression, head movement, and eyeball movement is generated, to realize high-fidelity, multi-modal collaborative driving of digital person expression.
[0139] In an optional embodiment, after generating the facial animation sequence of the target virtual avatar model, a body movement sequence can also be injected into the target virtual avatar model to enhance the overall expressiveness and interactive naturalness of the virtual avatar. The body movement sequence includes but is not limited to head movement, hand movement, and coordinated body movement of other parts of the body, which can be parameterized generated according to speech content, emotional state, and conversation rhythm.
[0140] Specifically, the head action can not only include nodding, shaking, and other semantic related actions driven by the speech rhythm and emotional features, but also superimpose low-frequency natural micro-motions to simulate the unconscious head swing of real humans during speech. For example, periodically superimposing horizontal rotation of ±5° left and right (simulating slight lateral head thinking) and pitch rotation of ±3° forward and backward (simulating head fluctuation under the rhythm of natural breathing) at a frequency of 0.5 Hz makes the virtual image action more lively and not stiff. This basic micro-motion is superimposed on the active nodding action triggered by the speech accent (such as slight nodding at the end of each sentence to indicate confirmation), forming a hierarchical head movement performance.
[0141] Further, hand and upper limb actions can be dynamically triggered according to semantic keywords (such as "big", "small", "start", "end") or emotional intensity, such as generating hand gesture opening action when expressing emphasis, and generating finger tapping or hands overlapping micro-motion when hesitating. All limb actions are parameterized calling based on a preset action primitive library, and are aligned with facial animation on the time axis to ensure that the lip shape, expression, eye contact, and body language are consistent, finally generating a full-body upper and lower body animation sequence covering face, head, and body, significantly improving the degree of personification and user interaction experience of digital people in weak network completion scenarios.
[0142] In step S140, the facial animation sequence is rendered in real time to generate a virtual image animation, and the virtual image animation is used to replace the real-time communication video stream.
[0143] In this step, the facial animation sequence can be rendered in real time to generate a virtual image animation, and the virtual image animation is used to replace the real-time communication video stream.
[0144] Specifically, a double-buffered rendering architecture can be used to render the facial animation sequence in real time to ensure that frame generation and frame display are decoupled asynchronously, avoid screen tearing, and maintain a stable output frame rate of 30 fps. After detecting that the communication quality does not meet the preset standard and the target virtual image model is loaded, a visual replacement process from the real-time communication video stream to the virtual image animation can be triggered. To achieve smooth transition and reduce the user's perception of abruptness, a Gaussian blur-based gradual fusion strategy is used: first, the last valid video frame of the real-time communication video stream is processed by Gaussian blur (kernel radius of 4 pixels) to generate a first processed picture; at the same time, the initial frame of the virtual image animation is processed by Gaussian blur with the same parameters to obtain a second processed picture; then, the first processed picture and the second processed picture are fused by linear interpolation or fade-in and fade-out to generate a first transition picture; finally, the first transition picture is used as a bridging medium to gradually transition to a clear virtual image animation, completing the seamless replacement of the real-time communication video stream.
[0145] In an alternative embodiment, when it is detected that the communication quality meets the preset quality standard, the recovered real-time communication video stream can also be used to replace the avatar animation.
[0146] Specifically, when the network transmission indicators meet the condition that the video stream packet loss rate is continuously lower than 5% and there is no freezing for 1 second, it can be determined that the communication quality has recovered, and at this time, the reverse switching process can be triggered. In order to avoid the visual abruptness caused by picture jumping, a Gaussian blur transition strategy symmetrical to the forward switching is adopted: first, the last valid frame of the avatar animation is subjected to Gaussian blur processing (kernel radius of 4 pixels) to generate a third processed picture; at the same time, the first clear picture of the recovered real-time communication video stream is subjected to Gaussian blur processing with the same parameters to obtain a fourth processed picture; then, the third processed picture and the fourth processed picture are gradually fused to generate a second transition picture; finally, the second transition picture is used as a visual bridge to gradually fade out the avatar animation and fade in the real video stream, realizing seamless switching from virtual to real.
[0147] In an alternative embodiment, after extracting the voiceprint features corresponding to the speech of the current speaker, the present disclosure also provides a scheme for updating the local avatar model library based on a new speaker. Referring to Figure 4 , Figure 4 A flowchart showing how to update the local avatar model library in the embodiments of the present disclosure is shown, including steps S401-S403:
[0148] In step S401, when it is detected that the communication quality meets the preset quality standard, it is determined whether the speech is from a new speaker.
[0149] In this step, when it is detected that the communication quality meets the preset quality standard, it can be determined whether the speech is from a new speaker.
[0150] Specifically, referring to Figure 5 , Figure 5 A flowchart showing how to determine whether the speech is from a new speaker in the embodiments of the present disclosure is shown, including steps S501-S502:
[0151] In step S501, the voiceprint features are encoded into a speaker representation vector.
[0152] In this step, the extracted voiceprint features can be mapped into high-dimensional vectors using a pre-trained deep voiceprint encoding model to generate a speaker representation vector with identity discrimination capability. Specifically, an ECAPA-TDNN (Emphasized Channel Attention, Propagation and Aggregation-Time Delay Neural Network) model can be used as a voiceprint encoder. The model effectively captures long-term temporal features related to speaker identity in speech through a multi-level feature aggregation mechanism. The input is the voiceprint features extracted from a speech segment of the current speaker lasting a certain period of time (e.g., 2-3 seconds). After forward inference by the model, a normalized speaker representation vector is output, which is used to represent the acoustic identity features of the user.
[0153] In step S502, the speaker representation vector is matched with the pre-stored reference speaker representation vector, and the similarity matching result is used to determine whether the speaking voice comes from a new speaker.
[0154] In this step, the current generated speaker representation vector can be used to calculate the cosine similarity with each registered identity vector in the local reference speaker representation vector library maintained by the terminal, obtaining a set of similarity scores. If the highest similarity score is lower than a pre-set similarity threshold (e.g., 0.75), it is determined that the current speaker is not registered locally and belongs to a new speaker. If the highest similarity score is greater than or equal to the similarity threshold, it is determined to be a known speaker, and the modeling update process does not need to be triggered. This mechanism effectively distinguishes between known users and potential new users by setting a reasonable similarity threshold, providing a decision basis for whether to start video collection and cloud modeling in the future, while avoiding resource waste caused by false triggering.
[0155] Reference is then made to Figure 4 In step S402, when it is determined that the speaking voice comes from a new speaker, video data of the new speaker is collected and uploaded to the cloud.
[0156] In this step, when it is determined that the speaking voice comes from a new speaker, 8 seconds of continuous video data can be collected and uploaded to the cloud.
[0157] Specifically, reference is made to Figure 6 , Figure 6 A flowchart showing how video data is uploaded to the cloud in the embodiments of the present disclosure is shown, including steps S601-S603:
[0158] In step S601, the video data is subjected to motion blur detection and brightness normalization processing to obtain a first processed video.
[0159] In this step, the video data can be preprocessed for quality enhancement to obtain a first processed video. Specifically, an energy function based on gradient amplitude variance or Laplacian operator can be used to evaluate the motion blur degree of the image frame by frame, and low-quality frames with a blur score below a threshold value can be removed. At the same time, the remaining frames are subjected to brightness normalization processing, and the light distribution is adjusted by histogram equalization or gamma correction method to eliminate the influence of overexposure or overexposure and improve the consistency of facial texture. After the above processing, a first processed video with stable image quality and clear face is generated, providing high-quality input for subsequent key frame extraction.
[0160] In step S602, key frame images are selected from the first processed video.
[0161] In this step, a clarity evaluation model based on the number and distribution density of SIFT (Scale-Invariant Feature Transform) feature points can be used to score each frame in the first processed video, and high-score frames with rich feature points, clear focus in the face region, and moderate pose changes are selected as candidate key frames. Further combined with the time uniform sampling strategy, the time interval is uniformly divided within 8 seconds of video length, and the frame with the highest SIFT score is preferentially selected in each interval, and finally 18 representative and quality-standard key frame images are selected. This method takes into account both image quality and facial pose diversity, ensuring the integrity of spatial information required for cloud modeling.
[0162] In step S603, image information is labeled for the key frame images, and the processed key frame images are uploaded to the cloud; wherein the image information includes the current timestamp and / or device identifier.
[0163] In this step, metadata tags can be added to each key frame image, and the image information includes at least the timestamp of the collection time, the device identifier of the collection device, and optional context information such as user ID and session number, which are used for cloud data tracing and identity association. All labeled key frame images are compressed and encoded (such as JPEG format), and then uploaded in batches to the cloud server through a secure transmission protocol (such as HTTPS or TLS encrypted channel). After uploading is completed, the local cache data can be automatically cleared to protect user privacy and security. The cloud receives the image set to build a personalized virtual avatar model and realize high-fidelity digital human generation.
[0164] In an optional embodiment, the cloud can generate the above virtual avatar model based on the following method. Referring to Figure 7 , Figure 7 A flowchart showing how the cloud generates a virtual avatar model in an embodiment of the present disclosure is shown, including steps S701-S704:
[0165] In step S701, face parameters of the new speaker are extracted based on the key frame images, and a face mesh of the new speaker is reconstructed based on the face parameters; the face parameters include a face geometry basis, face deformation control points, and expression parameters.
[0166] In this step, after receiving the key frame images and their metadata uploaded by the terminal, the cloud first aligns the image set through a multi-view face alignment algorithm (such as 3DMM fitting) to extract high-precision face features of the new speaker. A FLAME (Faces Learned with an Articulated Model and Expressions) learnable 3D face model is used as a basic mesh framework, and through a back-end rendering and optimization strategy, individualized face geometry bases, face deformation control points, and dynamic expression parameters (a total of 52 AU action unit coefficients) are jointly estimated from the 2D key frames. Based on the above parameters, a 3D face mesh model with unique identity and supporting expression driving is generated, providing a geometric basis for subsequent high-fidelity modeling.
[0167] In step S702, face neural modeling is performed on the face mesh to generate an initial virtual avatar model.
[0168] In this step, the reconstructed 3D face mesh can be used as a structural prior, combined with the multi-view key frame images uploaded by the terminal, and a face-specific modeling method based on NeRF (Neural Radiance Fields) (such as FaceNeRF or PixelNeRF) is used for neural implicit field modeling. By inputting the image set with pose estimation, a radiance field network is trained to learn the mapping relationship from 3D spatial coordinates and viewing direction to color and density, achieving fine restoration of high-order visual attributes such as skin texture, hair details, and light reflection. After training, the radiance field network is bound to the FLAME mesh to generate an initial virtual avatar model with photo-realistic realism, supporting arbitrary view rendering and dynamic lighting simulation.
[0169] In step S703, a speech-driven model is constructed according to the acoustic characteristics of the new speaker combined with the expression parameters; the speech-driven model is used to define the mapping relationship between the acoustic characteristics and the face animation parameters; the face animation parameters include mouth opening degree and mouth corner offset.
[0170] In this step, the LSTM (Long Short-Term Memory) network can be used to model the timing of the speech segments corresponding to the new speaker. First, the mel-spectral features and the fundamental frequency (F0), energy envelope, and other acoustic features of the speech are extracted as input sequences; at the same time, the expression parameter sequence in the corresponding time period is extracted, especially the AU06 (buccinator contraction) and AU12 (mouth corner lifting) coefficients related to the lip shape, which are mapped to continuous controllable parameters such as lip opening degree (0-100%) and mouth corner offset (±5mm). Through timing alignment and supervised training, an end-to-end mapping relationship from speech features to facial animation parameters is established to form a personalized speech-driven model, realizing the natural lip synchronization capability of "speak and move".
[0171] In step S704, the initial virtual avatar model is fused with the speech-driven model to generate a virtual avatar model.
[0172] In this step, the neural rendering network in the initial virtual avatar model can be integrated with the speech-driven model. Specifically, the parameters such as lip opening degree and mouth corner offset output by the speech-driven model can be used as control signals to dynamically modulate the expression coefficients of the FLAME grid and drive the NeRF model to generate realistic pictures in the corresponding lip shape state during rendering.
[0173] In an optional implementation, to further improve the deployment efficiency of the terminal, the fused model can be subjected to lightweight processing: 8-bit weight quantization is used to reduce the parameter precision, and the knowledge distillation technology with MobileNet as the teacher model is used to migrate the expression capability of the large model to the lightweight student network, finally generating a lightweight digital human model with a volume of less than 75MB and supporting end-side real-time inference. The model is delivered to the terminal through an encrypted channel and cached locally, and used for subsequent dynamic video completion under speech driving.
[0174] Next, referring to Figure 4 In step S403, the preset virtual avatar model library is updated according to the virtual avatar model generated by the cloud based on the video data.
[0175] In this step, the terminal can receive the virtual avatar model described above, and update the preset virtual avatar model library based on the virtual avatar model.
[0176] Specifically, referring to Figure 8 , Figure 8 A flowchart for updating the preset virtual avatar model library according to the virtual avatar model in the embodiments of the present disclosure is shown, including steps S801-S803:
[0177] In step S801, the model file of the encrypted virtual avatar model sent by the cloud is received.
[0178] In this step, the virtual avatar model file generated and encrypted by the cloud can be received. The model file can include lightweight 3D face mesh parameters, neural rendering weights, and voice-driven mapping models. At the same time, the model file header carries an identity ID bound to the speaker identity, which is used to identify the model ownership.
[0179] In step S802, the decrypted model file is stored in the preset virtual avatar model library, and the index identification corresponding to the model file is recorded.
[0180] In this step, the terminal can perform decryption operation on the encrypted model file to restore the original model data. Then the decrypted model file can be stored in the local preset virtual avatar model library in the form of file or database entry, and a unique index identification such as self-incrementing number or UUID is assigned to the model.
[0181] In step S803, the hash value corresponding to the voiceprint feature is calculated, and the mapping relationship between the hash value and the index identification is stored.
[0182] In this step, the terminal can call a hash algorithm (such as SHA-256 or lightweight Hash function) to encode the voiceprint feature vector of the current new speaker, generating a fixed-length voiceprint hash value. Then the corresponding relationship between the hash value and the model file index identification assigned in step S802 is established and persistently stored in the local mapping table. Through this mechanism, when the speaker's voice is detected subsequently, only the voiceprint feature and the hash value need to be extracted to quickly locate the corresponding virtual avatar model, realize low-delay and offline model matching and loading, and improve the response efficiency and user experience in weak network environment.
[0183] Reference Figure 9 , Figure 9 The overall flowchart of the video communication method in the embodiment of the present disclosure is shown, which includes steps S901-S912:
[0184] In step S901, start;
[0185] In step S902, real-time or timed analysis of voice content to extract voiceprint features; specifically, the terminal collects input voice signals in real time or at a fixed time, and extracts the voiceprint features of the speaker through advanced voiceprint recognition technology such as ECAPA-TDNN model. The feature vector is used for subsequent new speaker detection and identity matching process to ensure accurate identification of the speaker's identity;
[0186] In step S903, it is judged whether the communication quality is good; specifically, the current network communication quality is monitored in real time, including bandwidth, delay, packet loss rate and other key indicators, to evaluate whether it meets the preset quality standard (such as minimum bandwidth threshold, maximum delay limit). If the communication quality is good, go to step S904; otherwise, jump to step S908 to start the local digital human model matching process.
[0187] In step S904, in the case of a new speaker, video data is collected, key frames are extracted and uploaded to the cloud;
[0188] In step S905, the cloud generates a lightweight virtual avatar model based on the key frames;
[0189] In step S906, the cloud delivers the compressed model to the terminal;
[0190] In step S907, the terminal caches the model and updates the virtual avatar model library;
[0191] In step S908, load the model to drive lip movement / head movement to generate animation; specifically, the terminal matches and loads the corresponding virtual avatar model from the local model library according to the voiceprint characteristics of the current speaker, then combines the real-time voice input to dynamically adjust the lip opening degree, mouth corner offset and head movement trajectory using the voice-driven model, to generate a coherent and natural facial animation sequence;
[0192] In step S909, the animation is smoothly transitioned with the original video stream;
[0193] In step S910, it is judged whether the communication quality is restored;
[0194] In step S911, if it is restored, seamlessly switch to the restored communication video stream;
[0195] In step S912, end.
[0196] Based on the above technical solution, the present disclosure has at least the following technical effects:
[0197] First, by adopting the end-cloud collaborative digital human generation architecture, efficient resource allocation is achieved through "cloud generation + end rendering". When the communication quality meets the preset quality standard, the cloud uses high-performance computing to process the key frame image of the new speaker, reconstructs the face grid based on facial parameters, and generates an initial virtual avatar model through facial neural modeling, then combines with the speech characteristics to build a speech-driven model, fuses a lightweight virtual avatar model and delivers it to the terminal. When the communication quality does not meet the preset quality standard, the terminal matches and loads the target virtual avatar model based on the local cached virtual avatar model library, without relying on real-time cloud computing, avoiding video completion delay or failure caused by network anomalies, and ensuring continuous generation and smooth replacement of virtual avatar animation.
[0198] Second, through the on-demand trigger mechanism driven by voice, efficient resource utilization and precise identity matching are achieved. Video data collection and cloud modeling process are only triggered when a new speaker is detected. Specifically, the voiceprint features are encoded into speaker representation vectors, and similarity matching is performed with reference speaker representation vectors. When the similarity does not meet the preset similarity threshold, it is determined to be a new speaker, and then the camera collects video and uploads it to the cloud. At the same time, after the model is delivered, a mapping relationship between the hash value corresponding to the voiceprint feature and the virtual avatar model file index identifier is established, realizing fast model retrieval and loading based on voiceprint hash, and improving the accuracy and response efficiency of identity recognition.
[0199] Third, through the lightweight local rendering strategy, it adapts to the deployment needs of mid-low end terminal devices. The terminal pre-stores virtual avatar models and associated face animation templates locally, generates face animation sequences based on the rhythm and prosody features extracted from the speech, combines voiceprint features with face animation templates to generate face animation sequences, and renders virtual avatar animation in real time. Through double-buffer rendering technology, a stable frame rate of 30fps is guaranteed, and Gaussian blur transition technology is used in the replacement process to achieve seamless switching with real-time communication video stream, significantly reducing local computing load while ensuring smoothness and continuity of the picture.
[0200] Fourth, through the multi-modal emotion fusion driving scheme, the expression realism of digital human animation and user interaction satisfaction are improved. When generating face animation sequences, not only the rhythm and prosody features of speech, but also text emotion features are introduced, and fusion features are obtained through multi-modal feature fusion to drive micro-expression parameters. For example, according to the emotion polarity, adjust the motion intensity of AU12 (mouth corner up) or AU4 (eyebrow up); At the same time, combined with the eye movement parameters obtained through eye tracking, drive the digital human eyeball to rotate within the range of horizontal ±15° and vertical ±10°, realize the effect of gaze following; This scheme makes the expression and action of virtual avatar more consistent with the actual dialogue situation, and experimental verification can significantly improve user interaction satisfaction.
[0201] The present disclosure also provides a video communication device, Figure 10 A schematic diagram of a video communication device in an exemplary embodiment of the present disclosure is shown. As shown in the figure, Figure 10 The video communication device 1000 can include a first feature extraction module 1010, a model loading module 1020, a second feature extraction module 1030, an animation generation module 1040, and a model updating module 1050. Among them:
[0202] The first feature extraction module 1010 is configured to extract a voiceprint feature corresponding to a speech voice of a current speaker.
[0203] The model loading module 1020 is configured to, when it is detected that the communication quality does not meet the preset quality standard, match and load a target virtual image model corresponding to the current speaker from a preset virtual image model library based on the voiceprint feature.
[0204] The second feature extraction module 1030 is configured to extract a rhythm and prosody feature corresponding to the speech voice, and generate a face animation sequence of the target virtual image model based on the voiceprint feature, the rhythm and prosody feature, and a pre-stored face animation template.
[0205] The animation generation module 1040 is configured to render a virtual image animation by real-time rendering the face animation sequence, and replace a real-time communication video stream based on the virtual image animation.
[0206] In an exemplary embodiment of the present disclosure, the model loading module 1020 determines whether the communication quality meets the preset quality standard by the following method:
[0207] Monitoring a network transmission indicator of a communication link; the network transmission indicator includes at least one of a data packet loss rate, a transmission delay, a jitter, and a bandwidth utilization rate;
[0208] According to the comparison result between the network transmission indicator and a preset indicator threshold, it is determined whether the communication quality meets the preset quality standard.
[0209] In an exemplary embodiment of the present disclosure, the model loading module 1020 matches and loads a target virtual image model corresponding to the current speaker from a preset virtual image model library based on the voiceprint feature, including:
[0210] Calculating a hash value corresponding to the voiceprint feature;
[0211] According to the correspondence between the hash value and a pre-set hash value and file index of a virtual image model, the target virtual image model corresponding to the current speaker is matched and loaded.
[0212] In the example embodiment of the present disclosure, the second feature extraction module 1030 extracts the rhythm and prosody features corresponding to the spoken voice, including:
[0213] dividing the spoken voice into a plurality of unit-length voice segments;
[0214] extracting the voice energy, stress position and speech rate variation features corresponding to each voice segment based on Fourier transform;
[0215] wherein the voice energy and the mouth shape control parameter of the target virtual image model have a preset mapping relationship.
[0216] In the example embodiment of the present disclosure, after extracting the rhythm and prosody features corresponding to the spoken voice, the second feature extraction module 1030 is configured to:
[0217] extracting text sentiment features and acoustic features corresponding to the spoken voice; the text sentiment features and the micro-expression control parameter of the target virtual image model have a preset mapping relationship;
[0218] performing multi-modal feature fusion on the text sentiment features, the acoustic features, the voiceprint features, and the rhythm and prosody features to obtain fusion features;
[0219] generating a facial animation sequence of the target virtual image model according to the fusion features combined with the pre-stored facial animation template.
[0220] In the example embodiment of the present disclosure, the second feature extraction module 1030 generates a facial animation sequence of the target virtual image model according to the fusion features combined with the pre-stored facial animation template, including:
[0221] collecting eye movement parameters of the current speaker;
[0222] generating a facial animation sequence of the target virtual image model according to the fusion features, the eye movement parameters combined with the pre-stored facial animation template.
[0223] In the example embodiment of the present disclosure, after generating the facial animation sequence of the target virtual image model, the second feature extraction module 1030 is configured to:
[0224] injecting a body action sequence into the target virtual image model; the body action sequence includes at least one of the following: head action, hand action, and body action.
[0225] In the example embodiment of the present disclosure, the animation generation module 1040 generates a virtual image animation to replace a real-time communication video stream based on the virtual image animation, including:
[0226] Gaussian blur processing is performed on the last valid video frame of the real-time communication video stream to obtain a first processed picture;
[0227] Gaussian blur processing is performed on the initial frame of the virtual image animation to obtain a second processed picture;
[0228] According to the first processed picture and the second processed picture, a first transition picture is generated;
[0229] The first transition picture and the virtual image animation are used to replace the real-time communication video stream.
[0230] In an example embodiment of the present disclosure, the animation generation module 1040 is configured to:
[0231] When it is detected that the communication quality meets the preset quality standard, the virtual image animation is replaced with the recovered real-time communication video stream.
[0232] In an example embodiment of the present disclosure, the animation generation module 1040 replaces the virtual image animation with the recovered real-time communication video stream, including:
[0233] Gaussian blur processing is performed on the last valid frame of the virtual image animation to obtain a third processed picture;
[0234] Gaussian blur processing is performed on the initial frame of the recovered real-time communication video stream to obtain a fourth processed picture;
[0235] According to the third processed picture and the fourth processed picture, a second transition picture is generated;
[0236] The second transition picture and the recovered real-time communication video stream are used to replace the virtual image animation.
[0237] In an example embodiment of the present disclosure, after extracting the voiceprint features corresponding to the speech of the current speaker, the model updating module 1050 is configured to:
[0238] When it is detected that the communication quality meets the preset quality standard, it is determined whether the speech is from a new speaker;
[0239] In a case where it is determined that the speech is from the new speaker, video data of the new speaker is collected and uploaded to the cloud;
[0240] According to a virtual image model generated by the cloud based on the video data, the preset virtual image model library is updated.
[0241] In the example embodiment of the present disclosure, the model updating module 1050 determines whether the speech voice is from a new speaker, comprising:
[0242] encoding the voiceprint feature into a speaker representation vector;
[0243] performing similarity matching between the speaker representation vector and a pre-stored reference speaker representation vector, and determining whether the speech voice is from a new speaker according to the similarity matching result.
[0244] In the example embodiment of the present disclosure, the model updating module 1050 determines whether the speech voice is from a new speaker according to the similarity matching result, comprising:
[0245] if the similarity matching result does not satisfy a preset similarity threshold, determining that the speech voice is from a new speaker;
[0246] if the similarity matching result satisfies the preset similarity threshold, determining that the speech voice is not from the new speaker.
[0247] In the example embodiment of the present disclosure, the video data comprises continuous video data of a preset time length;
[0248] The model updating module 1050 uploads the video data to the cloud, comprising:
[0249] performing motion blur detection and brightness normalization processing on the video data to obtain a first processed video;
[0250] selecting key frame images from the first processed video;
[0251] labeling image information for the key frame images, and uploading the processed key frame images to the cloud; wherein the image information comprises a current timestamp and / or a device identifier.
[0252] In the example embodiment of the present disclosure, the cloud generates the virtual avatar model based on the following manner:
[0253] extracting facial parameters of the new speaker based on the key frame images, and reconstructing a face mesh of the new speaker based on the facial parameters; the facial parameters comprise facial geometric basis, facial deformation control points, and expression parameters;
[0254] performing face neural modeling on the face mesh to generate an initial virtual avatar model;
[0255] constructing a speech-driven model according to the acoustic characteristics of the new speaker combined with the expression parameters; the speech-driven model is used to define the mapping relationship between the acoustic characteristics and the facial animation parameters; the facial animation parameters comprise mouth opening degree and mouth corner offset.
[0256] fusing the initial virtual image model with the voice-driven model to generate the virtual image model.
[0257] In the example embodiment of the present disclosure, the model updating module 1050 updates the preset virtual image model library according to the virtual image model generated by the cloud based on the video data, including:
[0258] receiving the model file of the encrypted virtual image model sent by the cloud;
[0259] storing the decrypted model file into the preset virtual image model library, and recording the index identifier corresponding to the model file;
[0260] calculating the hash value corresponding to the voiceprint feature, and storing the mapping relationship between the hash value and the index identifier.
[0261] The specific details of the modules in the above video communication device have been described in detail in the corresponding video communication method, and therefore will not be described here.
[0262] It should be noted that, although several modules or units of the device for action execution are mentioned in the above detailed description, such division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into several modules or units.
[0263] In addition, although the steps of the method in the present disclosure are described in a specific order in the drawings, this does not require or imply that the steps must be performed in this specific order, or that all the steps shown must be performed to achieve the desired result. In addition or alternatively, some steps can be omitted, a plurality of steps can be combined into one step, and / or one step can be divided into a plurality of steps, etc.
[0264] From the above description of the embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or by software in combination with necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a U disk, a mobile hard disk, etc.) or a network, and includes a plurality of instructions to make a computing device (which can be a personal computer, a server, a mobile terminal, or a network device, etc.) execute the method according to the embodiments of the present disclosure.
[0265] The present disclosure also provides a computer readable storage medium, which can be included in the electronic device described in the above embodiments, or can exist separately without being assembled into the electronic device.
[0266] The computer readable storage medium may, for example, be, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the above. More specific examples of the computer readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer readable storage medium can be any tangible medium that contains or stores a program used or used with an instruction execution system, apparatus, or device.
[0267] The computer readable storage medium can send, propagate or transfer a program for use by or in connection with an instruction execution system, apparatus, or device. Program code contained on a computer readable storage medium can be transmitted using any suitable medium, including but not limited to wireless, wire line, optical fiber cable, RF, etc., or any suitable combination of the above.
[0268] The computer readable storage medium carries one or more programs, which, when executed by an electronic device, cause the electronic device to implement the methods described in the above embodiments.
[0269] In addition, an electronic device capable of implementing the above method is also provided in the embodiments of the present disclosure.
[0270] Those skilled in the art can understand that each aspect of the present disclosure can be implemented as a system, a method or a program product. Therefore, each aspect of the present disclosure can be specifically implemented as follows: a complete hardware embodiment, a complete software embodiment (including firmware, microcode, etc.), or an embodiment combining hardware and software aspects, which can be collectively referred to as "circuitry", "module" or "system" here.
[0271] The electronic device 1100 according to this embodiment of the present disclosure will be described below with reference to Figure 11 Figure 11 The displayed electronic device 1100 is only an example and should not impose any limitation on the functions and use range of the embodiments of the present disclosure.
[0272] As Figure 11 As shown, the electronic device 1100 is in the form of a general computing device. Components of the electronic device 1100 can include, but are not limited to, at least one processor 1110, at least one memory 1120, a bus 1130 connecting different system components including the memory 1120 and the processor 1110, a display 1140.
[0273] The memory stores program codes which can be executed by the processor 1110, so that the processor 1110 performs the steps according to various exemplary embodiments of the present disclosure described in the above “Exemplary Method” section of the present specification. For example, the processor 1110 can perform the steps as shown in FIG. 11, i.e., step S110, extracting a voiceprint feature corresponding to a speaking voice of a current speaker when detecting that a communication quality does not meet a preset quality standard; step S120, matching and loading a target virtual image model corresponding to the current speaker from a preset virtual image model library based on the voiceprint feature; step S130, extracting a rhythm and prosody feature corresponding to the speaking voice, generating a face animation sequence of the target virtual image model based on the voiceprint feature, the rhythm and prosody feature, and the pre-stored face animation template; and step S140, rendering a virtual image animation in real time by the face animation sequence, and taking over a real-time communication video stream based on the virtual image animation. Figure 1
[0274] The memory 1120 can include a readable medium in the form of volatile storage, such as a random access memory (RAM) 11201 and / or a cache memory 11202, and can further include a read-only memory (ROM) 11203.
[0275] The memory 1120 can further include program / utility 11204 having a set of programs / modules 11205, including, but not limited to, an operating system, one or more application programs, other program modules, and program data, and each of these examples, or some combination thereof, can include implementation of a network environment.
[0276] The bus 1130 can represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor or local bus using any of a variety of bus architectures.
[0277] The electronic device 1100 can also communicate with one or more external devices 1200 such as a keyboard, a pointing device, a Bluetooth device, etc.; and one or more devices that enable a user to interact with the electronic device 1100; and / or one or more devices (e.g. routers, modems, etc.) that enable the electronic device 1100 to communicate with one or more other computing devices. Such communication can occur via Input / Output (I / O) interface 1150. Still yet, the electronic device 1100 can communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or the Internet) through network adapter 1160. As depicted, network adapter 1160 communicates with the other components of electronic device 1100 via bus 1130. It should be appreciated that although not shown, other hardware and / or software modules could be used in conjunction with the electronic device 1100. Such modules include, but are not limited to, microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archival storage systems, etc.
[0278] Other embodiments of the disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the features disclosed herein. The disclosure is intended to cover any variations, uses, or adaptations of the disclosure following, in general, the principles of the disclosure and including such features to the extent that they are not disclosed in the prior art. The specification and examples are to be regarded as illustrative only, and the true scope and spirit of the disclosure are indicated by the appended claims.
Claims
1. A video communication method, characterized in that, include: Extract the voiceprint features corresponding to the current speaker's speech; When the communication quality is detected to be unsatisfactory according to the preset quality standard, the target virtual avatar model corresponding to the current speaker is matched and loaded from the preset virtual avatar model library based on the voiceprint features. Extract the rhythm and prosody features corresponding to the spoken voice, and generate a facial animation sequence for the target virtual image model based on the voiceprint features, the rhythm and prosody features, and a pre-stored facial animation template. The facial animation sequence is rendered in real time to generate a virtual character animation, which then replaces the real-time communication video stream.
2. The method according to claim 1, characterized in that, Whether the communication quality meets the preset quality standard is determined by the following methods: Monitor network transmission metrics of the communication link; the network transmission metrics include at least one of packet loss rate, transmission delay, jitter, and bandwidth utilization. Based on the comparison between the network transmission index and the preset index threshold, it is determined whether the communication quality meets the preset quality standard.
3. The method according to claim 1, characterized in that, The step of matching and loading the target virtual avatar model corresponding to the current speaker from a preset virtual avatar model library based on the voiceprint features includes: Calculate the hash value corresponding to the voiceprint feature; Based on the hash value and the pre-set correspondence between the hash value and the file index of the virtual avatar model, the target virtual avatar model corresponding to the current speaker is matched and loaded.
4. The method according to claim 1, characterized in that, The extraction of rhythmic and prosodic features corresponding to the spoken speech includes: The spoken voice is divided into multiple voice segments of unit duration; Based on Fourier transform, the speech energy, stress position, and speech rate variation features corresponding to each speech segment are extracted; The speech energy and the lip-sync control parameters of the target virtual avatar model have a preset mapping relationship.
5. The method according to any one of claims 1 to 4, characterized in that, After extracting the rhythmic and prosodic features corresponding to the spoken speech, the method further includes: Extract the textual emotional features and acoustic features corresponding to the spoken voice; the textual emotional features and the micro-expression control parameters of the target virtual avatar model have a preset mapping relationship; Multimodal feature fusion is performed on the text sentiment features, acoustic features, voiceprint features, and rhythmic features to obtain fused features; Based on the fusion features and the pre-stored facial animation template, a facial animation sequence of the target virtual character model is generated.
6. The method according to claim 5, characterized in that, The step of generating a facial animation sequence for the target virtual character model based on the fusion features and the pre-stored facial animation template includes: Collect the eye movement parameters of the current speaker; Based on the fusion features, the eye movement parameters, and the pre-stored facial animation template, a facial animation sequence for the target virtual character model is generated.
7. The method according to claim 1, characterized in that, After generating the facial animation sequence of the target virtual avatar model, the method further includes: Injecting a sequence of body movements into the target virtual avatar model; the sequence of body movements includes at least one of the following: head movements, hand movements, and movements of other parts of the body.
8. The method according to claim 1, characterized in that, The method of replacing the real-time communication video stream with the virtual character animation includes: Gaussian blurring is applied to the last valid video frame of the real-time communication video stream to obtain the first processed image; The initial frame of the virtual character animation is subjected to Gaussian blur processing to obtain the second processed image; A first transition screen is generated based on the first processing screen and the second processing screen; The first transition screen and the virtual character animation are used to replace the real-time communication video stream.
9. The method according to claim 1, characterized in that, The method further includes: When the communication quality is detected to meet the preset quality standard, the restored real-time communication video stream is used to replace the virtual character animation.
10. The method according to claim 9, characterized in that, The step of replacing the virtual character animation with the restored real-time communication video stream includes: Gaussian blur is applied to the last valid frame of the virtual character animation to obtain the third processed screen; Gaussian blurring is applied to the initial frame of the recovered real-time communication video stream to obtain the fourth processed frame; A second transition screen is generated based on the third and fourth processing screens; The virtual character animation is replaced by the second transition screen and the restored real-time communication video stream.
11. The method according to claim 1, characterized in that, After extracting the voiceprint features corresponding to the current speaker's speech, the method further includes: When the communication quality is detected to meet the preset quality standard, it is determined whether the spoken voice comes from a new speaker; If it is determined that the spoken voice comes from the new speaker, the video data of the new speaker is collected and uploaded to the cloud; The preset virtual avatar model library is updated based on the virtual avatar model generated in the cloud based on the video data.
12. The method according to claim 11, characterized in that, The determination of whether the spoken voice comes from a new speaker includes: The voiceprint features are encoded into a speaker representation vector; The speaker representation vector is matched with a pre-stored reference speaker representation vector for similarity, and the similarity matching result is used to determine whether the spoken speech comes from a new speaker.
13. The method according to claim 12, characterized in that, Determining whether the spoken voice comes from a new speaker based on the similarity matching result includes: If the similarity matching result does not meet the preset similarity threshold, it is determined that the spoken voice comes from a new speaker; If the similarity matching result meets the preset similarity threshold, it is determined that the spoken voice does not come from the new speaker.
14. The method according to claim 11, characterized in that, The video data includes continuous video data of a preset duration; Uploading the video data to the cloud includes: The video data is subjected to motion blur detection and brightness normalization to obtain the first processed video; Select keyframe images from the first processed video; Image information is labeled for the keyframe image, and the processed keyframe image is uploaded to the cloud; wherein, the image information includes the current timestamp and / or device identifier.
15. The method according to claim 14, characterized in that, The cloud platform generates the virtual avatar model based on the following method: The facial parameters of the new speaker are extracted based on the keyframe image, and the facial mesh of the new speaker is reconstructed based on the facial parameters; the facial parameters include facial geometric basis, facial deformation control points, and expression parameters. Facial neural network modeling is performed on the facial mesh to generate an initial virtual avatar model; A speech-driven model is constructed based on the acoustic features of the new speaker and the facial expression parameters; the speech-driven model is used to define the mapping relationship between acoustic features and facial animation parameters; the facial animation parameters include mouth opening degree and corner of mouth offset; The initial virtual avatar model is fused with the voice-driven model to generate the virtual avatar model.
16. The method according to claim 11, characterized in that, The step of updating the preset virtual avatar model library based on the virtual avatar model generated in the cloud based on the video data includes: Receive the encrypted model file of the virtual avatar model sent from the cloud; The decrypted model file is stored in the preset virtual avatar model library, and the index identifier corresponding to the model file is recorded; Calculate the hash value corresponding to the voiceprint feature and store the mapping relationship between the hash value and the index identifier.
17. A video communication device, characterized in that, include: The first feature extraction module is used to extract the voiceprint features corresponding to the current speaker's speech. The model loading module is used to match and load the target virtual image model corresponding to the current speaker from the preset virtual image model library based on the voiceprint features when the communication quality is detected to be unsatisfactory. The second feature extraction module is used to extract the rhythm and prosody features corresponding to the spoken voice, and generate a facial animation sequence of the target virtual image model based on the voiceprint features, the rhythm and prosody features and a pre-stored facial animation template. An animation generation module is used to render the facial animation sequence in real time to generate a virtual character animation, which then replaces the real-time communication video stream.
18. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the video communication method according to any one of claims 1 to 16.
19. An electronic device, characterized in that, include: processor; as well as Memory for storing the executable instructions of the processor; The processor is configured to execute the video communication method according to any one of claims 1 to 16 by executing the executable instructions.