Method and system for matching facial expressions and actions of digital cat puppet image through human voice

Through the voice-driven digital cat doll facial expression generation method, combined with emotion analysis and 3D modeling technology, the problem of insufficient accuracy and coherence of emotion analysis in the existing technology is solved, and a natural and real-time virtual character interaction experience is achieved, which improves the user experience and application scope.

CN120581023APending Publication Date: 2025-09-02CHINA UNICOM WO MUSIC & CULTURE CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510401092.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

The prior art has problems such as insufficient accuracy of emotion analysis, poor coherence of action, insufficient real-timeness and inconsistent multi-view rendering in the voice-driven expression generation of digital cat puppets, which affects the user experience and the application effect of virtual characters in complex interactive scenarios.

Method used

Through voice input and preprocessing, speech recognition, sentiment analysis, facial action mapping and real-time rendering technology, combined with 3D modeling tools and game engines, the natural, coherent and real-time action generation of digital cat puppet facial expressions is achieved, and the system response speed is optimized using multi-threaded and asynchronous programming technology, and iterative optimization is performed through user feedback.

Benefits of technology

It realizes a natural, real-time and emotionally driven virtual character interaction experience, improves users' immersion and realism in virtual interaction scenarios, expands the application scope of virtual characters, and provides higher flexibility and efficiency for digital content creation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120581023A_ABST
    Figure CN120581023A_ABST
Patent Text Reader

Abstract

The invention discloses a method for matching facial expressions and actions of digital cat puppet images through human voice, which comprises the following steps of: S1, acquiring voice data of a user, and preprocessing the voice data; s2, converting the preprocessed voice data into text data, performing sentiment analysis, and extracting sentiment information; s3, constructing an action library, pre-defining various facial actions of the digital cat puppet, and distributing an emotion label for each action; according to a sentiment analysis result, selecting proper facial expressions and actions for mapping; and S4, generating a corresponding facial action according to a mapping result, and rendering the generated facial action to a model of the digital cat puppet in real time. According to the method, dynamic interaction between a virtual character and user voice input is realized through voice emotion driven digital cat puppet face action matching and real-time rendering, based on voice preprocessing and emotion analysis technologies, emotion information in voice can be accurately extracted, and the naturalness and coherence of digital cat puppet face actions are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to artificial intelligence-driven virtual character interaction technology, specifically to speech emotion recognition, emotion-driven facial action mapping, and real-time animation rendering technology methods, and is mainly used in the dynamic expression generation and interaction system of virtual characters. Background Art

[0002] A digital puppet is a virtual character created through technologies such as computer graphics, motion capture, image rendering, and artificial intelligence. A digital puppet possesses human-like appearance, behavior, and interactive capabilities, existing in virtual environments such as computer screens, virtual reality (VR), or augmented reality (AR) devices. Its appearance, movements, and voice can be customized to suit different needs and scenarios. Technological advancements have enabled digital puppets to mimic other virtual puppet forms, such as digital cat puppets. Existing digital cat puppets typically rely on core technologies such as speech recognition, emotion analysis, and expression mapping to achieve voice-driven expression changes. For example, common existing approaches include deep learning-based speech emotion recognition (SER), which identifies emotions by extracting features from speech signals (such as Mel-frequency cepstral coefficients). Additionally, some technologies utilize audio-driven facial animation generation, such as Nvidia's Audio2Face technology, which converts audio into facial animation using pre-trained neural networks. This allows for real-time animation generation and can guide virtual characters to express key emotions.

[0003] However, the existing technology still has some obvious shortcomings and technical problems. First, the accuracy of emotion analysis in the existing technology still needs to be improved, especially when dealing with complex emotions or speech with large background noise, the accuracy and robustness of emotion recognition are insufficient. Secondly, the existing technology often lacks naturalness and consistency when mapping emotion recognition results to the facial movements of virtual characters. For example, although the model-based generation method can control the movement parameters, it still has shortcomings in synthesizing complete talking head videos under full 3D head control. In addition, real-time and movement consistency are also pain points of the existing technology. Since movements are usually generated in segments, there may be discontinuities between movements, which affects the user experience. At the same time, the existing technology also has shortcomings in multi-perspective rendering and dynamic adaptability, resulting in inconsistent performance of virtual characters under different perspectives.

[0004] In summary, existing technologies have significant shortcomings in terms of the naturalness of virtual character expressions, the accuracy of emotion analysis, the coherence of movements, and real-time performance, which limits the effectiveness of virtual characters in complex interactive scenarios. Therefore, it is necessary to improve and optimize the algorithmic processing of digital cat puppets' facial expressions and movements. Summary of the Invention

[0005] In view of the above problems, the present invention provides a method for matching facial expressions of digital cat images through human voice, comprising the following steps:

[0006] S1, voice input and preprocessing, collects user voice data and performs voice data preprocessing;

[0007] S2, speech recognition, converts the pre-processed speech data into text data, performs sentiment analysis, and extracts sentiment information from the text data;

[0008] S3, facial action mapping, builds an action library, pre-defines various facial actions of the digital cat (such as blinking, smiling, frowning, etc.), and assigns an emotion label to each action; based on the results of emotion analysis, selects appropriate facial expression actions for mapping;

[0009] S4, facial action generation and rendering, generates corresponding facial actions according to the mapping results, and renders the generated facial actions onto the digital cat puppet model in real time.

[0010] As a further illustration of the present invention, in step S1, a Python library such as pyaudio or sounddevice is used for voice acquisition; and librosa or pydub is used for audio preprocessing.

[0011] Furthermore, in step S2, speech-to-text conversion is performed using a speech recognition API such as Google Speech-to-Text, Microsoft Azure Speech Service, or an open source tool such as DeepSpeech; and sentiment analysis is performed using a natural language processing (NLP) tool such as spaCy, NLTK, or a pre-trained sentiment analysis model (such as BERT).

[0012] Furthermore, in step S3, a facial action library of the digital cat puppet is created using a 3D modeling tool such as Blender or Maya; and the emotion tags are mapped to the actions in the action library using a Python script or configuration file.

[0013] Furthermore, in step S4, a game engine such as Unity or Unreal Engine is used to generate and render facial movements; and skeletal animation or blendshape technology is used to achieve detailed control of facial movements.

[0014] Furthermore, in step S4, user feedback is collected and iterative optimization is performed.

[0015] On the other hand, the present invention also provides a system for matching facial expressions and movements of a digital cat puppet through human voice, comprising a voice input and preprocessing module, a voice recognition module, a facial movement mapping module, a movement generation module and a real-time feedback module;

[0016] The voice input and preprocessing module is used to collect the user's voice signal and preprocess and optimize the voice signal to provide high-quality audio data for subsequent processing;

[0017] The speech recognition module is used to convert the collected speech signals into text data, and perform sentiment analysis on the text data to extract sentiment information;

[0018] The facial action mapping module is used to build a facial action library and select appropriate facial expression actions for mapping based on the results of emotion analysis;

[0019] The action generation module is used to generate corresponding facial actions according to the mapping results, and render the generated facial actions on the model of the digital cat puppet in real time;

[0020] The real-time feedback module is used to collect user feedback on the action matching effect and iteratively optimize the system.

[0021] Furthermore, the voice input and preprocessing module, voice recognition module, facial action mapping module, action generation module and real-time feedback module are integrated into a complete system, using message queues (such as RabbitMQ) or event-driven architecture (such as Node.js) for communication between modules, and using multi-threading or asynchronous programming technology to improve the response speed of the system.

[0022] Furthermore, the system is deployed to a target platform (such as a PC, mobile device, or cloud) using Docker, and continuous integration and continuous deployment are performed using CI / CD tools such as Jenkins or GitHub Actions.

[0023] Beneficial effects of the present invention:

[0024] The present invention realizes dynamic interaction between virtual characters and user voice input through voice emotion-driven digital cat puppet facial movement matching and real-time rendering. Based on voice preprocessing and emotion analysis technology, it can accurately extract emotional information from voice; based on the emotion-to-action mapping mechanism combined with 3D modeling and real-time rendering technology, it ensures the naturalness and consistency of the digital cat puppet's facial movements and ensures real-time interactive performance, thereby realizing a natural, real-time and emotion-driven virtual character interaction experience, enhancing the user's immersion and sense of reality in virtual interactive scenes, expanding the application scope of virtual characters, and providing higher flexibility and efficiency for digital content creation. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 This is a flow chart of the method for matching the facial expressions and movements of a digital cat puppet through human voice in the present invention;

[0026] Figure 2 This is a system module structure diagram for matching the facial expressions and movements of a digital cat puppet through human voice in the present invention. DETAILED DESCRIPTION

[0027] The following is a detailed description of the embodiments of the present invention with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.

[0028] In the description of the present invention, it should be understood that the terms "upper", "lower", "front", "back", "left", "right", "top", "bottom", "inside", "outside", "first", "second", etc., indicating directions or positions or sequential relationships, are based on the directions or positions or sequential relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific direction, be constructed and operated in a specific direction, and therefore cannot be understood as limiting the present invention.

[0029] A method for matching facial expressions and movements of a digital cat puppet through human voice is invented, comprising the following steps:

[0030] S1, voice input and preprocessing, collects user voice data and performs voice data preprocessing. For example, a microphone or other audio input device can be used to collect the user's voice, and the collected voice is preprocessed by noise reduction, normalization, etc. to improve the accuracy of subsequent processing;

[0031] S2, speech recognition, converts the pre-processed speech data into text data, performs sentiment analysis, and extracts sentiment information from the text data. After text conversion, existing sentiment analysis technology can be used to extract sentiment information, such as happiness, sadness, anger, etc.

[0032] S3, facial action mapping, builds an action library, pre-defines various facial actions of the digital cat (such as blinking, smiling, frowning, etc.), and assigns an emotion label to each action; based on the results of emotion analysis, selects appropriate facial expression actions for mapping;

[0033] S4, facial action generation and rendering, generates corresponding facial actions according to the mapping results, and renders the generated facial actions onto the digital cat puppet model in real time.

[0034] As a further illustration of the present invention, in step S1, a Python library such as pyaudio or sounddevice is used for voice acquisition; and librosa or pydub is used for audio preprocessing.

[0035] Through the above method, dynamic interaction between virtual characters and user voice input can be achieved. Based on voice preprocessing and emotion analysis technology, emotional information in the voice can be accurately extracted; based on the emotion-to-action mapping mechanism combined with 3D modeling and real-time rendering technology, the naturalness and consistency of the digital cat puppet's facial movements are ensured, ensuring real-time interaction performance, thereby achieving a natural, real-time and emotion-driven virtual character interaction experience, enhancing the user's immersion and sense of reality in virtual interaction scenes, expanding the application scope of virtual characters, and providing higher flexibility and efficiency for digital content creation.

[0036] Specifically, in step S2 of the above method, speech recognition APIs such as Google Speech-to-Text, Microsoft Azure Speech Service, or open source tools such as DeepSpeech can be used to convert speech to text; natural language processing (NLP) tools such as spaCy, NLTK, or pre-trained sentiment analysis models (such as BERT) can be used to perform sentiment analysis.

[0037] In step S3 of the above method, a facial action library of a digital cat puppet is created using a 3D modeling tool such as Blender or Maya; and a Python script or configuration file is used to map the emotion labels to the actions in the action library.

[0038] In step S4 of the above method, a game engine such as Unity or Unreal Engine is used to generate and render facial movements; and skeletal animation or blendshape technology is used to achieve detailed control of facial movements.

[0039] Preferably, after the facial expressions and movements of the digital cat puppet image are generated by rendering the human voice signal, user feedback can be collected and iterative optimization can be performed.

[0040] To achieve the above-mentioned method of matching facial expressions and movements of a digital cat doll image through human voice, the present invention provides a system for matching facial expressions and movements of a digital cat doll image through human voice, which specifically includes a voice input and preprocessing module, a voice recognition module, a facial action mapping module, an action generation module, and a real-time feedback module; the voice input and preprocessing module is used to collect the user's voice signal and preprocess and optimize the voice signal to provide high-quality audio data for subsequent processing; the voice recognition module is used to convert the collected voice signal into text data and perform sentiment analysis on the text data to extract sentiment information; the facial action mapping module is used to build a facial action library and select appropriate facial expression movements for mapping based on the results of sentiment analysis; the action generation module is used to generate corresponding facial movements based on the mapping results and render the generated facial movements in real time onto the digital cat doll model; the real-time feedback module is used to collect user feedback on the action matching effect and iteratively optimize the system. The system of this embodiment collects and optimizes the voice signal through the voice input and preprocessing module, providing high-quality audio data for subsequent processing. Utilizing speech recognition and sentiment analysis technologies, the system converts speech into text and extracts emotional information, serving as the core basis for driving the digital cat's facial movements. Leveraging a pre-built digital cat facial movement library and a precise emotion-to-action mapping mechanism, the system selects appropriate facial movements based on sentiment analysis results. Using efficient real-time rendering technology, these movements are then smoothly rendered onto the digital cat model, achieving a natural flow of speech emotion and virtual character expression.

[0041] In a preferred embodiment, the voice input is integrated into a complete system with a pre-processing module, a speech recognition module, a facial action mapping module, an action generation module, and a real-time feedback module. Message queues (such as RabbitMQ) or event-driven architectures (such as Node.js) are used to carry out inter-module communication to ensure that the data between each module is smoothly transmitted. Multithreading or asynchronous programming techniques are used to improve the response speed of the system, and the system is optimized for performance to ensure real-time performance and user experience. After system integration, each module of the system is functionally tested. Automated testing tools such as pytest can be used to perform functional testing to ensure that each module can work properly. As mentioned above, in actual applications, feedback can be collected after rendering and generating digital cat doll facial expression movements, and iterative optimization can be performed to optimize the user experience of the system.

[0042] In practical applications, the system of the present invention that matches the facial expressions and movements of a digital cat puppet through human voice can be deployed to a target platform. For example, Docker can be used for containerized deployment to deploy the system to a PC, mobile device, or cloud to ensure consistency in different environments. CI / CD tools such as Jenkins or GitHub Actions can be used for continuous integration and continuous deployment.

[0043] The above description is merely an explanation of the preferred embodiments of the present invention and should not be construed as limiting the claims. The present invention is not limited to the above embodiments, and variations in the specific structure are permitted. In short, all variations made within the scope of the independent claims of the present invention are also within the scope of protection of the present invention.

Claims

1. A method for matching facial expressions and movements of a digital cat puppet through human voice, characterized in that: The following steps are involved: S1, voice input and preprocessing, collects user voice data and performs voice data preprocessing; S2, speech recognition, converts the pre-processed speech data into text data, performs sentiment analysis, and extracts sentiment information from the text data; S3, facial action mapping, builds an action library, pre-defines various facial actions of the digital cat puppet, and assigns an emotion label to each action; based on the results of emotion analysis, selects appropriate facial expression actions for mapping; S4, facial action generation and rendering, generates corresponding facial actions according to the mapping results, and renders the generated facial actions onto the digital cat puppet model in real time.

2. The method for matching facial expressions of digital cat dolls by human voice according to claim 1, wherein: In step S1, the Python library is used for voice acquisition; librosa or pydub is used for audio preprocessing.

3. The method for matching facial expressions of digital cat dolls by human voice according to claim 1, wherein: In step S2, speech-to-text conversion is performed using a speech recognition API or open source tools; Perform sentiment analysis using natural language processing tools or pre-trained sentiment analysis models.

4. The method for matching facial expressions of digital cat dolls by human voice according to claim 1, wherein: In step S3, a facial action library of the digital cat puppet is created using a 3D modeling tool; and the emotion tags are mapped to the actions in the action library using a Python script or configuration file.

5. The method for matching facial expressions of digital cat dolls through human voice according to claim 1, characterized in that: In step S4, a game engine is used to generate and render facial movements; and skeletal animation or blendshape technology is used to achieve detailed control of facial movements.

6. The method for matching facial expressions of digital cat dolls through human voice according to claim 1, characterized in that: The step S4 includes collecting user feedback and performing iterative optimization steps.

7. A system for matching facial expressions and movements of digital cat puppets with human voice, characterized by: It includes speech input and preprocessing module, speech recognition module, facial action mapping module, action generation module and real-time feedback module; The voice input and preprocessing module is used to collect the user's voice signal and preprocess and optimize the voice signal to provide high-quality audio data for subsequent processing; The speech recognition module is used to convert the collected speech signals into text data, and perform sentiment analysis on the text data to extract sentiment information; The facial action mapping module is used to build a facial action library and select appropriate facial expression actions for mapping based on the results of emotion analysis; The action generation module is used to generate corresponding facial actions according to the mapping results, and render the generated facial actions on the model of the digital cat puppet in real time; The real-time feedback module is used to collect user feedback on the action matching effect and iteratively optimize the system.

8. The system for matching facial expressions and movements of a digital cat puppet through human voice according to claim 7, characterized in that: The speech input and preprocessing module, speech recognition module, facial action mapping module, action generation module and real-time feedback module are integrated into a complete system, using message queues or event-driven architecture for communication between modules, and using multi-threading or asynchronous programming technology to improve the response speed of the system.

9. The system for matching facial expressions and movements of a digital cat puppet through human voice according to claim 7, characterized in that: Use Docker to deploy the system to the target platform and use CI / CD tools for continuous integration and continuous deployment.