A voice-driven approach to facial expressions of digital curling players

By establishing a personalized expression parameter library for digital curling players based on curling game videos and using a neural network model to drive facial expressions, the problem of the inability to achieve personalized customization in existing technologies is solved, and realistic simulation and interaction of digital curling players are achieved.

CN119359873BActive Publication Date: 2025-09-19HARBIN INST OF TECH +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411564299.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-05
Publication Date
2025-09-19
Estimated Expiration
2044-11-05

AI Technical Summary

Technical Problem

Existing technologies are unable to achieve personalized customization of digital curling athletes' facial expressions. The modeling process is cumbersome and cannot meet the needs of different users or scenarios.

Method used

Based on curling game videos, a digital human image of a curling player is created. Through multi-view 3D reconstruction and neural network models, personalized expression parameters are learned and driven to achieve online voice-driven facial expressions.

Benefits of technology

It enhances the user's immersion and participation, realizes realistic simulation and interaction of digital curling players, and meets the personalized needs of different users or scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119359873B_ABST
    Figure CN119359873B_ABST
Patent Text Reader

Abstract

This invention proposes a voice-driven method for the facial expressions of a digital curling athlete. The method creates a digital human image of the curling athlete based on curling match videos and establishes a library of personalized emotional parameters for the curling athlete. Using a neural network model, the curling match audio is converted into personalized expression parameters, and then these are converted into three-dimensional facial animations, enabling personalized voice-driven digital human creation. The modeling process for the digital curling athlete takes into account individual characteristics, such as facial features and expression traits. Furthermore, the digital human is modeled solely using match videos, and voice-driven online creation is implemented. This helps developers quickly implement personalized customization in digital human development to meet the needs of different users or scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of metaverse, digital humans, and artificial intelligence technology, and in particular to a voice-driven method for the facial expressions of a digital curling player. Background Art

[0002] With the development of 3D modeling technology, personalized digital human technology has become a hot research topic in virtual sports. By creating 3D digital human avatars that closely resemble real-life athletes, and using voice and motion control, digital humans can simulate the behaviors and expressions of real athletes, creating supplementary perspectives and providing viewers with a multi-angle and immersive virtual viewing experience.

[0003] In the sport of curling, personalized digital human modeling based on curling athlete competition videos has become an innovative and efficient method. Using deep learning algorithms, a speech-to-expression mapping system has been developed to accurately recognize athlete speech and convert it into facial expression control parameters. This system, combined with facial features from curling competition videos, enables realistic simulation. This method not only improves the efficiency of digital human creation but also provides new opportunities for virtual training and competition simulations for curling athletes, delivering a richer and more realistic experience. However, existing methods lack personalized customization, and the digital human modeling technology is cumbersome, making it difficult to meet the needs of different users or scenarios. Summary of the Invention

[0004] The present invention aims to address the problems of the prior art by proposing a voice-driven method for the facial expressions of a digital curling athlete. The method creates a digital human image of the curling athlete based on a curling match video and establishes a library of personalized emotion parameters for the curling athlete. A neural network model is then used to convert the curling match audio into personalized expression parameters, which are then converted into three-dimensional facial animations, enabling personalized voice-driven digital human animation.

[0005] The present invention is implemented by the following technical solution. The present invention proposes a voice-driven method for the facial expression of a digital curling player, the method comprising the following steps:

[0006] Step 1: Create a digital human image of the curling athlete based on a curling competition video. A textured 3D facial model of the curling athlete is obtained through multi-view 3D reconstruction. A digital human template with complete skeleton and facial control parameters is created. This template is aligned with the 3D facial model, and the skeleton points and facial control parameters are calibrated to obtain a digital facial image of the curling athlete with the athlete's unique characteristics.

[0007] Step 2: Build a personalized emotional parameter library for curlers using curling competition videos. Use a neural network model to learn emotional vectors from the curlers' calls. This process is then used to learn and output emotional vectors within the neural network model. The personalized facial and emotional features of the curlers during their expressions and calls are then extracted, trained, and inferred separately.

[0008] Step 3: Use a neural network model to learn expression parameters from the curling game audio, combine the emotion coefficient obtained in step 2 with the data in the personalized emotion parameter library, and superimpose it with the expression parameters to convert the curling game audio into personalized expression parameters, realizing an online voice-driven method for facial expressions, and converting personalized expression parameters into three-dimensional facial animation to realize personalized digital human voice-driven.

[0009] Furthermore, the step 1 is specifically as follows:

[0010] Step 1.1: Collect a series of overlapping multi-view facial images of curling athletes from a curling competition video. Based on a multi-view stereoscopic 3D reconstruction method, use 3D modeling software to generate a 3D facial model of the curling athlete with texture mapping.

[0011] Step 1.2: Create digital human templates with complete skeleton and facial control parameters, including male and female templates with various face shapes. Select the face shape most similar to the curling athlete in the curling competition video from the digital human templates and align it with the obtained 3D facial model of the curling athlete. This means fitting its vertices to the reconstructed 3D facial model without changing the original topology or number of vertices.

[0012] Step 1.3: Calibrate the skeleton points and facial control parameters of the digital human template after mesh vertex alignment. Convert the texture map of the 3D facial model into a texture material, replace the facial skin of the digital human template, modify the hairstyle, and add glasses and accessories to obtain a digital curling athlete's facial image with his or her own characteristics.

[0013] Furthermore, in step 2, based on the curling game video, emotional pictures of "happy", "sad", "afraid", "horrified", "surprised", "angry" and "disgusted" in six situations where curling players shouted "wipe the ice quickly", "stop wiping the ice", "a few seconds", "off the line", "top line" and "whisper" were collected, and each emotion was divided into two categories: "soothing" and "intense".

[0014] Furthermore, in step 2, a personalized emotion parameter library is established as follows: 68 key point sets {P1, P2, P3...P68} are detected for each frame of the curling game video, the facial expressions of the curling players are reconstructed using the eos library, the key points are fitted to obtain the Blendshape parameter values ​​of the facial expressions, the Blendshape parameter values ​​of each emotion are averaged, which is the parameter of the curling player under this emotion, and all the emotion parameters are saved to obtain a personalized emotion parameter library.

[0015] Furthermore, in step 2, a neural network model is constructed to learn and output emotion vectors as follows: a personalized emotion parameter library is used to design a neural network model, where the input layer receives the audio of curling athletes shouting, giving commands, and other communications; in the hidden layer, a convolutional neural network is used to extract deep features, and an attention mechanism is introduced to focus on key areas; an emotion vector learning module is added to learn and output emotion vectors from the features, representing the weights of the emotions contained in the curling game audio at this time in terms of happiness, sadness, fear, terror, surprise, anger, and disgust, to obtain a 1*7 vector; by training the model to minimize the error between the predicted and true emotion vectors, personalized emotion features are extracted from the curling athletes' shouts and expressions in real time.

[0016] Furthermore, in step 3, a neural network model is used to extract audio features from the curling game audio clip and perform feature representation; an expression network is used to analyze expression features; and a fully connected layer expands the expression parameter vector into a BlendShape parameter weight.

[0017] Furthermore, in step 3, the emotion vector obtained in step 2 is multiplied by the BlendShape parameter stored in the personalized emotion parameter library to obtain the emotion base. The emotion base is weightedly added to the expression parameters parsed by the neural network from the curling game audio to obtain personalized expression parameters with emotion. The overall loss function value Loss is designed as:

[0018]

[0019] Among them, loss P Indicates the minimum mean square error between the predicted BlendShape parameters and the true value, loss M Indicates the coherence stability value between the previous and next frames; n is the number of channels of the BlendShape parameter, y (i) (x) is the network output, is the expected output, Δy (i) (x) is the difference in network output for the corresponding frame, is the difference in expected output for the corresponding frame.

[0020] Furthermore, in step 3, the BlendShape parameters are sent to the 3D modeling software, which receives the BlendShape parameters in real time through LiveLink, controls the facial expressions of the digital curling player, and renders the animation in real time. After receiving the data, the parameter data is converted into an animation of the digital curling player's facial expressions using Blueprint, thereby driving the digital curling player to make corresponding facial movements in coordination with the voice, thereby realizing online voice-driven facial expressions of the digital curling player.

[0021] The present invention also proposes an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of the voice-driven method for the facial expressions of a digital curling player are implemented.

[0022] The present invention also proposes a computer-readable storage medium for storing computer instructions, which, when executed by a processor, implement the steps of the voice-driven method for the facial expressions of a digital curling player.

[0023] Compared with the prior art, the present invention has the following beneficial effects:

[0024] This invention provides a voice-driven method for simulating the facial expressions of a digital curling athlete. This method enhances user immersion and engagement through realistic simulation and interaction with the digital curling athlete, providing a more authentic and natural interactive experience. The modeling process for the digital curling athlete takes into account individual characteristics such as body shape, movement habits, and facial expressions, helping developers achieve better customization in digital human development to meet the needs of different users or scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0026] Figure 1 This is a flow chart of a voice-driven method for the facial expressions of a digital curling athlete described in the present invention. DETAILED DESCRIPTION

[0027] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0028] Combine Figure 1 The present invention proposes a voice-driven method for facial expressions of digital curling players, the method comprising the following steps:

[0029] Step 1: Create a digital human image of the curling athlete based on a curling competition video. A textured 3D facial model of the curling athlete is obtained through multi-view 3D reconstruction. A digital human template with complete skeleton and facial control parameters is created. This template is aligned with the 3D facial model, and the skeleton points and facial control parameters are calibrated to obtain a digital facial image of the curling athlete with the athlete's unique characteristics.

[0030] The step 1 is specifically as follows:

[0031] Step 1.1: Collect a series of overlapping multi-view facial images of curling athletes from a curling competition video. Based on a multi-view stereoscopic 3D reconstruction method, use 3D modeling software to generate a 3D facial model of the curling athlete with texture mapping.

[0032] Step 1.2: Create digital human templates with complete skeleton and facial control parameters, including male and female templates with various face shapes. Select the face shape most similar to the curling athlete in the curling competition video from the digital human template and align it with the obtained 3D facial model of the curling athlete. This means fitting its vertices to the reconstructed 3D facial model without changing the original topology or number of vertices.

[0033] Step 1.3: Calibrate the skeleton points and facial control parameters of the digital human template after mesh vertex alignment. Convert the texture map of the 3D facial model into a texture material, replace the facial skin of the digital human template, modify the hairstyle, and add glasses and accessories to obtain a digital curling athlete's facial image with his or her own characteristics.

[0034] Step 2: Build a personalized emotional parameter library for curlers using curling competition videos. Use a neural network model to learn emotional vectors from the curlers' calls. This process is then used to learn and output emotional vectors within the neural network model. The personalized facial and emotional features of the curlers during their expressions and calls are then extracted, trained, and inferred separately.

[0035] In step 2, based on the curling game video, emotional images of "happy", "sad", "afraid", "horrified", "surprised", "angry" and "disgusted" in six situations where curling players shouted "quickly wipe the ice", "stop wiping the ice", "a few seconds", "off the line", "top line" and "whisper" were collected, and each emotion was divided into two categories: "soothing" and "intense".

[0036] In step 2, a personalized emotion parameter library is established as follows: 68 key point sets {P1, P2, P3...P68} are detected for each frame of the curling game video, the facial expressions of the curling players are reconstructed using the eos library, the key points are fitted to obtain the Blendshape parameter values ​​of the facial expressions, the Blendshape parameter values ​​of each emotion are averaged, which is the parameter of the curling player under this emotion, and all the emotion parameters are saved to obtain a personalized emotion parameter library.

[0037] In step 2, a neural network model is constructed to learn and output emotion vectors. Specifically, the following steps are performed: a personalized emotion parameter library is used to design a neural network model, where the input layer receives audio of curling athletes shouting, giving commands, and performing other communications; in the hidden layer, a convolutional neural network is used to extract deep features, and an attention mechanism is introduced to focus on key areas; an emotion vector learning module is added to learn and output emotion vectors from the features, representing the weights of the emotions contained in the curling game audio at this time in terms of happiness, sadness, fear, terror, surprise, anger, and disgust, to obtain a 1*7 vector; the model is trained to minimize the error between the predicted and true emotion vectors, and personalized emotion features are extracted from the curling athletes' shouts and expressions in real time.

[0038] Step 3: Use a neural network model to learn expression parameters from the curling game audio, combine the emotion vector obtained in step 2 with the data in the personalized emotion parameter library, and superimpose it with the expression parameters to convert the curling game audio into personalized expression parameters, realizing an online voice-driven method for facial expressions, and converting personalized expression parameters into three-dimensional facial animation to realize personalized digital human voice-driven.

[0039] In step 3, a neural network model is used to extract audio features from the curling game audio clip and perform feature representation; an expression network is used to analyze expression features; and a fully connected layer expands the expression parameter vector into BlendShape parameter weights.

[0040] In step 3, the emotion coefficient obtained in step 2 is multiplied by the BlendShape parameter stored in the personalized emotion parameter library to obtain the emotion base. The emotion base is weighted and added to the expression parameters parsed by the neural network from the curling game audio to obtain the personalized expression parameters with emotion. The overall loss function value of the design is:

[0041]

[0042] Among them, loss P Indicates the minimum mean square error between the predicted BlendShape parameters and the true value, loss M Indicates the coherence stability value between the previous and next frames; n is the number of channels of the BlendShape parameter, y (i) (x) is the network output, is the expected output, Δy (i) (x) is the difference in network output for the corresponding frame, is the difference in expected output for the corresponding frame.

[0043] In step 3, the BlendShape parameters are sent to the 3D modeling software, which receives the BlendShape parameters in real time via LiveLink, controls the facial expressions of the digital curling player, and renders the animation in real time. After receiving the data, the parameter data is converted into an animation of the digital curling player's facial expressions using Blueprint, thereby driving the digital curling player to make corresponding facial movements in coordination with the voice, thus realizing online voice-driven facial expressions of the digital curling player.

[0044] The present invention also proposes an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of the voice-driven method for the facial expressions of a digital curling player are implemented.

[0045] The present invention also proposes a computer-readable storage medium for storing computer instructions, which, when executed by a processor, implement the steps of the voice-driven method for the facial expressions of a digital curling player.

[0046] The memory in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. The non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus RAM (DRRAM). It should be noted that the memory of the methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0047] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments may be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated therein. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a high-density digital video disc (DVD)), or a semiconductor medium (eg, a solid state disc (SSD)).

[0048] During implementation, each step of the above method can be completed by an integrated logic circuit of the hardware in the processor or by instructions in the form of software. The steps of the method disclosed in conjunction with the embodiments of the present application can be directly embodied as being executed by a hardware processor, or can be executed by a combination of hardware and software modules in the processor. The software module can be located in a storage medium mature in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in conjunction with its hardware. To avoid repetition, it will not be described in detail here.

[0049] It should be noted that the processor in the embodiments of the present application can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method embodiment can be completed by an integrated logic circuit of the hardware in the processor or by instructions in the form of software. The above processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component. The various methods, steps, and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in the embodiments of the present application can be directly embodied as being executed by a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium mature in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, registers, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware.

[0050] The above is a detailed introduction to the voice-driven method for facial expressions of a digital curling player proposed in the present invention. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the ideas of the present invention, there may be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as limiting the present invention.

Claims

1. A voice-driven method for facial expressions of digital curling players, characterized in that: The method comprises the following steps: Step 1: Create a digital human image of the curling athlete based on a curling competition video. A textured 3D facial model of the curling athlete is obtained through multi-view 3D reconstruction. A digital human template with complete skeleton and facial control parameters is created. This template is aligned with the 3D facial model, and the skeleton points and facial control parameters are calibrated to obtain a digital facial image of the curling athlete with the athlete's unique characteristics. Step 2: Build a personalized emotional parameter library for curlers using curling competition videos. Use a neural network model to learn emotional vectors from the curlers' calls. This process is then used to learn and output emotional vectors within the neural network model. The personalized facial and emotional features of the curlers during their expressions and calls are then extracted, trained, and inferred separately. Step 3: Use a neural network model to learn expression parameters from the curling game audio. Combine the emotion vector obtained in Step 2 with the data in the personalized emotion parameter library and superimpose it with the expression parameters. Convert the curling game audio into personalized expression parameters to implement an online voice-driven method for facial expressions. Convert the personalized expression parameters into a 3D facial animation to achieve personalized digital human voice-driven. In step 2, based on curling game videos, we collected emotional images of "happy", "sad", "afraid", "fear", "surprise", "angry", and "disgusted" in six situations: "quickly wipe the ice", "stop wiping the ice", "a few seconds", "off the line", "top line", and "whisper". Each emotion was then categorized into "soothing" and "intense"; In step 2, a personalized emotion parameter library is established by detecting 68 key point sets {P1, P2, P3…P68} for each frame of the curling game video, reconstructing the curling player's facial expression using the eos library, fitting the key points to obtain the Blendshape parameter values ​​of the facial expression, taking the average of the Blendshape parameter values ​​of each emotion, which is the parameter of the curling player under this emotion, and saving all the emotion parameters to obtain a personalized emotion parameter library; In step 2, a neural network model is constructed to learn and output emotion vectors. Specifically, the following steps are performed: A personalized emotion parameter library is used to design a neural network model. The input layer receives audio of curling athletes shouting, giving commands, and performing other communications. In the hidden layer, a convolutional neural network is used to extract deep features, and an attention mechanism is introduced to focus on key areas. An emotion vector learning module is added to learn and output an emotion vector from the features. This vector represents the weight of the emotions contained in the curling audio at that time, such as happiness, sadness, fear, terror, surprise, anger, and disgust, resulting in a 1x7 vector. The model is trained to minimize the error between the predicted and true emotion vectors, extracting personalized emotion features from the curling athletes' shouts and expressions in real time. In step 3, a neural network model is used to extract audio features from the curling game audio clip and perform feature representation; an expression network is used to analyze expression features; and a fully connected layer expands the expression parameter vector into BlendShape parameter weights. In step 3, the emotion vector obtained in step 2 is multiplied by the BlendShape parameter stored in the personalized emotion parameter library to obtain the emotion base. The emotion base is weightedly added to the expression parameters parsed by the neural network from the curling game audio to obtain the personalized expression parameters with emotion. The overall loss function value of the design is: in, Indicates the minimum mean square error between the predicted BlendShape parameters and the true value, Indicates the coherence stability value between the previous and next frames; n is the number of channels of the BlendShape parameter, is the network output, is the expected output, is the difference in network output for the corresponding frame, is the difference in expected output for the corresponding frame.

2. The method according to claim 1, characterized in that The step 1 is specifically as follows: Step 1.1: Collect a series of overlapping multi-view facial images of curling athletes from a curling competition video. Based on a multi-view stereoscopic 3D reconstruction method, use 3D modeling software to generate a 3D facial model of the curling athlete with texture mapping. Step 1.2: Create digital human templates with complete skeleton and facial control parameters, including male and female templates with various face shapes. Select the face shape most similar to the curling athlete in the curling competition video from the digital human templates and align it with the obtained 3D facial model of the curling athlete. This means fitting its vertices to the reconstructed 3D facial model without changing the original topology or number of vertices. Step 1.3: Calibrate the skeleton points and facial control parameters of the digital human template after mesh vertex alignment. Convert the texture map of the 3D facial model into a texture material, replace the facial skin of the digital human template, modify the hairstyle, and add glasses and accessories to obtain a digital curling athlete's facial image with his or her own characteristics.

3. The method according to claim 2, characterized in that In step 3, the BlendShape parameters are sent to the 3D modeling software, which receives the BlendShape parameters in real time via LiveLink, controls the facial expressions of the digital curling player, and renders the animation in real time. After receiving the data, the parameter data is converted into an animation of the digital curling player's facial expressions using Blueprint, thereby driving the digital curling player to make corresponding facial movements in coordination with the voice, thus realizing online voice-driven facial expressions of the digital curling player.

4. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 3 are implemented.

5. A computer-readable storage medium for storing computer instructions, characterized in that: When the computer instructions are executed by a processor, the steps of the method according to any one of claims 1 to 3 are implemented.

Citation Information

Patent Citations

  • Facial expression capturing method and device, storage medium and terminal

    CN115294624A

  • Implementation method of curling virtual reality competition

    CN116650940A