Facial motion generation method and device, electronic equipment and readable storage medium

By generating and fusing facial motion data, the high cost and unnatural interaction problems of facial expression control in bionic robots have been solved, enabling flexible and natural facial motion generation and improving the human-computer interaction experience.

CN119516587BActive Publication Date: 2026-04-21UBTECH ROBOTICS CORP LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
UBTECH ROBOTICS CORP LTD
Filing Date
2024-09-30
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

In existing technologies, controlling facial expressions in bionic robots requires a large amount of manpower and time, and lacks adaptability, making it difficult to achieve a natural and smooth interactive experience.

Method used

By acquiring the emotion category and audio text of the target emotion, first facial motion data expressing the target emotion is generated, and second facial motion data for reading aloud is integrated. Facial motion is controlled using expression locators to achieve the generation of target facial motion data.

Benefits of technology

It achieves flexible control of target emotions and a natural, immersive human-computer interaction experience, reducing manual and time costs and improving the naturalness and fluency of facial movements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119516587B_ABST
    Figure CN119516587B_ABST
Patent Text Reader

Abstract

This application provides a facial motion generation method, apparatus, electronic device, and readable storage medium. The method includes: acquiring the emotion category of a target emotion and audio text; determining the emotion change parameters corresponding to the emotion category, and generating first facial motion data expressing the target emotion based on the emotion change parameters; determining second facial motion data for reading the audio text aloud; and fusing the first facial motion data and the second facial motion data based on a first expression locator and a second expression locator to obtain target facial motion data. Through this application, the first facial motion data expressing the target emotion can be incorporated into the second facial motion data used for reading aloud.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to computer technology, and more particularly to a method, apparatus, electronic device, and readable storage medium for generating facial movements. Background Technology

[0002] Robotics is an interdisciplinary field that integrates knowledge and skills from key disciplines such as mechanical engineering, electrical engineering, computer science, and artificial intelligence. Among them, bionic robotics is a key branch of robotics. By creating blended shape data of the target object to drive the robot's facial expression system, the robot's facial expressions become more natural and diverse.

[0003] In related technologies, fine-grained control of the target object's facial expressions is achieved by creating a blendshape model of the target object and using motion capture technology. However, this method requires a large investment of manpower and time to create facial expression keyframes, has high requirements for hardware configuration, and the target object's motion data must be pre-recorded. This means that in practical applications, the target object can only mechanically repeat fixed actions and expressions, lacking the ability to adapt to changing circumstances, thus making it difficult to achieve a natural and smooth interactive experience. Summary of the Invention

[0004] This application provides a facial motion generation method, apparatus, electronic device, and readable storage medium, which can incorporate first facial motion data expressing target emotions into second facial motion data used for reading aloud.

[0005] The technical solution of this application embodiment is implemented as follows:

[0006] This application provides a method for generating facial movements, the method comprising:

[0007] Obtain the emotion category and audio text of the target emotion;

[0008] Determine the emotion change parameters corresponding to the emotion category, and generate first facial motion data expressing the target emotion based on the emotion change parameters, wherein the emotion change parameters are used to control the emotion change of the first facial motion data;

[0009] Second facial motion data for reading the audio text is determined, wherein the first facial motion data includes a first expression locator that is the same as the second expression locator portion included in the second facial motion data;

[0010] Based on the first expression locator and the second expression locator, the first facial motion data and the second facial motion data are fused to obtain the target facial motion data.

[0011] This application provides a facial motion generation device, the device comprising:

[0012] The data acquisition module is used to obtain the emotion category of the target emotion and the audio text;

[0013] A facial motion data generation module is used to determine the emotion change parameters corresponding to the emotion category, and generate first facial motion data expressing the target emotion based on the emotion change parameters, wherein the emotion change parameters are used to control the emotion change of the first facial motion data;

[0014] The facial motion data generation module is further configured to determine second facial motion data for reading the audio text, wherein the first facial motion data includes a first expression locator that is the same as the second expression locator included in the second facial motion data;

[0015] The data fusion module is used to fuse the first facial motion data and the second facial motion data based on the first expression locator and the second expression locator to obtain the target facial motion data.

[0016] This application provides an electronic device, the electronic device comprising:

[0017] Memory is used to store executable instructions or computer programs.

[0018] The processor, when executing computer-executable instructions or computer programs stored in the memory, implements the facial motion generation method provided in the embodiments of this application.

[0019] This application provides a computer-readable storage medium storing a computer program or computer-executable instructions for implementing the facial motion generation method provided in this application when executed by a processor.

[0020] This application provides a computer program product, including a computer program or computer executable instructions. When the computer program or computer executable instructions are executed by a processor, they implement the facial motion generation method provided in this application.

[0021] The embodiments of this application have the following beneficial effects:

[0022] By generating first facial motion data that expresses the target emotion through the emotion change parameters corresponding to the emotion category of the target emotion, more flexible and simple control of the first facial motion data of the target emotion is achieved. Through audio text, second facial motion data for reading the audio text is determined. Based on the first expression locator and the second expression locator, the first facial motion data and the second facial motion data are fused to obtain the target facial motion data. The first facial motion data that expresses the target emotion is then incorporated into the second facial motion data used for reading, achieving a more natural and immersive human-computer interaction experience. Attached Figure Description

[0023] Figure 1 This is a schematic diagram of the facial motion generation system architecture provided in the embodiments of this application;

[0024] Figure 2 This is a schematic diagram of the structure of an electronic device for generating facial movements provided in an embodiment of this application;

[0025] Figure 3 This is a first flowchart illustrating the facial motion generation method provided in this application embodiment;

[0026] Figure 4 This is a second flowchart illustrating the facial motion generation method provided in this application embodiment;

[0027] Figure 5 This is a schematic diagram of the third process of the facial motion generation method provided in the embodiments of this application;

[0028] Figure 6 This is a schematic diagram of the fourth process of the facial motion generation method provided in the embodiments of this application;

[0029] Figure 7 This is a schematic diagram of the fifth process of the facial motion generation method provided in the embodiments of this application;

[0030] Figure 8 This is a schematic diagram of the sixth process of the facial motion generation method provided in the embodiments of this application;

[0031] Figure 9 This is a schematic diagram illustrating the first principle of the facial motion generation method provided in this application embodiment;

[0032] Figure 10 This is a schematic diagram illustrating the second principle of the facial motion generation method provided in the embodiments of this application.

[0033] It should be noted that the terms "first" and "second" mentioned above are only used to distinguish between different options and do not represent the degree of superiority or inferiority of the options or their priority in the implementation process. Detailed Implementation

[0034] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0035] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0036] In the following description, the terms "first, second, third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0037] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0038] Unless otherwise defined, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in the embodiments of this application is for the purpose of describing the embodiments of this application only and is not intended to limit this application.

[0039] In the implementation of this application, the collection and processing of relevant data should strictly comply with the requirements of relevant laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.

[0040] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.

[0041] 1) BlendShape Model: This is a technique used in computer graphics and animation that allows complex facial expressions to be created by blending the weight values ​​of multiple facial locators. It is often used to simulate the facial movements of living organisms, enabling virtual characters to display a wide range of emotions and expressions.

[0042] 2) BlendShape data: In computer graphics and animation, blendshape data refers to facial deformation data used to create blendshape models. This data typically includes a series of expression targets and corresponding weights. By adjusting the weights of the expression targets, countless different facial expressions can be created.

[0043] 3) Expression Targets: These are key points or markers used in facial animation to identify and control specific facial muscles or areas. These targets can be virtual points or physical markers. In the Blendshape model system provided by Augmented Reality Kit (ARKit), which is based on 52 expression targets, the following are some examples of expression targets: left eye blink (eyeBlinkLeft), left eye looking down (eyeLookDownLeft), left eye looking at the tip of the nose (eyeLookInLeft), left eye looking to the left (eyeLookOutLeft), and left eye looking up (eyeLookUpLeft).

[0044] 4) Frame Rate Value: This is a metric used to measure the update speed of video, animation, or image sequences. The unit is frames per second (fps), which is the number of frames displayed per second. Frame rate value is an important concept in video production and game development, and is crucial for ensuring the smooth playback of videos or animations.

[0045] This application provides a facial motion generation method, apparatus, electronic device, computer-readable storage medium, and computer program product, which can incorporate first facial motion data expressing target emotions into second facial motion data used for reading aloud.

[0046] The electronic device for facial motion generation provided in this application can be various types of terminals or servers. The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smart TV, smartwatch, etc., but is not limited thereto. The terminal and server can be directly or indirectly connected via wired or wireless communication, which is not limited herein.

[0047] See Figure 1 , Figure 1 This is a schematic diagram of the architecture of the facial motion generation system provided in this application embodiment. In the facial motion generation system 10 provided in this application embodiment, in order to support a facial motion generation application, the terminal 400 connects to the server 200 through the network 300. The network 300 can be a wide area network or a local area network, or a combination of the two.

[0048] Terminal 400 can be used to obtain a facial motion generation request that includes the emotion category of the target emotion and the audio text.

[0049] In some embodiments, a facial motion generation plugin may be embedded in the client running in the terminal 400 to implement the facial motion generation method locally on the client. For example, the terminal 400 calls the facial motion generation plugin to implement the facial motion generation method, generates first facial motion data expressing the target emotion through the emotion change parameters corresponding to the emotion category of the target emotion, determines second facial motion data for reading the audio text through the audio text, and fuses the first facial motion data and the second facial motion data based on the first expression locator and the second expression locator to obtain the target facial motion data.

[0050] In some embodiments, after the terminal 400 obtains a facial motion generation request that includes the emotion category of the target emotion and the audio text, it calls the facial motion generation interface of the server 200 (which can be provided as a cloud service). The server 200 implements a facial motion generation method through a facial motion generation plugin. It generates first facial motion data expressing the target emotion through the emotion change parameters corresponding to the emotion category of the target emotion, determines second facial motion data for reading the audio text through the audio text, and fuses the first facial motion data and the second facial motion data based on the first expression locator and the second expression locator to obtain the target facial motion data.

[0051] The facial motion generation method provided in this application embodiment can be implemented by a terminal or a server alone, or by a terminal and a server working together. For example, the terminal can undertake the facial motion generation method described below alone, or the terminal can send a facial motion generation request to the server, which includes obtaining the emotion category of the target emotion and the audio text. The server executes the facial motion generation method according to the received facial motion generation request, which includes obtaining the emotion category of the target emotion and the audio text. It generates first facial motion data expressing the target emotion through the emotion change parameters corresponding to the emotion category of the target emotion, determines second facial motion data for reading the audio text through the audio text, and fuses the first facial motion data and the second facial motion data based on the first expression locator and the second expression locator to obtain the target facial motion data.

[0052] In some embodiments, the terminal or server can implement the facial motion generation method provided in this application by running a computer program. For example, the computer program can be a native program or software module in an operating system; it can be a native application (APP), that is, a program that needs to be installed in the operating system to run, such as a live streaming application; it can also be a mini-program, that is, a program that only needs to be downloaded to a browser environment to run; or it can be a mini-program that can be embedded in any APP. In short, the above-mentioned computer program can be any form of application, module or plugin.

[0053] See Figure 2 , Figure 2 This is a schematic diagram of the structure of an electronic device for facial motion generation provided in an embodiment of this application. Figure 2 The electronic device 500 shown can be Figure 1 The terminal 400 or server 200 in the electronic device 500 includes at least one processor 510, a memory 550, at least one network interface 520, and a user interface 530. The various components in the electronic device 500 are coupled together via a bus system 540. It is understood that the bus system 540 is used to implement communication between these components. In addition to a data bus, the bus system 540 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in... Figure 2 The general labeled all buses as Bus System 540.

[0054] The processor 510 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.

[0055] User interface 530 includes one or more output devices 531 that enable the presentation of media content, including one or more speakers and / or one or more visual displays. User interface 530 also includes one or more input devices 532, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.

[0056] The memory 550 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state storage, hard disk drives, optical disk drives, etc. The memory 550 may optionally include one or more storage devices physically located away from the processor 510.

[0057] The memory 550 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 550 described in this application embodiment is intended to include any suitable type of memory.

[0058] In some embodiments, memory 550 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, as illustrated below.

[0059] Operating system 551 includes system programs for handling various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic business functions and handling hardware-based tasks;

[0060] The network communication module 552 is used to reach other computing devices via one or more (wired or wireless) network interfaces 520, exemplary network interfaces 520 including: Bluetooth, WiFi, and Universal Serial Bus (USB), etc.

[0061] Presentation module 553 is configured to enable the presentation of information (e.g., a user interface for operating peripheral devices and displaying content and information) via one or more output devices 531 (e.g., a display screen, a speaker, etc.) associated with user interface 530;

[0062] The input processing module 554 is used to detect and translate one or more user inputs or interactions from one or more input devices 532.

[0063] In some embodiments, the apparatus provided in this application can be implemented in software. Figure 2 A facial motion generation device 555 stored in memory 550 is shown. This device can be software in the form of programs and plug-ins, and includes the following software modules: a data acquisition module 5551, a facial motion data generation module 5552, a data fusion module 5553, and a facial motion data application module 5554. These modules are logically connected and can therefore be arbitrarily combined or further separated according to the functions they implement. The functions of each module will be described below.

[0064] In other embodiments, the apparatus provided in this application can be implemented in hardware. As an example, the apparatus provided in this application can be a processor in the form of a hardware decoding processor, which is programmed to execute the facial motion generation method provided in this application. For example, the processor in the form of a hardware decoding processor can be one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.

[0065] See Figure 3 , Figure 3 This is a first flowchart illustrating the facial motion generation method provided in this application embodiment, which will be combined with... Figure 3 The steps shown are explained below. The facial motion generation method provided in this application embodiment can be implemented by the server or terminal alone, or by the server and terminal working together. Therefore, the executing entity of each step will not be repeated below.

[0066] In step 101, the emotion category of the target emotion and the audio text are obtained.

[0067] As an example, the audio data corresponding to the facial movements to be generated for the target object is obtained. Speech recognition processing is performed on the audio data to obtain audio text. Emotion recognition processing is then performed on the audio text to obtain K emotion recognition results for the facial movements to be generated. The emotion category of the target emotion is determined from these K emotion recognition results. The emotion recognition results include the emotion category and the duration corresponding to the emotion category. Determining the emotion category of the target emotion from the K emotion recognition results can be achieved as follows: the k-th emotion recognition result is taken as the target emotion, and the emotion category of the k-th emotion recognition result is taken as the emotion category of the target emotion. Here, K is a positive integer, k is an incrementing integer, and 1 ≤ k ≤ K. Speech recognition processing and emotion recognition processing can be implemented through services provided by third-party service providers or through locally trained artificial intelligence models; there are no restrictions here. Emotion category is a classification of human emotional experiences, including happiness, sadness, anger, and grief. The target object can be a digital virtual object or a robot object; there are no restrictions here.

[0068] For example, the audio data corresponding to the facial movements to be generated is obtained, and speech recognition processing is performed on the audio data to obtain the audio text "The weather is so nice today, but we can't go out for a spring outing." Then, emotion recognition processing is performed on the audio text to obtain the emotion recognition results: the emotion category of "The weather is so nice today" is happiness, with a duration of 2 seconds, and the emotion category of "But we can't go out for a spring outing" is sadness, with a duration of 3 seconds. The emotion category of happiness corresponding to "The weather is so nice today" in the emotion recognition results is determined as the emotion category of the target emotion.

[0069] In step 102, the emotion change parameters corresponding to the emotion category are determined, and based on the emotion change parameters, the first facial motion data expressing the target emotion is generated.

[0070] As an example, the emotion change parameters corresponding to different emotion categories can be the same or different, which is not limited here. The first facial motion data includes at least one frame of blendshape data arranged in temporal order. The blendshape data includes at least one-dimensional expression locator and the weight value corresponding to the expression locator. The weight value is a real number that is greater than or equal to 0 and less than or equal to 1.

[0071] In some embodiments, the first facial motion data includes N expression segments, where N is an integer greater than 1, see [link to relevant documentation]. Figure 4 , Figure 4 This is a schematic diagram of the second process of the facial motion generation method provided in the embodiments of this application. Figure 3The step 102 shown, "generating first facial motion data expressing the target emotion based on emotion change parameters", can be achieved through the following steps 1021 to 1022, which are explained in detail below.

[0072] In step 1021, the nth initial expression segment of the first facial motion data is determined based on the emotion category.

[0073] Where n is an integer that increases sequentially, 1≤n≤N.

[0074] As an example, the same emotion category includes at least one emotion expression method, and different emotion expression methods of the same emotion category correspond to different standard expression templates. That is, one emotion category corresponds to at least one standard expression template. From the at least one standard expression template corresponding to the emotion category, a standard expression template for constructing the nth initial expression segment is determined. The initial expression segment includes at least one frame of standard expression template arranged in chronological order. The at least one frame of standard expression template included in the initial expression segment is the same standard expression template. The sum of the frame numbers of the N expression segments is less than or equal to the number of frames expressing the target emotion.

[0075] A standard expression template is a specific blendshape data used to express a specific emotion category. For example, in the blendshape model system based on 52 expression locators provided by the Augmented Reality Kit (ARKit), the standard expression template corresponding to the sad category can be represented as a 52-dimensional blendshape data [0, 0.7, 0.7, 0.8, ..., 0, 0]. The standard expression template corresponding to the sad category uses expression locators with indices 1, 2, and 3 from the 52 expression locators. The initial weight value of the expression locator with index 1 is 0.7, the initial weight value of the expression locator with index 2 is 0.7, and the initial weight value of the expression locator with index 3 is 0.8 (the weight value of the expression locators corresponding to the other indices is 0). The list of relevant dimensions of the standard expression template can be represented as [1, 2, 3].

[0076] In some embodiments, the sum of the trigger probabilities of multiple standard emoji templates corresponding to the same emotion category is 1. When an emotion category corresponds to multiple standard emoji templates, "determining the standard emoji template for constructing the nth initial emoji fragment" can be achieved through the following steps: determining the trigger probabilities of the multiple standard emoji templates corresponding to the emotion category respectively; selecting a target standard emoji template based on the trigger probabilities of the multiple standard emoji templates; and using the target standard emoji template as the standard emoji template for constructing the nth initial emoji fragment.

[0077] For example, the emotion category of sadness corresponds to two different standard emoji templates, namely standard emoji template 1 and standard emoji template 2. The trigger probability of standard emoji template 1 is 0.6, and the trigger probability of standard emoji template 2 is 0.4. Based on the trigger probabilities of standard emoji template 1 and standard emoji template 2, the selected target standard emoji template has a probability of 0.6 of being standard emoji template 1 and a probability of 0.4 of being standard emoji template 2.

[0078] In this embodiment of the application, the same emotion category includes at least one way of expressing emotion. Different ways of expressing emotion within the same emotion category correspond to different standard expression templates. Different standard expression templates provide different forms of expression for each basic emotion category. The different forms of expression of the standard expression templates are reflected in the fact that, for the emotion category of anger, standard expression template A expresses anger by frowning and squinting, while standard expression template B expresses anger by glaring and downward-pointing eyebrows. This makes the target object's emotional expression more delicate and realistic, and increases the richness of expression.

[0079] In some embodiments, the "number of frames for expressing the target emotion" can be determined by the following steps: obtaining the frame rate value of the target facial motion playback; and determining the number of frames for expressing the target emotion based on the frame rate value and the duration of the target emotion.

[0080] As an example, the frame rate value for the target facial motion playback refers to the number of frames per second that the target object's facial motion is rendered. The unit of the frame rate value is frames per second (fps). The number of frames that express the target emotion is determined based on the product of the frame rate value and the duration of the target emotion.

[0081] For example, if the frame rate is 60fps and the duration of the target emotion is 2 seconds, then the number of frames used to express the target emotion is the product of the frame rate and the duration, which is 120 frames.

[0082] In step 1022, based on the emotion change parameters, the nth initial expression fragment is deformed to obtain the nth expression fragment.

[0083] As an example, the initial expression fragment includes at least one frame of standard expression templates arranged in chronological order. Based on the emotion change parameter, each standard expression template included in the nth initial expression fragment is deformed to obtain the nth expression fragment, wherein the nth expression fragment includes at least one frame of standard expression templates arranged in chronological order.

[0084] In this embodiment, by using standardized expression templates, the rendering performance of the target object's facial expressions can be optimized, improving the utilization efficiency of computing resources. Furthermore, it eliminates the need to create new templates from scratch for each new emotion category; existing standard expression templates can be adjusted to quickly create new standard expression templates for different emotion categories, thus saving development time and human resources. In addition, by adjusting the emotion change parameters, more flexible and simple control over the expression changes of the nth expression segment can be achieved.

[0085] In some embodiments, the emotion change parameters include the degree of emotion change, the magnitude of emotion change, the number of emotion change frames, and the number of emotion holding frames. The nth expression segment includes an emotion change segment and an emotion holding segment. See [link to documentation]. Figure 5 , Figure 5 This is a schematic diagram of the third process of the facial motion generation method provided in the embodiments of this application. Figure 4 The step 1022 shown can be achieved by performing the following steps 201 to 203 on the standard expression template included in the nth initial expression fragment, as explained in detail below.

[0086] In step 201, the weight value of the third expression locator included in the I-frame standard expression template is determined based on the degree of emotion change, the amplitude of emotion change, and the number of emotion change frames.

[0087] Where I represents the number of frames showing emotional changes.

[0088] As an example, the emotional change segment includes segments of emotional change from weak to strong and segments of emotional change from strong to weak. The emotional change frame count includes the number of emotional change frames from weak to strong and the number of emotional change frames from strong to weak. When the emotional change segment is a segment of emotional change from weak to strong, the weight value of the third expression locator included in the I-frame standard expression template is determined based on the degree of emotional change, the amplitude of emotional change, and the number of emotional change frames from weak to strong. When the emotional change segment is a segment of emotional change from strong to weak, the weight value of the third expression locator included in the I-frame standard expression template is determined based on the degree of emotional change, the amplitude of emotional change, and the number of emotional change frames from strong to weak.

[0089] In some embodiments, the third expression locator includes a fifth expression locator, a sixth expression locator, and a seventh expression locator. The fifth and sixth expression locators have a facial symmetry relationship, and the seventh expression locator is a third expression locator that is different from both the fifth and sixth expression locators. Figure 5Step 201 shown can be achieved through the following steps: based on the magnitude of the emotion change and the number of emotion change frames, determine the weight values ​​of the fifth and seventh expression locators included in the I-frame standard expression template; based on the weight value of the fifth expression locator, determine the weight value of the sixth expression locator.

[0090] As an example, facial symmetry refers to the left-right symmetry of a biological facial structure. For instance, in the blendshape model system provided by ARKit, which is based on 52 expression locators, the expression locator "left eye blink" and the expression locator "right eye blink" have a left-right symmetry relationship in facial structure. After determining the weight value of the fifth expression locator "left eye blink" included in the I-frame standard expression template based on the amplitude of emotional change and the number of emotional change frames, it is not necessary to calculate the weight value of the sixth expression locator "right eye blink". The weight value of the fifth expression locator "left eye blink" can be directly used as the weight value of the sixth expression locator "right eye blink".

[0091] In this embodiment, by keeping the weight values ​​of expression locators with existing facial symmetry consistent, it is possible to ensure that the left and right sides of the target object's face are consistent, avoiding visual unnaturalness caused by asymmetrical expressions. At the same time, it reduces the number of expression locators that need to be controlled independently, which can save computing resources and improve the efficiency of expression control.

[0092] In some embodiments, see Figure 6 , Figure 6 This is a schematic diagram of the fourth process of the facial motion generation method provided in the embodiments of this application. Figure 5 Step 201 shown can be implemented through the following steps 2011 to 2013, which are explained in detail below.

[0093] In step 2011, based on the number of emotion change frames, the standard expression template of frame I is determined from the nth initial expression segment.

[0094] As an example, when the emotional change segment is a segment where the emotion changes from weak to strong, a first range of emotional change frames is obtained, and I frames of emotional change from weak to strong are randomly selected from the first range of emotional change frames. When the emotional change segment is a segment where the emotion changes from strong to weak, a second range of emotional change frames is obtained, and I frames of emotional change from strong to weak are randomly selected from the second range of emotional change frames. From the nth initial expression segment, I frames of standard expression templates are determined. The first range of emotional change frames and the second range of emotional change frames can be the same or different, and there is no restriction here.

[0095] For example, when the emotional change segment is a segment where the emotion changes from weak to strong, the first emotional change frame number range [4, 10] is obtained, and 6 emotional change frames from weak to strong are randomly selected from the first emotional change frame number range. From the nth initial expression segment, 6 standard expression templates are determined.

[0096] In step 2012, the range of emotional change is determined based on the degree of emotional change, the magnitude of emotional change, and the initial weight value of the third expression locator included in the I-frame standard expression template.

[0097] As an example, during an emotional change, the emotional change range refers to the range of change of the weight value of the third expression locator between the emotional minimum and the emotional maximum. The emotional minimum is the minimum value that the weight value of the third expression locator may reach, and the emotional maximum is the maximum value that the weight value of the third expression locator may reach. The emotional change amplitude is the value used to determine the emotional maximum.

[0098] In this embodiment of the application, by randomly determining the number of emotional change frames within the range of emotional change frames, the emotional changes can be made more natural and unpredictable, avoiding mechanical and formulaic animation performance.

[0099] In some embodiments, Figure 6 Step 2012 shown can be achieved through the following steps: determining the range of maximum emotional value based on the degree of emotional change, the magnitude of emotional change, and the initial weight value; determining the random number determined from the range of maximum emotional value as the maximum emotional value; determining the weight value of the expression locator included in the last frame of the standard expression template of the (n-1)th expression segment as the minimum emotional value; and constructing the range of emotional change based on the maximum and minimum emotional values.

[0100] As an example, a random number determined from the range of maximum emotional values ​​is identified as the maximum emotional value, and the weight value of the expression locator included in the last frame of the standard expression template of the (n-1)th expression segment is identified as the minimum emotional value. Based on the maximum and minimum emotional values, the range of emotional changes is constructed. For the first expression segment, the minimum emotional value is a random number that is 0 or close to 0.

[0101] For example, when the target emotion category is sadness, the standard expression template corresponding to sadness is [0, 0.7, 0.7, 0.8, ..., 0, 0]. Taking the third expression locator with index 1 as an example, the minimum value of the maximum emotion range of the third expression locator is the initial weight value of 0.7, and the emotion change range is a coefficient sequence of 52 coefficients [1.0, 0.4, ..., 0.5, ..., 0.2, 0.2, ..., 0.5, 0.5]. The initial weight value of 0.7 and the value of the expression with index 1 in the emotion change range are calculated. The product of the coefficient 0.4 is used to determine the minimum value of the range of maximum emotional value of the third expression locator. The range of maximum emotional value is from the minimum value of 0.28 to the maximum value of 0.7. The maximum emotional value is randomly determined to be 0.5 from the range of maximum emotional value [0.28, 0.7]. The weight value of 0.2 of the expression locator of the last frame of the standard expression template of the n-1th expression segment is determined as the minimum emotional value, and the range of emotional value [0.2, 0.5] is obtained.

[0102] In some embodiments, "determining the range of maximum emotional value based on the degree of emotional change, the magnitude of emotional change, and the initial weight value" can be achieved through the following steps: when the degree of emotional change is less than the threshold of emotional change, the product of the magnitude of emotional change, the initial weight value, and the first coefficient is determined as the first minimum value; based on the first product, the first ratio, and the first minimum value, the first maximum value is determined; and based on the first maximum value and the first minimum value, the range of maximum emotional value is formed. When the degree of emotional change is greater than or equal to the threshold of emotional change, the product of the magnitude of emotional change and the initial weight value is determined as the second minimum value; based on the product of the first difference and the second difference, the threshold of emotional change, and the second minimum value, the second maximum value is determined; and based on the second maximum value and the second minimum value, the range of maximum emotional value is formed.

[0103] Among them, the first product is the product of the emotional change amplitude, the initial weight value and the second coefficient, the first ratio is the ratio of the emotional change degree to the emotional change degree threshold, the first difference is the difference between the initial weight value and the second minimum value, and the second difference is the difference between the emotional change degree and the emotional change degree threshold.

[0104] As an example, the range of maximum emotional value variation is determined by comparing the relationship between the degree of emotional change and the threshold of emotional change, where the threshold of emotional change is set to distinguish the strength of emotional change.

[0105] When the degree of emotional change is less than the threshold for emotional change, the range of the maximum emotional value is composed of the first minimum value and the first maximum value. The first minimum value is determined by the product of the emotional change amplitude, the initial weight value, and the first coefficient using the formula (1.1) as follows:

[0106] min_val=scale*value*α (1.1)

[0107] Wherein, min_val is the first minimum value, scale is the magnitude of sentiment change, value is the initial weight value, and α is the first coefficient, which can be set to 0.8.

[0108] The first maximum value is determined using formula (1.2), based on the first product, the first ratio, and the first minimum value. Formula (1.2) is as follows:

[0109]

[0110] Where max_val is the first maximum value, min_val is the first minimum value, (1-α)*scale*value is the first product, scale is the magnitude of sentiment change, value is the initial weight value, and (1-α) is the second coefficient; The first ratio is emo_level, which represents the degree of emotional change, and emo_level_cut is the threshold for the degree of emotional change.

[0111] When the degree of emotional change is greater than or equal to the threshold of emotional change, the range of the maximum emotional value is composed of the second minimum value and the second maximum value. The second minimum value is determined by multiplying the emotional change amplitude and the initial weight value using the formula (1.3) as follows:

[0112] min_val = scale * value (1.3)

[0113] Where min_val is the second minimum value, scale is the magnitude of sentiment change, and value is the initial weight value.

[0114] The second maximum value is determined using formula (1.4), based on the product of the first and second differences, the threshold for the degree of emotional change, and the second minimum value. Formula (1.4) is as follows:

[0115]

[0116] Where max_val is the second maximum value, min_val is the second minimum value, value-min_val is the first difference, value is the initial weight value, emo_level-emo_level_cut is the second difference, emo_level is the degree of emotional change, and emo_level_cut is the threshold for the degree of emotional change.

[0117] In step 2013, the weight values ​​of the third expression locators included in the I-frame standard expression template are determined based on the range of emotional changes.

[0118] As an example, based on the range of emotional changes, I weight values ​​are determined frame by frame for each third expression locator included in the I-frame standard expression template.

[0119] In some embodiments, Figure 6 Step 2013 shown can be achieved by performing the following steps for each dimension of the third expression locator: determining I random values ​​from the range of emotional changes; sorting the I random values ​​according to the order of emotional changes; and determining the sorted I random values ​​as the weight values ​​of the third expression locator.

[0120] As an example, the order of emotion change includes an order of emotion change from weak to strong and an order of emotion change from strong to weak. I random values ​​are determined from the range of emotion change. When the order of emotion change is from weak to strong, the I random values ​​are sorted from smallest to largest. When the order of emotion change is from strong to weak, the I random values ​​are sorted from largest to smallest. The sorted I random values ​​are then determined as the weight values ​​of the third expression locator.

[0121] For example, six random values ​​of 0.3, 0.33, 0.40, 0.25, 0.29, and 0.48 are determined from the range of emotion changes [0.2, 0.5]. Assuming the emotion changes in order from weakest to strongest, these six random values ​​are sorted from smallest to largest, resulting in 0.25, 0.29, 0.3, 0.33, 0.40, and 0.48. These six sorted random values ​​are then used as the weight values ​​for the third expression locator.

[0122] In this embodiment of the application, by randomly determining the maximum emotional value from the range of emotional maximum value changes and randomly determining the weight value of the third expression locator from the range of emotional changes, the randomness of the intensity of emotional changes is increased.

[0123] In step 202, the weight value of the fourth expression locator included in the J-frame standard expression template is determined based on the number of emotion-preserving frames.

[0124] Where J represents the number of frames that preserve emotion.

[0125] As an example, the emotion-preserving segment includes an emotion peak-preserving segment and an emotion silence-preserving segment, and the emotion-preserving frame count includes the emotion peak-preserving frame count and the emotion silence-preserving frame count. When the emotion-preserving segment is an emotion peak-preserving segment, the weight value of the fourth expression locator included in the J-frame standard expression template is determined based on the emotion peak-preserving frame count. When the emotion-preserving segment is an emotion silence-preserving segment, the weight value of the fourth expression locator included in the J-frame standard expression template is determined based on the emotion silence-preserving frame count. Here, emotion preservation refers to the state in which the weight value of the fourth expression locator remains unchanged, emotion peak preservation refers to the state in which the weight value of the fourth expression locator remains unchanged at the emotion maximum, and emotion silence preservation refers to the state in which the weight value of the fourth expression locator remains unchanged at the emotion minimum.

[0126] In some embodiments, see Figure 7 , Figure 7 This is a schematic diagram of the fifth process of the facial motion generation method provided in the embodiments of this application. Figure 5 Step 202 shown can be achieved through the following steps 2021 to 2023, which are explained in detail below.

[0127] In step 2021, based on the number of emotion-preserving frames, the standard expression template for frame J is determined from the nth initial expression fragment.

[0128] As an example, when the emotion-maintaining segment is an emotion peak-maintaining segment, the range of emotion peak-maintaining frames is obtained, and J frames of emotion peak-maintaining frames are randomly selected from the range of emotion peak-maintaining frames; when the emotion-maintaining segment is an emotion silence-maintaining segment, the range of emotion silence-maintaining frames is obtained, and J frames of emotion silence-maintaining frames are randomly selected from the range of emotion silence-maintaining frames, and J frames of standard expression templates are determined from the nth initial expression segment.

[0129] For example, when the emotion-maintaining segment is an emotion peak-maintaining segment, the range of emotion peak-maintaining frames [8, 20] is obtained, 15 emotion peak-maintaining frames are randomly selected from the range, and 15 standard expression templates are determined from the nth initial expression segment; when the emotion-maintaining segment is an emotion silence-maintaining segment, the range of emotion silence-maintaining frames [30, 120] is obtained, 60 emotion silence-maintaining frames are randomly selected from the range, and 60 standard expression templates are determined from the nth initial expression segment.

[0130] In step 2022, a target standard expression template is determined that is located before and adjacent to the standard expression template of frame J.

[0131] As an example, when the emotion-maintaining segment is an emotion peak-maintaining segment, the target standard expression template is the standard expression template of the last frame of the segment where the emotion changes from weak to strong in the nth expression segment. When the emotion-maintaining segment is an emotion silence-maintaining segment, the target standard expression template is the standard expression template of the last frame of the segment where the emotion changes from strong to weak in the nth expression segment.

[0132] In step 2023, the weight value of the expression locator included in the target standard expression template is determined as the weight value of the fourth expression locator included in the J-frame standard expression template.

[0133] As an example, the weight value of each expression locator included in the target standard expression template is determined frame by frame as the weight value of the corresponding fourth expression locator in the J-frame standard expression template.

[0134] In this embodiment, by randomly selecting frames within a specific range of emotion retention frames, facial expression changes become more natural and unpredictable, avoiding fixed patterns and thus enhancing the randomness of emotional change segments. While randomly selecting emotion retention frames, the naturalness of emotional transitions is maintained by selecting adjacent target standard expression templates, ensuring the continuity of emotional changes.

[0135] In step 203, an emotion change segment is constructed based on the weight value of the third expression locator, and an emotion retention segment is constructed based on the weight value of the fourth expression locator.

[0136] See Figure 9 , Figure 9 This is a schematic diagram of the first principle of the facial motion generation method provided in the embodiments of this application. Based on the weight value of the third expression locator (expression locator with an index value equal to 0), the change segment 901 of emotion from weak to strong and the change segment 902 of emotion from strong to weak are respectively constructed. Based on the weight value of the fourth expression locator, the emotion peak maintenance segment 903 and the emotion silence maintenance segment 904 are respectively constructed.

[0137] In step 103, second facial motion data for reading the audio text is determined.

[0138] The first facial motion data includes the same first expression locator as the second facial motion data includes the same second expression locator.

[0139] As an example, all expression locators used for facial motion generation are divided into expression locators that are only related to lip-sync, expression locators that are only related to emotion, expression locators that are related to both lip-sync and emotion, and expression locators that are not related to either lip-sync or emotion. The first facial motion data includes expression locators that are only related to emotion and expression locators that are related to both lip-sync and emotion. The second facial motion data is used to read audio text and includes expression locators that are only related to lip-sync and expression locators that are related to both lip-sync and emotion. The first expression locators included in the first facial motion data are partially the same as the second expression locators included in the second facial motion data, both including expression locators that are related to both lip-sync and emotion.

[0140] In step 104, based on the first expression locator and the second expression locator, the first facial motion data and the second facial motion data are fused to obtain the target facial motion data.

[0141] In some embodiments, see Figure 8 , Figure 8 This is a schematic diagram of the sixth process of the facial motion generation method provided in the embodiments of this application. Figure 3 Step 104 shown can be implemented through the following steps 1041 to 1042, which are explained in detail below.

[0142] In step 1041, the sum of the first weight value and the second weight value is determined as the target weight value of the first target expression locator.

[0143] Wherein, the first target expression locator is the same expression locator as the second expression locator in the first expression locator, the first weight value is the weight value of the first target expression locator included in the first facial motion data, and the second weight value is the weight value of the first target expression locator included in the second facial motion data.

[0144] As an example, when the sum of the first weight value and the second weight value is less than or equal to 1, the sum is determined as the target weight value of the first target expression locator; when the sum of the first weight value and the second weight value is greater than 1, the target weight value of the first target expression locator is determined as 1.

[0145] For example, the first facial motion data includes expression locator 1, expression locator 2, and expression locator 3, and the second facial motion data includes expression locator 3, expression locator 4, and expression locator 5. The first target expression locator is expression locator 3. The first weight value is the weight value sequence of expression locator 3 included in the first facial motion data [1.0, 0.4, 0.5, 0.2, 0.2, 0.5, 0.5]. The second weight value is the weight value of expression locator 3 included in the second facial motion data [0.2, 0.5, 0.5, 1.0, 0.4, 0.5, 0.2]. The first weight value and the second weight value are summed frame by frame to obtain the target weight value sequence of the first target expression locator [1.0, 0.9, 1.0, 1.0, 0.6, 1.0, 0.7].

[0146] In step 1042, the target weight value of the first target expression locator and the weight value of the second target expression locator are combined to determine the target facial motion data.

[0147] Among them, the second target expression locator is the expression locator other than the first target expression locator among the first expression locator and the second expression locator.

[0148] For example, the first facial motion data includes expression locator 1, expression locator 2, and expression locator 3; the second facial motion data includes expression locator 3, expression locator 4, and expression locator 5; the first target expression locator is expression locator 3; and the second target expression locator includes expression locators 1 and 2 from the first facial motion data, and expression locators 4 and 5 from the second facial motion data. The target weight value of the first target expression locator and the weight value of the second target expression locator are combined to determine the target facial motion data. The target facial motion data includes expression locators 1 and 2 from the first facial motion data, expression locators 4 and 5 from the second facial motion data, and the target weight value of the first target expression locator.

[0149] In this embodiment, by combining the weight values ​​of two different facial motion data, more flexible emotional expression is achieved, enabling the target object to exhibit rich emotional changes while speaking, thus improving the flexibility of emotional changes.

[0150] In some embodiments, Figure 3 After step 104 shown, the following steps may also be performed: smoothing the target facial motion data to obtain smoothed target facial motion data; applying the smoothed target facial motion data to the face of the target object.

[0151] As an example, the frame rate value of the facial motion playback is obtained, and the smoothing frame number m is determined based on the frame rate value, where the frame rate value is proportional to the smoothing frame number; the target frame in the target facial motion data is determined, and based on the smoothing frame number m, the standard expression templates of the target frame, the m frames before the target frame, and the m frames after the target frame are smoothed to obtain the smoothed target facial motion data; the smoothed target facial motion data is applied to the face of the target object, where the target frame is any frame of data in the target facial motion data.

[0152] In this embodiment, by smoothing the target facial motion data and applying weighted smoothing of the preceding and following frames in each dimension, it can be ensured that each frame of motion can utilize the information of the preceding and following frames, thereby achieving a smoother and more natural transition. This not only makes the virtual human's movements more natural and fluid, but also makes the movements of the driving hardware more natural and fluid, improving the driving quality of the target object's facial movements and making the facial movements more natural and fluid, thereby enhancing the user experience. In addition, it also avoids damage to the hardware motor due to sudden large movements.

[0153] The following describes an exemplary application of the embodiments of this application in a blendshape model system based on 52 facial locators provided by an actual ARKit.

[0154] With the continuous advancement of artificial intelligence technology, the application scope of digital humans is expanding rapidly. Whether in animation production, film special effects, game characters, advertising, customer service, or live streaming, digital humans are playing an increasingly important role. However, in practical applications, simply building models is far from sufficient. Digital humans also need the ability to dynamically capture movements and expressions to communicate and interact with humans more naturally and effectively. Furthermore, robotics technology is booming, with bionic robots emerging as a key branch. Enabling robots to perform lively and expressive actions and display rich emotions has become one of the important trends in bionic robot research and development.

[0155] However, in related technologies, the motion data of virtual humans must be pre-recorded, which means that in practical applications, virtual humans can only mechanically repeat fixed actions and expressions, lacking the ability to adapt to changing circumstances, making it difficult to achieve a natural and smooth interactive experience. In addition, a lot of manpower and time costs are required to produce expression keyframes, and high requirements are also placed on hardware configuration.

[0156] To address the aforementioned issues, this application proposes a facial motion generation method that can drive a virtual human face to display corresponding emotional expressions. Through integration and connection with hardware, it enables precise control of the robot's facial emotions.

[0157] I. Incomplete decoupling of content and emotion.

[0158] By performing incomplete decoupling of content and sentiment for the weights of the 52 blendshapes (i.e., there exists a shared dimension between content and sentiment, equivalent to the sentiment identifier mentioned above), the 52-dimensional blendshapes are split into four categories, resulting in:

[0159] 1) Dimensions that are only related to the content of speech, i.e., dimensions that are driven by lip movements (equivalent to the facial expression locators that are only driven by lip movements mentioned above).

[0160] 2) Dimensions that are only related to emotions, namely, dimensions related to eyebrows, eyes, and cheeks (equivalent to the expression locators that are only related to emotion-driven factors mentioned above).

[0161] 3) Dimensions that are related to both the content and emotion of speech, namely the dimension related to the corner of the mouth (equivalent to the facial expression locator mentioned above that is related to both mouth shape and emotion).

[0162] 4) Dimensions that are not closely related to the content and emotion of speech, namely the blendshape dimension corresponding to actions that the robot hardware cannot perform (equivalent to the facial expression locator mentioned above that is not related to lip-sync and emotion-sync). Even if there is data for these dimensions, the robot's facial movements and expressions will not be much different.

[0163] Because different emotional expressions result in varying facial changes in terms of area and intensity, different combinations of dimensions may be needed to express different emotions.

[0164] 2. Obtain the set of dimensions used by the standard template emoticons (equivalent to the standard emoticon templates mentioned above).

[0165] Create multiple standard template emojis according to the required emotion category.

[0166] Since the same emotion may involve multiple different forms of expression and different degrees of emotional expression, a single emotion can be used to create multiple standard template expressions for different facial expressions. The different degrees of emotion will be introduced later.

[0167] Different expressions of the same emotion may use different blendshape dimensions. Therefore, the list of related dimensions of an emotion is a set of dimensions used by the standard template emoticons corresponding to various different expressions of an emotion category. The set of dimensions of the standard template emoticons is a set of dimensions of the list of related dimensions of multiple emotion categories.

[0168] For example, for the three emotion categories of happiness, sadness, and anger, happiness uses a standard template emoji with one expression form, and the list of related dimensions for happiness includes: [0, 1, 4, 7, 8, 15, 16, 20, 21] dimensions; sadness has two different standard template emojis with two different expressions, and the list of related dimensions for sadness includes: [0, 1, 4, 35, 36, 40, 41] and [1, 2, 3, 40, 41] dimensions; anger uses a standard template emoji with one expression form, and the list of related dimensions for anger includes: [8, 9, 10, 15, 16, 20, 21] dimensions.

[0169] The set of dimensions for the standard template emoji is the set of dimensions mentioned above, namely [0, 1, 4, 7, 8, 15, 16, 20, 21, 35, 36, 40, 41, 2, 3, 9, 10]. The index values ​​corresponding to the dimensions mentioned above do not represent the actual index values ​​of the 52-dimensional blendshape in ARKit, but are only used to illustrate the relationship between the dimensions used to create the standard template emoji and the dimensions used for a single emoji.

[0170] III. Constructing standard template emoji data.

[0171] Four standard template emojis were created, categorized into three emotion types.

[0172] Based on the blendshape dimension set used in the standard template emoji, the values ​​corresponding to each dimension when the emotional intensity reaches its maximum value under different emotional expression forms of each category are obtained (equivalent to the weight values ​​mentioned above). The blendshape data of the standard template emoji is 52-dimensional, and the value range corresponding to the used dimensions is [0, 1.0]. The value corresponding to the unused dimensions is set to 0, resulting in four 52-dimensional standard template emoji data.

[0173] IV. Control of changes in emotion-related dimensions.

[0174] The control of changes in the weight values ​​corresponding to the emotion-related dimensions can be achieved through the magnitude and range of emotion changes, the number of frames from weak to strong emotion, the change curve from weak to strong emotion, the number of frames maintained in the strong emotion state, the number of frames from strong to weak emotion, the change curve from strong to weak emotion, the frequency of emotion triggering, and the trigger probability of different expressions of the same emotion.

[0175] See Figure 9 The emotion-related dimensions include 42 dimensions such as [0, 1, ..., 41]. Figure 9In the coordinate system, the horizontal direction represents the number of frames in the blendshape data, and the vertical direction represents the value of the corresponding dimension. The expression of emotion includes emotion enhancement, emotion maintenance, emotion weakening, and no emotion state. The duration of emotion expression varies, and the change curve of the blendshape data also varies.

[0176] For example, for the emotion category of sadness, there are two different forms of expression, using the dimensions [0, 1, 4, 35, 36, 40, 41] and [1, 2, 3, 40, 41] respectively;

[0177] The standard template for the emotion category of sadness is data1 = [0.6, 0.7, ..., 0.9, ..., 0.8, 0.8, ..., 0.5, 0.5, ...], which has a total of 52 dimensions. The values ​​of the omitted dimensions are all 0.

[0178] The standard template for the emotion category of sadness is data2 = [0, 0.7, 0.7, 0.8, ..., 0.8, 0.8, ...], which has a total of 52 dimensions. The values ​​of the omitted dimensions are all 0.

[0179] See Figure 10 , Figure 10 This is a schematic diagram of the second principle of the facial motion generation method provided in the embodiments of this application. The following will take the standard template expression data1 of the sad emotion category as an example to illustrate the numerical changes of the emotion-related dimensions in the 52-dimensional data included in data1 (hereinafter referred to as the emotion-related dimensions).

[0180] Step 1, set the amplitude and range of emotion variation: Emotion varies between its minimum and maximum values. That is, each emotion-related dimension has a minimum and maximum value. The range of emotion variation is from the minimum to the maximum value. The range of the maximum value is [data1*s1, data1*1.0], where s1 is a coefficient (equivalent to the emotion variation amplitude mentioned above), s1 = [1.0, 0.4, ..., 0.5, ..., 0.2, 0.2, ..., 0.5, 0.5, ...], a total of 52 dimensions. The minimum value is the value at the end of the last emotion trigger (the minimum value is 0 at the first emotion trigger).

[0181] Step 2, set the number of frames from weak to strong emotion: set the range of the number of frames nframe1 from weak to strong emotion to [4, 10]. When nframe1 = 6, it takes 6 frames of data to transition from the minimum value of emotion to the maximum value of this emotion. Therefore, the transition frame number of emotion-related dimensions is 6 frames. Through these 6 frames of transition, the emotion-related dimensions reach the maximum value of this emotion expression. The maximum value of each emotion expression is not necessarily the same, thereby enhancing the diversity of emotion changes.

[0182] Step 3, set the data for the change curve from weak to strong: When it takes 6 frames for the emotion to change from weak to strong, the value corresponding to each dimension of the emotion is a random number between the minimum and maximum values. Arrange the 6 random numbers of each dimension of the emotion-related dimensions from small to large to form the change curve from weak to strong for that dimension.

[0183] Step 4, set the number of frames to maintain the emotional intensity state: The range of the number of frames to maintain the emotional intensity state is [8, 20]. When nframe_emo = 15, the values ​​corresponding to the emotional-related dimensions are maintained for 15 frames.

[0184] Step 5, set the number of frames from strong to weak emotion: The range of the number of frames nframe2 from strong to weak emotion is [8, 16]. See the steps above for setting the number of frames from weak to strong emotion, which will not be repeated here.

[0185] Step 6, set the change curve from strong to weak: Refer to the steps above for setting the change curve data from weak to strong, which will not be repeated here.

[0186] Step 7, set the frequency of emotion triggering: The frequency of emotion triggering refers to the frequency at which the emotion weakens and remains for a certain number of frames before returning to the state of emotion. The range of nframe_normal is set to [30, 120]. When nframe_normal = 60, the values ​​corresponding to the emotion-related dimensions remain for 60 frames after the last emotion ends. That is, the emotion-related dimensions in the 52-dimensional data remain for 60 frames each. The data that is maintained for the emotion is the data that was used when the last emotion ended.

[0187] It also allows setting the trigger probability for different expressions of the same emotion: for example, sadness has two different expressive forms, with probabilities set to p1 = 0.4 and p2 = 0.6 respectively. If the same emotion has multiple different expressions, the sum of the probabilities is 1. This ensures that different expressions of the same emotion occur with a certain probability.

[0188] Through steps 1 to 7 above, emotion-related dimension data is constructed, namely emotion data (equivalent to the first facial action data mentioned above).

[0189] V. Controlling the numerical changes of emotion-related dimensions, and adjusting them symmetrically from left to right.

[0190] Since a normal human face is basically symmetrical, the control of facial expressions must also take into account left and right symmetry. Based on this, from the emotion-related dimensions, we can identify those dimensions a (equivalent to the fifth expression locator above) and b (equivalent to the sixth expression locator above) that include both sides. We can control only the value corresponding to dimension a (the value on one side of the face), and the value corresponding to dimension b can directly reuse the value corresponding to dimension a (the data on the other side uses the same value).

[0191] VI. Add emotional expression to basic lip-sync data.

[0192] From the 52-dimensional blendshape data that drives the virtual human's facial expressions and movements, dimensions that are only related to the spoken content are extracted. During algorithmic processing, when the same text is spoken under different emotions, corresponding lip-shape data (equivalent to the second facial movement data mentioned above) is constructed. The lip-shape data that is only related to the spoken content remains unchanged, thereby maintaining its basic lip shape to ensure consistency between the spoken content and the lip shape. The lip-shape data is also 52-dimensional, with the numerical range of the dimension related to the spoken content being [0, 1.0], and the values ​​of other dimensions being 0.

[0193] The data on lip movements and the data on emotion-related dimensions are integrated. The emotion-related dimensions include dimensions that are only related to emotion, and dimensions that are related to both the content of the speech and the emotion.

[0194] For dimensions that are related to both speech content and emotion, the values ​​corresponding to the above dimensions are the sum of the emotion-related dimension data and the dimension data that are only related to speech content. That is, the values ​​corresponding to the above dimensions of lip shape data and the above dimensions of emotion data are added together. When the sum exceeds 1.0, it is set to 1.0.

[0195] For example, when the dimensions used for lip-reading data are [0, 1, 48, 49], and the emotion-related dimensions include [0, 1, 4, 7, 8, 15, 16, 20, 21, 35, 36, 40, 41, 2, 3, 9, 10], it can be seen that the 48th and 49th dimensions use the basic lip-reading data unchanged, the 0th and 1st dimensions use the superposition of the two, and the remaining dimensions [4, 7, 8, 15, 16, 20, 21, 35, 36, 40, 41, 2, 3, 9, 10] use the values ​​corresponding to the emotion-related dimensions.

[0196] VII. Expression and control of emotions when there is no content to speak.

[0197] Similar to the handling of lip-sync dimensions during speech, this is a special case where the values ​​of all lip-sync dimensions are 0, which will not be elaborated upon here.

[0198] 8. Data smoothing processing.

[0199] Smooth the values ​​corresponding to each emotion-related dimension and lip-shape-related dimension in the standard template emoji dimension set.

[0200] For example, if the emotion-related dimensions include [0, 1, 4, 35, 36, 40, 41], then smoothing is performed on each of these 7 dimensions. Considering that some of the dimensions are symmetrical, one side of the symmetrical dimension pair can be smoothed, and the corresponding dimension on the other side can be given the same value.

[0201] The smoothing method is as follows: obtain the frame rate value of the blendshape data, and determine the number of frames to be smoothed based on the frame rate value. That is, when the frame rate value is high, more frames need to be smoothed at the same time, and when the frame rate value is low, fewer frames are smoothed. When the frame rate value is 30pfs, a total of 7 frames of data are used for smoothing, and when the frame rate value is 10pfs, 3 frames of data are used for smoothing.

[0202] Taking a frame rate of 30 as an example, smoothing is performed using the data from the first and last 7 frames. Starting from the first frame, the data is smoothed. For the nth frame, the data from the (n-3)th frame, (n-2)th frame, (n-1)th frame, nth frame, (n+1)th frame, (n+2)th frame, and (n+3)th frame are smoothed with weights of 0.04, 0.07, 0.22, 0.34, 0.22, 0.07, and 0.04 respectively.

[0203] For the first few frames of the blendshape data:

[0204] When n=1, the data in the first frame, the second frame, the third frame, and the fourth frame are smoothed with weights of 0.45, 0.3, 0.15, and 0.1, respectively.

[0205] When n=2, the data in the first frame, second frame, third frame, fourth frame, and fifth frame are smoothed with weights of 0.2, 0.4, 0.2, 0.1, and 0.1, respectively.

[0206] When n=3, the data in the first frame, second frame, third frame, fourth frame, fifth frame, and sixth frame are smoothed with weights of 0.1, 0.2, 0.3, 0.2, 0.1, and 0.1, respectively.

[0207] In the last few frames of the blendshape data, use similar weight setting steps as described above to swap the order of the weights to ensure that the weights of the smoothed data used sum to 1, with the middle data having the highest weight and the farther away the data has, the lower the weight.

[0208] The facial motion generation method of this application, through a simple and easily configurable implementation mechanism, achieves multi-dimensional emotional control of the target object without interfering with lip movements. It can generate rich and nuanced emotional fluctuations, giving the target object more realistic and natural emotional expressions. Furthermore, the facial motion generation method of this application is highly scalable, allowing for the easy introduction of new emotion categories. By simply creating corresponding standard template expressions, the emotion library can be quickly expanded. Simultaneously, the parameter design is flexible, allowing for detailed adjustments based on the actual effect of the target object's emotional expression to adapt to diverse emotional change control needs, thereby enhancing the target object's expressiveness and interactive realism.

[0209] The following description continues to illustrate the exemplary structure of the facial motion generation device 555 provided in the embodiments of this application as a software module. In some embodiments, such as Figure 2 As shown, the software module stored in the facial motion generation device 555 in the memory 550 may include:

[0210] The data acquisition module 5551 is used to acquire the emotion category of the target emotion and the audio text.

[0211] The facial motion data generation module 5552 is used to determine the emotion change parameters corresponding to the emotion category, and generate first facial motion data expressing the target emotion based on the emotion change parameters, wherein the emotion change parameters are used to control the emotion change expressed by the first facial motion data.

[0212] The facial motion data generation module 5552 is further configured to determine second facial motion data for reading the audio text, wherein the first facial motion data includes a first expression locator that is the same as the second expression locator included in the second facial motion data.

[0213] The data fusion module 5553 is used to fuse the first facial motion data and the second facial motion data based on the first expression locator and the second expression locator to obtain target facial motion data.

[0214] In some embodiments, the first facial motion data includes N expression segments, where N is an integer greater than 1. The facial motion data generation module 5552 is further configured to determine the nth initial expression segment of the first facial motion data based on the emotion category, where n is an integer that increments sequentially, 1≤n≤N; and to perform deformation processing on the nth initial expression segment based on the emotion change parameters to obtain the nth expression segment.

[0215] In some embodiments, the emotion change parameters include the degree of emotion change, the amplitude of emotion change, the number of emotion change frames, and the number of emotion retention frames. The nth expression segment includes an emotion change segment and an emotion retention segment. The facial motion data generation module 5552 is further configured to perform the following processing on the standard expression template included in the nth initial expression segment: based on the degree of emotion change, the amplitude of emotion change, and the number of emotion change frames, determine the weight value of the third expression locator included in the I-frame standard expression template, where I is the number of emotion change frames; based on the number of emotion retention frames, determine the weight value of the fourth expression locator included in the J-frame standard expression template, where J is the number of emotion retention frames; based on the weight value of the third expression locator, construct the emotion change segment, and based on the weight value of the fourth expression locator, construct the emotion retention segment.

[0216] In some embodiments, the facial motion data generation module 5552 is further configured to: determine the I-frame standard expression template from the nth initial expression segment based on the number of emotion change frames; determine the range of emotion change based on the degree of emotion change, the amplitude of emotion change, and the initial weight value of the third expression locator included in the I-frame standard expression template; and determine the weight value of the third expression locator included in the I-frame standard expression template based on the range of emotion change.

[0217] In some embodiments, the facial motion data generation module 5552 is further configured to determine the range of maximum emotional value changes based on the degree of emotional change, the amplitude of emotional change, and the initial weight value; determine the random number determined from the range of maximum emotional value changes as the maximum emotional value; determine the weight value of the expression locator included in the last frame of the standard expression template of the (n-1)th expression segment as the minimum emotional value; and construct the range of emotional change based on the maximum emotional value and the minimum emotional value.

[0218] In some embodiments, the facial motion data generation module 5552 is further configured to: when the degree of emotional change is less than a threshold value for emotional change, determine the product of the amplitude of emotional change, the initial weight value, and the first coefficient as a first minimum value; determine a first maximum value based on the first product, the first ratio, and the first minimum value; and construct the range of maximum emotional value variation based on the first maximum value and the first minimum value, wherein the first product is the product of the amplitude of emotional change, the initial weight value, and the second coefficient, and the first ratio is the ratio of the degree of emotional change to the threshold value for emotional change; when the degree of emotional change is greater than or equal to the threshold value for emotional change, determine the product of the amplitude of emotional change and the initial weight value as a second minimum value; determine a second maximum value based on the product of the first difference and the second difference, the threshold value for emotional change, and the second minimum value; and construct the range of maximum emotional value variation based on the second maximum value and the second minimum value, wherein the first difference is the difference between the initial weight value and the second minimum value, and the second difference is the difference between the degree of emotional change and the threshold value for emotional change.

[0219] In some embodiments, the facial motion data generation module 5552 is further configured to perform the following processing for each dimension of the third expression locator: determine I random values ​​from the range of emotional changes; sort the I random values ​​according to the order of emotional changes, and determine the sorted I random values ​​as the weight values ​​of the third expression locator.

[0220] In some embodiments, the facial motion data generation module 5552 is further configured to determine the J-frame standard expression template from the nth initial expression segment based on the emotion retention frame number; determine a target standard expression template located before and adjacent to the J-frame standard expression template; and determine the weight value of the expression locator included in the target standard expression template as the weight value of the fourth expression locator included in the J-frame standard expression template.

[0221] In some embodiments, the data fusion module 5553 is further configured to sum the first weight value and the second weight value to determine the target weight value of the first target expression locator, wherein the first target expression locator is the same expression locator as the second expression locator among the first expression locators, the first weight value is the weight value of the first target expression locator included in the first facial motion data, and the second weight value is the weight value of the first target expression locator included in the second facial motion data; and to combine the target weight value of the first target expression locator and the weight value of the second target expression locator to determine the target facial motion data, wherein the second target expression locator is the expression locator other than the first target expression locator among the first expression locator and the second expression locator.

[0222] In some embodiments, the facial motion data application module 5554 is used to smooth the target facial motion data to obtain smoothed target facial motion data; and to apply the smoothed target facial motion data to the face of the target object.

[0223] This application provides a computer program product, which includes a computer program or computer-executable instructions stored in a computer-readable storage medium. The processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the facial motion generation method described above in this application.

[0224] This application provides a computer-readable storage medium storing computer-executable instructions or a computer program. When the computer-executable instructions or the computer program are executed by a processor, the processor will execute the facial motion generation method provided in this application. For example, ... Figures 3 to 8 The method for generating facial movements is shown.

[0225] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a variety of devices including one or any combination of the above-mentioned memories.

[0226] In some embodiments, computer-executable instructions may take the form of programs, software, software modules, scripts, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as stand-alone programs or as modules, components, subroutines, or other units suitable for use in a computing environment.

[0227] As an example, computer-executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple co-located files (e.g., files that store one or more modules, subroutines, or code sections).

[0228] As an example, computer-executable instructions can be deployed to execute on a single electronic device, or on multiple electronic devices located at one location, or on multiple electronic devices distributed across multiple locations and interconnected via a communication network.

[0229] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.

Claims

1. A method for generating facial movements, characterized in that, The method includes: Obtain the emotion category and audio text of the target emotion; Determine the emotion change parameters corresponding to the emotion category, wherein the emotion change parameters are used to control the emotion change expressed by the first facial motion data, and the emotion change parameters include the degree of emotion change, the amplitude of emotion change, the number of emotion change frames, and the number of emotion holding frames; Based on the emotion category, N initial expression segments are determined. For the standard expression template included in the nth initial expression segment among the N initial expression segments, the following processing is performed: Based on the degree of emotion change, the amplitude of emotion change, and the number of emotion change frames, the weight value of the third expression locator included in the I-frame standard expression template is determined; based on the number of emotion retention frames, the weight value of the fourth expression locator included in the J-frame standard expression template is determined; based on the weight value of the third expression locator, an emotion change segment is constructed, and based on the weight value of the fourth expression locator, an emotion retention segment is constructed; the nth expression segment is constructed from the emotion change segment and the emotion retention segment; where N is an integer greater than 1, n is an incrementing integer, 1≤n≤N, I is the number of emotion change frames, and J is the number of emotion retention frames; The first facial movement data, which consists of N facial expression fragments, constitutes the target emotion; Determine second facial motion data for reading the audio text aloud, wherein the first facial motion data includes a first expression locator that is the same as the second expression locator portion included in the second facial motion data; Based on the first expression locator and the second expression locator, the first facial motion data and the second facial motion data are fused to obtain the target facial motion data.

2. The method according to claim 1, characterized in that, The determination of the weight value of the third expression locator included in the I-frame standard expression template based on the degree of emotional change, the amplitude of emotional change, and the number of emotional change frames includes: Based on the number of emotion change frames, the standard expression template for frame I is determined from the nth initial expression segment; The range of emotional change is determined based on the degree of emotional change, the amplitude of emotional change, and the initial weight value of the third expression locator included in the I-frame standard expression template; Based on the range of emotional changes, the weight value of the third expression locator included in the I-frame standard expression template is determined.

3. The method according to claim 2, characterized in that, The process of determining the range of emotional change based on the degree of emotional change, the amplitude of emotional change, and the initial weight value of the third expression locator included in the I-frame standard expression template includes: Based on the degree of emotional change, the magnitude of emotional change, and the initial weight value, the range of variation of the maximum emotional value is determined; The random number determined from the range of the maximum emotional value is defined as the maximum emotional value; The weight value of the expression locator included in the last frame of the standard expression template of the (n-1)th expression fragment is determined as the minimum value of emotion. The range of emotional variation is defined by the maximum and minimum emotional values.

4. The method according to claim 3, characterized in that, The process of determining the range of variation of the maximum emotional value based on the degree of emotional change, the magnitude of emotional change, and the initial weight value includes: When the degree of emotional change is less than the threshold of emotional change, the product of the amplitude of emotional change, the initial weight value, and the first coefficient is determined as the first minimum value. Based on the first product, the first ratio, and the first minimum value, the first maximum value is determined. Based on the first maximum value and the first minimum value, the range of the maximum emotional value is formed. The first product is the product of the amplitude of emotional change, the initial weight value, and the second coefficient. The first ratio is the ratio of the degree of emotional change to the threshold of emotional change. When the degree of emotional change is greater than or equal to the threshold value of emotional change, the product of the amplitude of emotional change and the initial weight value is determined as the second minimum value. Based on the product of the first difference and the second difference, the threshold value of emotional change, and the second minimum value, the second maximum value is determined. Based on the second maximum value and the second minimum value, the range of the maximum emotional value is formed. The first difference is the difference between the initial weight value and the second minimum value, and the second difference is the difference between the degree of emotional change and the threshold value of emotional change.

5. The method according to claim 2, characterized in that, The step of determining the weight value of the third expression locator included in the I-frame standard expression template based on the range of emotional changes includes: Perform the following processing on the third expression locator for each dimension: I random values ​​are determined from the range of emotional changes; The I random values ​​are sorted according to the order of emotional changes, and the sorted I random values ​​are determined as the weight values ​​of the third expression locator.

6. The method according to claim 1, characterized in that, The step of determining the weight value of the fourth expression locator included in the J-frame standard expression template based on the emotion preservation frame number includes: Based on the number of emotion-preserving frames, the standard expression template for frame J is determined from the nth initial expression segment; Determine the target standard expression template that is located before and adjacent to the standard expression template of frame J; The weight values ​​of the expression locators included in the target standard expression template are determined as the weight values ​​of the fourth expression locator included in the J-frame standard expression template.

7. The method according to any one of claims 1-6, characterized in that, The step of fusing the first facial expression locator and the second facial expression locator with the first facial motion data and the second facial motion data to obtain target facial motion data includes: The sum of the first weight value and the second weight value is determined as the target weight value of the first target expression locator, wherein the first target expression locator is the same expression locator as the second expression locator in the first expression locator, the first weight value is the weight value of the first target expression locator included in the first facial motion data, and the second weight value is the weight value of the first target expression locator included in the second facial motion data. The target weight value of the first target expression locator and the weight value of the second target expression locator are combined to determine the target facial action data, wherein the second target expression locator is the expression locator other than the first target expression locator among the first expression locator and the second expression locator.

8. The method according to any one of claims 1-6, characterized in that, After fusing the first facial motion data and the second facial motion data based on the first expression locator and the second expression locator to obtain the target facial motion data, the method further includes: The target facial motion data is smoothed to obtain smoothed target facial motion data. The smoothed facial motion data is applied to the target object's face.

9. A facial motion generation device, characterized in that, The device includes: The data acquisition module is used to obtain the emotion category of the target emotion and the audio text; A facial motion data generation module is used to determine the emotion change parameters corresponding to the emotion category. These emotion change parameters control the emotion changes expressed by the first facial motion data and include the degree of emotion change, the amplitude of emotion change, the number of emotion change frames, and the number of emotion maintenance frames. Based on the emotion category, N initial expression segments are determined. For the nth initial expression segment among the N initial expression segments, the following processing is performed on the standard expression template: based on the degree of emotion change, the amplitude of emotion change, and the number of emotion change frames, a third expression is determined from the I-frame standard expression template. The weight value of the locator; based on the number of emotion-preserving frames, determine the weight value of the fourth expression locator included in the J-frame standard expression template; based on the weight value of the third expression locator, construct an emotion change segment, and based on the weight value of the fourth expression locator, construct an emotion preservation segment; the emotion change segment and the emotion preservation segment constitute the nth expression segment; where N is an integer greater than 1, n is an integer that increases sequentially, 1≤n≤N, I is the number of emotion change frames, and J is the number of emotion preservation frames; the N expression segments constitute the first facial motion data of the target emotion; The facial motion data generation module is further configured to determine second facial motion data for reading the audio text, wherein the first facial motion data includes a first expression locator that is the same as the second expression locator included in the second facial motion data; The data fusion module is used to fuse the first facial motion data and the second facial motion data based on the first expression locator and the second expression locator to obtain the target facial motion data.

10. An electronic device, characterized in that, The electronic device includes: Memory is used to store executable instructions or computer programs. A processor, when executing computer-executable instructions or computer programs stored in the memory, implements the facial motion generation method according to any one of claims 1 to 8.

11. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or computer program are executed by a processor, the facial motion generation method according to any one of claims 1 to 8 is implemented.

12. A computer program product comprising computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or computer program are executed by a processor, the facial motion generation method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Mouth shape driving method, device and equipment of digital human and storage medium

    CN117935807A