A method, device and equipment for generating eyebrow animation and a storage medium
By segmenting audio into audio segments with sound, without sound, with expression, and without expression, and using expression-driven AU rules to generate eyebrow animation curves, the problem of environmental factors and rigidity in the generation of virtual digital human eyebrow animation is solved, thus improving the realism of expressions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TENCENT TECHNOLOGY (SHENZHEN) CO LTD
- Filing Date
- 2022-08-15
- Publication Date
- 2026-05-29
AI Technical Summary
In existing technologies, the generation of eyebrow animations for 3D virtual digital humans is easily affected by unclear environmental factors, and the expressions are stiff when there is no expression or sound, making it difficult to accurately express emotions.
The audio to be processed is segmented into segments with sound, without sound, with expression, and without expression. The eyebrow animation curve for each target time segment is obtained through expression-driven AU rules and then spliced together to generate a more accurate eyebrow animation.
By eliminating the influence of unclear environmental factors and simplifying the generation process, we ensure that the eyebrow animation matches the audio scene and enhances the realism of the virtual object's facial expressions.
Smart Images

Figure CN115359160B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device, and storage medium for generating eyebrow animation. Background Technology
[0002] With the rapid development of deep learning, especially computer vision technology, computer vision technology has been widely applied in many fields such as security, healthcare, and entertainment. Virtual reality technology, as a more advanced computer vision technology, has become a current research hotspot.
[0003] Currently, common 3D virtual digital humans can communicate normally with humans through voice recognition technology. However, how to make the virtual image of a 3D virtual digital human have realistic expressions, smooth and natural facial features, and non-mechanical movement changes remains a challenge in terms of intelligence and graphics.
[0004] The common facial expression driving technology JALI is based on a layered implementation. The eyebrow animation layer needs to use the auxiliary language paralingual to exclude other information from the semantic information of words in the audio, such as tone of voice, intonation, and breath, and then combine it with the character positioning to generate facial expression animation.
[0005] However, in paralingual, the quantitative indicators that can represent the context, especially tone, are unclear and easily affected by some unclear environmental factors, which can lead to deviations in the generated eyebrow animation. In addition, when the character is in a calm state, the character positioning does not take into account the current eyebrow movement state, so the positioning of the character may also lead to judgment bias, resulting in the generated facial expression driving not being able to accurately express emotions and not being vivid enough. Summary of the Invention
[0006] This application provides a method, apparatus, device, and storage medium for generating eyebrow animation. It fragments different emotional states to more accurately obtain the corresponding segmented eyebrow animation curve for each emotional state, generating eyebrow animation corresponding to the audio to be processed. This not only eliminates the influence of unclear environmental factors on the overall facial expression animation and simplifies the automatic generation process of eyebrow animation, but also considers the natural movement state of expressionless or silent eyebrows based on expression-driven AU rules. This avoids stiff facial expressions when expressionless or silent, ensuring that the eyebrow animation fits the scene setting of the audio to be processed, thereby enhancing the realism of the virtual object's facial expression-driven animation.
[0007] One embodiment of this application provides a method for generating eyebrow animation, including:
[0008] The audio to be processed is divided into audio segments and silent segments. Each audio segment corresponds to an audio start and end timestamp, and each silent segment corresponds to a silent start and end timestamp.
[0009] Based on the emoji tags corresponding to the audio to be processed, the audio to be processed is divided into segments with emojis and segments without emojis. The emoji tags include the emoji type and the start and end timestamps of each emoji type. Each emoji segment corresponds to an emoji type and an emoji start and end timestamp, and each segment without emojis corresponds to a segment without emojis.
[0010] Based on the start and end timestamps of audio, silent, facial expression, and no facial expression, the intersection of audio segments, silent segments, facial expression segments, and no facial expression segments is obtained to obtain the emotional state corresponding to each target time period. The emotional state includes no facial expression and no audio state, facial expression and no audio state, no facial expression and audio state, or facial expression and audio state.
[0011] Based on the expression-driven AU rule, determine the AU of the eyebrows associated with the emotional state for each target time period;
[0012] Interpolate the AU of the eyebrows associated with the emotional state for each target time period to obtain the eyebrow animation curve for each segment corresponding to the emotional state.
[0013] The eyebrow animation curves of each segment are spliced together to obtain the target eyebrow animation curve;
[0014] Based on the target eyebrow animation curve, generate the eyebrow animation corresponding to the audio to be processed.
[0015] This application also provides an apparatus for generating eyebrow animation, comprising:
[0016] The processing unit is used to divide the audio to be processed into audio segments and silent segments, wherein each audio segment corresponds to an audio start and end timestamp, and each silent segment corresponds to a silent start and end timestamp.
[0017] The processing unit is also used to divide the audio to be processed into expressive segments and non-expressive segments based on the expression tags corresponding to the audio to be processed. The expression tags include expression types and expression start and end timestamps corresponding to each expression type. Each expression segment corresponds to an expression type and an expression start and end timestamp, and a non-expressive segment corresponds to a non-expressive start and end timestamp.
[0018] The acquisition unit is used to acquire the intersection of audio segments, silent segments, expressive segments, and expressionless segments based on the audio start and end timestamps, silent start and end timestamps, expressive start and end timestamps, and expressionless start and end timestamps, to obtain the emotional state corresponding to the target time period. The emotional state includes an expressionless and silent state, an expressive and silent state, an expressionless and audio state, or an expressive and audio state.
[0019] The unit is defined to determine the AU of the eyebrows associated with the emotional state for each target time period based on the expression-driven AU rule.
[0020] The processing unit is also used to interpolate the AU of the eyebrows associated with the emotional state corresponding to each target time period to obtain the eyebrow animation curve of each segment corresponding to the emotional state.
[0021] The processing unit is also used to splice the eyebrow animation curves of each segment to obtain the target eyebrow animation curve;
[0022] The processing unit is also used to generate eyebrow animation corresponding to the audio to be processed based on the target eyebrow animation curve.
[0023] In one possible design, in another implementation of the embodiments of this application, the determining unit may specifically be used for:
[0024] Multiple first random seed points on the segment corresponding to the expressionless and silent state are calculated using a random function;
[0025] For each of the multiple first random seed points, extract the extended segment corresponding to each first random seed point from the segment corresponding to the expressionless and silent state;
[0026] Based on the expression-driven AU rule, the default AU of the eyebrow movement state associated with the expressionless state is selected as the AU of the eyebrow.
[0027] Specifically, the processing unit can be used to perform interpolation calculations based on each first random seed point, the extended segment corresponding to each first random seed point, and each default AU to obtain the eyebrow animation curve corresponding to the expressionless and silent state.
[0028] In one possible design, in another implementation of the embodiments of this application, the determining unit may specifically be used for:
[0029] Multiple second random seed points on the segment corresponding to the expressionless but silent state are calculated using a random function;
[0030] For each of the multiple second random seed points, extract the extended segment corresponding to each second random seed point from the segment corresponding to the expressionless and silent state;
[0031] Based on the expression-driven AU rule, the first expression AU of the eyebrow movement state associated with the expression type in the expressionless state is selected as the AU of the eyebrow.
[0032] The processing unit can specifically be used for:
[0033] By using fuzzy rules, the weight of the first AU corresponding to each first expression AU is obtained;
[0034] Interpolation calculations are performed based on each second random seed point, the corresponding extended segment of each second random seed point, and each first expression AU. The interpolation calculation results are then weighted based on the weight of each first AU to obtain the eyebrow animation curve corresponding to the expressionless state.
[0035] In one possible design, in another implementation of the embodiments of this application, the determining unit may specifically be used for:
[0036] Multiple third random seed points on the segment corresponding to the expressionless but vocal state are calculated using a random function;
[0037] For each of the multiple third random seed points, extract the extended segment corresponding to each third random seed point from the segment corresponding to the expressionless but vocal state;
[0038] Based on the expression-driven AU rule, the default AU of the eyebrow movement state associated with the expressionless state is selected as the AU of the eyebrow.
[0039] The processing unit can specifically be used for:
[0040] Extract the sound signal from the segment corresponding to the audio state;
[0041] The sound signal is smoothed to obtain the smoothed value, and the smoothed value is used as the AU adjustment weight.
[0042] Interpolation calculations are performed based on each third random seed point, the corresponding extended segment of each third random seed point, and each default AU. The interpolation calculation results are then weighted based on the AU adjustment weights to obtain the eyebrow animation curve corresponding to the expressionless yet vocal state.
[0043] In one possible design, in another implementation of the embodiments of this application, the determining unit may specifically be used for:
[0044] Multiple fourth random seed points on the segment corresponding to the expressive and vocal state are calculated using a random function;
[0045] For each of the multiple fourth random seed points, extract the extended segment corresponding to each fourth random seed point from the segment corresponding to the expressive and vocal state;
[0046] Based on the expression-driven AU rule, the second expression AU of the eyebrow movement state associated with the expression type in the expression and voice state is selected as the AU of the eyebrow.
[0047] The processing unit can specifically be used for:
[0048] Extract the sound signal from the segment corresponding to the audio state;
[0049] The sound signal is smoothed to obtain the smoothed value, and the smoothed value is used as the AU adjustment weight.
[0050] Interpolation calculations are performed based on each fourth random seed point, the corresponding extended segment of each fourth random seed point, and each second expression AU. The interpolation calculation results are then weighted based on the AU adjustment weights to obtain the eyebrow animation curve of the segment corresponding to the expressive and vocal state.
[0051] In one possible design, in another implementation of the embodiments of this application, the determining unit may specifically be used for:
[0052] Extract the sound intensity signal from the segment corresponding to the sound state;
[0053] Specifically, the unit can be used for:
[0054] The sound intensity signal is filtered by a low-pass filter to obtain a smooth distribution of the sound intensity signal;
[0055] The smooth distribution of the sound intensity signal is normalized to obtain the smooth value corresponding to the sound intensity signal, and the smooth value is used as the AU adjustment weight.
[0056] In one possible design, in another implementation of the embodiments of this application, the determining unit may specifically be used for:
[0057] Extract the pitch signal from the segment corresponding to the audio state;
[0058] Specifically, the unit can be used for:
[0059] The pitch signal is filtered by a low-pass filter to obtain a smooth distribution of the pitch signal;
[0060] The smooth distribution of the pitch signal is normalized to obtain the smooth value corresponding to the pitch signal, and the smooth value is used as the AU adjustment weight.
[0061] In one possible design, in another implementation of the embodiments of this application, the determining unit may specifically be used for:
[0062] Extract the intensity and pitch signals from the segments corresponding to the audio state;
[0063] Specifically, the unit can be used for:
[0064] By using a low-pass filter, the intensity signal and the pitch signal are filtered separately to obtain a smooth distribution of the intensity signal and a smooth distribution of the pitch signal.
[0065] The smooth distribution of the sound intensity signal is normalized to obtain the smooth value corresponding to the sound intensity signal, and the smooth value is used as the adjustment weight of the first AU.
[0066] The smooth distribution of the pitch signal is normalized to obtain the smooth value corresponding to the pitch signal, and the smooth value is used as the adjustment weight of the second AU.
[0067] In one possible design, in another implementation of the embodiments of this application,
[0068] The acquisition unit is also used to extract accented segments and intensity signals from the segments corresponding to the audio state.
[0069] The processing unit is also used to use the ratio of the intensity signal in the accented segment to the intensity threshold as an enhancement coefficient;
[0070] The processing unit is also used to adjust the AU adjustment weight based on the enhancement coefficient to obtain the target AU adjustment weight.
[0071] In one possible design, in another implementation of the embodiments of this application, the processing unit may specifically be used for:
[0072] According to the order of the target time periods corresponding to each segment of the eyebrow animation curve, the eyebrow animation curves of each segment are spliced together to obtain the spliced eyebrow animation curve.
[0073] For each pair of adjacent eyebrow animation curves in the spliced eyebrow animation curve, the corresponding segment curve at the splicing point is obtained according to the basic window length, and the segment curve is Gaussian smoothed to obtain the target eyebrow animation curve.
[0074] This application also provides a computer device, including: a memory, a processor, and a bus system;
[0075] The memory is used to store programs;
[0076] The processor implements the methods described above when executing a program in memory;
[0077] Bus systems are used to connect memory and processor to enable communication between them.
[0078] Another aspect of this application provides a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the methods described above.
[0079] As can be seen from the above technical solutions, the embodiments of this application have the following beneficial effects:
[0080] By segmenting the audio to be processed into audio segments and silent segments, and further segmenting it into expressive segments and non-expressive segments, the intersection of the audio segments, silent segments, expressive segments, and non-expressive segments is obtained to acquire the emotional state corresponding to each target time period. The emotional state includes a non-expressive and silent state, an expressive and silent state, a non-expressive and expressive state, or an expressive and expressive state. Then, based on the expression-driven AU rule, the eyebrow animation curve corresponding to each emotional state is obtained and spliced together to form the target eyebrow animation curve. Based on the target eyebrow animation curve, the eyebrow animation corresponding to the audio to be processed is generated. The above method allows for the segmentation of the audio to be processed, dividing it into audio segments with sound, silent segments, expressive segments, and expressionless segments. The intersection of these segments yields the emotional state corresponding to each target time period. Then, based on expression-driven AU rules, different emotional states are segmented to more accurately obtain the corresponding eyebrow animation curves for each emotional state. These eyebrow animation curves are then concatenated to form the target eyebrow animation curve. Based on this target curve, the corresponding eyebrow animation for the audio to be processed is generated. This not only eliminates the influence of ambiguous environmental factors on the overall facial expression animation and simplifies the automatic generation process of eyebrow animation, but also considers the natural movement of eyebrows in expressionless or silent situations based on expression-driven AU rules. This avoids stiff facial expressions in these cases, ensuring that the eyebrow animation fits the scene setting of the audio to be processed, thus enhancing the realism of the virtual object's facial expression-driven animation. Attached Figure Description
[0081] Figure 1 This is a schematic diagram of the architecture of the animation data control system in an embodiment of this application;
[0082] Figure 2 This is a flowchart of one embodiment of the eyebrow animation generation method in this application;
[0083] Figure 3 This is a flowchart of another embodiment of the eyebrow animation generation method in this application;
[0084] Figure 4 This is a flowchart of another embodiment of the eyebrow animation generation method in this application;
[0085] Figure 5 This is a flowchart of another embodiment of the eyebrow animation generation method in this application;
[0086] Figure 6 This is a flowchart of another embodiment of the eyebrow animation generation method in this application;
[0087] Figure 7 This is a flowchart of another embodiment of the eyebrow animation generation method in this application;
[0088] Figure 8 This is a flowchart of another embodiment of the eyebrow animation generation method in this application;
[0089] Figure 9 This is a flowchart of another embodiment of the eyebrow animation generation method in this application;
[0090] Figure 10 This is a flowchart of another embodiment of the eyebrow animation generation method in this application;
[0091] Figure 11 This is a flowchart of another embodiment of the eyebrow animation generation method in this application;
[0092] Figure 12 This is a schematic diagram illustrating the principle of the eyebrow animation generation method in the embodiments of this application;
[0093] Figure 13 This is a schematic diagram of one embodiment of the eyebrow animation generation device in this application;
[0094] Figure 14 This is a schematic diagram of one embodiment of the computer device described in this application. Detailed Implementation
[0095] This application provides a method, apparatus, device, and storage medium for generating eyebrow animation. It fragments different emotional states to more accurately obtain the corresponding segmented eyebrow animation curve for each emotional state, generating eyebrow animation corresponding to the audio to be processed. This not only eliminates the influence of unclear environmental factors on the overall facial expression animation and simplifies the automatic generation process of eyebrow animation, but also considers the natural movement state of expressionless or silent eyebrows based on expression-driven AU rules. This avoids stiff facial expressions when expressionless or silent, ensuring that the eyebrow animation fits the scene setting of the audio to be processed, thereby enhancing the realism of the virtual object's facial expression-driven animation.
[0096] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and drawings of this application are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented, for example, in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “corresponding to,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0097] To facilitate understanding, some terms or concepts involved in the embodiments of this application will be explained first.
[0098] 1. JALI: A speech-driven facial animation solution based on a visual pixel model, which integrates elements such as facial expression animation rules, human observation, and AI.
[0099] 2. Visual Pixel Model: A model of different basic pronunciation mouth shapes.
[0100] 3. Facial Action Coding System (FACS)
[0101] Facial motion coding system is a system that classifies and encodes facial muscle movements based on facial expressions.
[0102] 4. Action Unit (AU)
[0103] An AU is a basic model derived from the basic movements of a single muscle or a group of muscles. Different combinations of AUs can produce different facial expressions.
[0104] 5. Time alignment tool (Montreal Forced Aligner, MFA):
[0105] MFA is a tool that attempts to time-align input spoken audio with corresponding script phonemes.
[0106] It is understood that in the specific implementation of this application, data such as audio to be processed and emoji tags are involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0107] It is understandable that the eyebrow animation generation method disclosed in this application specifically involves Intelligent Vehicle Infrastructure Cooperative Systems (IVICS). The following is a further introduction to IVICS. IVICS, or vehicle-road cooperative systems for short, is a development direction of Intelligent Transportation Systems (ITS). IVICS utilizes advanced wireless communication and next-generation Internet technologies to implement comprehensive, real-time dynamic information interaction between vehicles and roads. Based on the collection and fusion of dynamic traffic information across all times and spaces, it conducts active vehicle safety control and cooperative road management, fully realizing effective coordination between people, vehicles, and roads, ensuring traffic safety, and improving traffic efficiency, thereby forming a safe, efficient, and environmentally friendly road traffic system.
[0108] Understandably, the eyebrow animation generation method disclosed in this application also involves artificial intelligence (AI) technology, which will be further introduced below. Artificial intelligence is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results. In other words, artificial intelligence is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new type of intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling them to possess perception, reasoning, and decision-making functions.
[0109] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0110] Secondly, Natural Language Processing (NLP) is an important field within computer science and artificial intelligence. It studies the theories and methods that enable effective communication between humans and computers using natural language. NLP is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field involves natural language—the language people use in daily life—and thus it has a close connection with linguistic research. NLP technologies typically include text processing, semantic understanding, machine translation, question answering, and knowledge graphs.
[0111] Secondly, Machine Learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory, among others. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instructional learning.
[0112] With the research and advancement of artificial intelligence (AI) technology, AI is being studied and applied in various fields, such as smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, autonomous driving, drones, robots, smart healthcare, and smart customer service. It is believed that with the development of technology, AI will be applied in more fields and play an increasingly important role.
[0113] It should be understood that the eyebrow animation generation method provided in this application can be applied to various scenarios, including but not limited to artificial intelligence, maps, smart transportation, cloud technology, etc., to generate eyebrow animation through eyebrow animation curves, complete the facial expression animation of virtual objects, and can be applied to scenarios such as virtual idols, role-playing games, 3D animation production of virtual digital humans, virtual anchors on live streaming platforms, virtual customer service for smart shopping, virtual butlers for smart transportation navigation, or virtual butlers for smart homes.
[0114] To address the aforementioned problems, this application proposes a method for generating eyebrow animation, which is applied to... Figure 1 Please refer to the animation object control system shown. Figure 1 , Figure 1 This is a schematic diagram of the architecture of the animation object control system in an embodiment of this application, such as... Figure 1As shown, the server obtains the audio to be processed provided by the terminal device, segments the audio into audio segments with sound and silent segments, and further segments the audio into segments with facial expressions and segments without facial expressions. Then, it obtains the intersection of the audio segments with sound, silent segments, segments with facial expressions, and segments without facial expressions to obtain the emotional state corresponding to each target time period. The emotional state includes a silent state without facial expressions, a silent state with facial expressions, a voiceless state without facial expressions, or a voiceless state with facial expressions. Then, based on the expression-driven AU rule, it obtains the eyebrow animation curve of the segment corresponding to each emotional state and splices them into the target eyebrow animation curve. Based on the target eyebrow animation curve, it generates the eyebrow animation corresponding to the audio to be processed. The above method allows for the segmentation of the audio to be processed, dividing it into audio segments with sound, silent segments, expressive segments, and expressionless segments. The intersection of these segments yields the emotional state corresponding to each target time period. Then, based on expression-driven AU rules, different emotional states are segmented to more accurately obtain the corresponding eyebrow animation curves for each emotional state. These eyebrow animation curves are then concatenated to form the target eyebrow animation curve. Based on this target curve, the corresponding eyebrow animation for the audio to be processed is generated. This not only eliminates the influence of ambiguous environmental factors on the overall facial expression animation and simplifies the automatic generation process of eyebrow animation, but also considers the natural movement of eyebrows in expressionless or silent situations based on expression-driven AU rules. This avoids stiff facial expressions in these cases, ensuring that the eyebrow animation fits the scene setting of the audio to be processed, thus enhancing the realism of the virtual object's facial expression-driven animation.
[0115] Understandable, Figure 1 Only one type of terminal device is shown in the diagram. In real-world scenarios, many more types of terminal devices can participate in the data processing. These include, but are not limited to, mobile phones, computers, smart voice interaction devices, smart home appliances, and in-vehicle terminals. The specific number and types depend on the actual scenario and are not limited here. Furthermore, Figure 1 The diagram shows one server, but in real-world scenarios, multiple servers can be involved, especially in scenarios involving multi-model training and interaction. The number of servers depends on the specific scenario and is not limited here.
[0116] It should be noted that in this embodiment, the server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. Terminal devices and servers can be directly or indirectly connected via wired or wireless communication, and terminal devices and servers can be connected to form a blockchain network; this application does not impose any limitations on this.
[0117] Based on the above introduction, the method for generating eyebrow animation in this application will be described below. Please refer to [link / reference]. Figure 2 One embodiment of the eyebrow animation generation method in this application includes:
[0118] In step S101, the audio to be processed is divided into audio segments and silent segments, wherein each audio segment corresponds to an audio start and end timestamp, and each silent segment corresponds to a silent start and end timestamp.
[0119] In this embodiment, in order to better refine the eyebrow movement state under different conditions and thus obtain a more accurate eyebrow animation, the audio to be processed can be first processed by audio segmentation, that is, the audio to be processed can be divided into audio segments and silent segments.
[0120] The audio to be processed consists of pre-captured audio that can be used to generate eyebrow animation and facial expression animation for virtual objects. Each audio segment corresponds to an audio start and end timestamp, and each silent segment corresponds to a silent start and end timestamp.
[0121] Specifically, such as Figure 12 As shown, the audio to be processed is divided into audio segments and silent segments (such as...). Figure 12 (The extraction of audio and silent segments as shown) can be achieved by using the Praat tool to segment the audio signal to be processed into audio and silent segments, or by using other audio editing tools to segment audio and silent segments; no specific limitations are made here.
[0122] For example, suppose the start timestamp of an audio file to be processed is "00:00:00" and the end timestamp is "10:00:00". If the audio file to be processed can be divided into an audio segment and a silent segment, the start and end timestamps of the audio segment are "00:00:00" and "05:00:00" respectively, and the start and end timestamps of the silent segment are "05:00:01" and "10:00:00" respectively.
[0123] In step S102, based on the expression tags corresponding to the audio to be processed, the audio to be processed is divided into expression segments and expressionless segments. The expression tags include expression types and the start and end timestamps of each expression type. Each expression segment corresponds to an expression type and an expression start and end timestamp, and the expressionless segment corresponds to an expressionless start and end timestamp.
[0124] In this embodiment, in order to better refine the eyebrow movement state under different conditions and thus obtain a more accurate eyebrow animation, the audio to be processed can be segmented into expressive segments and non-expressive segments based on the expression tags corresponding to the audio to be processed.
[0125] The emoticon tags are pre-defined based on actual application requirements. Each emoticon tag can include the emoticon type, a start and end timestamp for each emoticon type, and the emoticon intensity for each emoticon type. These tags can be stored in an XML markup file. Emoticon intensity reflects the presentation of each emoticon on the virtual object's facial animation. Emoticon intensity can be categorized as very small (vs), small (s), medium (m), large (b), and very large (vb), etc. For example, the intensity of a raised eyebrow corresponding to the happy emoticon type is large (b). Each emoticon segment corresponds to an emoticon type and an emoticon start and end timestamp, while each expressionless segment corresponds to an expressionless start and end timestamp.
[0126] Specifically, such as Figure 12 As shown, based on the emoji tags corresponding to the audio to be processed (such as...) Figure 12 The audio to be processed is divided into expressive segments and non-expressive segments according to the start and end timestamps of each expressive type (the indicated emoji tags).
[0127] For example, suppose the start timestamp of the audio to be processed is "00:00:00" and the end timestamp is "10:00:00". If the emoji tag contains two emoji types, and the start and end timestamps of each emoji type are "00:00:00" and "02:00:00" and "04:00:01" and "08:00:00" respectively, then the start and end timestamps of the expressionless segment are "02:00:01" and "04:00:00" and "08:00:01" and "10:00:00". That is, the audio to be processed is divided into two segments with expressions and two segments without expressions.
[0128] In step S103, based on the start and end timestamps of the audio segment, the start and end timestamps of the silent segment, the start and end timestamps of the facial expression segment, and the start and end timestamps of the expressionless segment, the intersection of the audio segment, the silent segment, the facial expression segment, and the expressionless segment is obtained to obtain the emotional state corresponding to each target time period. The emotional state includes the expressionless and silent state, the facial expressionless and silent state, the expressionless and audio state, or the facial expression and audio state.
[0129] In this embodiment, in order to better enrich the expression-driven presentation and avoid the expression state from being rigid when there is no expression or no sound, the intersection of the audio segment, the silent segment, the expression segment, and the expressionless segment can be obtained based on the audio start and end timestamp, the silent start and end timestamp, the expression start and end timestamp, and the expressionless start and end timestamp to obtain the emotional state corresponding to each target time period.
[0130] As shown in Table 1 below, emotional states include expressionless and silent state (state 1 as shown in Table 1), expressionful and silent state (state 2 as shown in Table 1), expressionless and vocal state (state 3 as shown in Table 1), or expressionful and vocal state (state 4 as shown in Table 1).
[0131] expressionless With expressions silent State 1 State 2 Audio State 3 State 4
[0132] Table 1
[0133] The target time period is the period formed by the intersection of the start and end timestamps of the audio start and end timestamps, the silent start and end timestamps, the facial expression start and end timestamps, and the expressionless start and end timestamps. For example, if an audio start and end timestamp is "00:00:00" and "05:00:00", and an facial expression start and end timestamp is "00:00:00" and "02:00:00", the start and end timestamps obtained by taking the intersection are "00:00:00" and "02:00:00", meaning the target time period is from "00:00:00" to "02:00:00".
[0134] Specifically, such as Figure 12 As shown, after obtaining the audio segments with sound, without sound, with expression, and without expression, further finer segmentation can be performed on the audio to be processed. Specifically, this can be done by taking the intersection of the audio segments, the silent segments, the expressive segments, and the expressionless segments based on the start and end timestamps of the audio segments, the silent segments, the expressive segments, and the expressionless segments (e.g., ...). Figure 12 The intersection of the indicated facial expressions and audio clips is used to obtain the emotional state corresponding to each target time period (e.g., Figure 12 (The indicated state is 1, state 2, state 3, or state 4).
[0135] For example, continuing with the above example, suppose the start and end timestamps for facial expressions are "00:00:00" and "02:00:00", "04:00:01" and "08:00:00", the start and end timestamps for expressionless faces are "02:00:01" and "04:00:00", "08:00:01" and "10:00:00", the start and end timestamps for spoken segments are "00:00:00" and "05:00:00", and the start and end timestamps for silent segments are "05:00:01" and "10:00:00". Taking the intersection of the time periods corresponding to each start and end time period, we can obtain 5 target time periods and 5 time intervals. The target time period corresponds to the following emotional states: "expression and voice" for the target time period from "00:00:00" to "02:00:00", "expressionless but voice" for the target time period from "02:00:01" to "04:00:00", "expressionless but voice" for the target time period from "04:00:01" to "05:00:00", "expression and voice" for the target time period from "05:00:01" to "08:00:00", "expressionless but voiceless" for the target time period from "08:00:01" to "10:00:00".
[0136] In step S104, based on the expression-driven AU rule, the AU of the eyebrows associated with the emotional state corresponding to each target time period is determined;
[0137] In this embodiment, after obtaining the emotional state corresponding to each target time period, the AU (emotional signature) of the eyebrows associated with the emotional state corresponding to each target time period can be obtained based on the expression-driven AU rule. Figure 12 The rules for eyebrow animation (AU) are shown to enable subsequent generation of eyebrow animations more accurately and better based on the expression-driven AU that represents the eyebrow movement state associated with each emotional state, while maintaining the independence of eyebrow animations.
[0138] Specifically, the expression-driven AU rules are AU-driven rules designed based on the facial muscle movement specifications defined in FACS to represent various expressions under seven emotions. These seven emotions can be specifically expressed as follows: Joy: including feelings of liking, happiness, fondness, fondness, delight, and joy; Anger: including feelings of rage, annoyance, fury, resentment, and bitterness; Sorrow: including feelings of sadness, grief, grief, pity, sorrow, melancholy, and grief; Pleasure: referring to a joyful, pleasant, and happy feeling; Surprise: referring to feelings of astonishment, astonishment, panic, fear, surprise, amazement, and surprise; Fear: referring to feelings of panic, fear, apprehension, worry, and dread; and Longing: referring to feelings of yearning, remembrance, and yearning. FACS is a system that classifies and encodes facial muscle movements based on facial expressions.
[0139] Expression-driven AU rules can define the corresponding AU intensity for different expression types and intensities. For example, the AU rules for the "confused" expression are shown in Table 2 below:
[0140] Facial intensity AU1 ... AU4 ... AU17 ... AU23 ... vs s ... m ... vs ... vs ... s s ... m ... vs ... s ... m m ... m ... s ... m ... b m ... b ... m ... b ... vb b ... vb ... b ... m ...
[0141] Table 2
[0142] Among them, very small (vs), small (s), medium (m), big (b), and very big (vb) are understood to mean that, as shown in Table 2, when the expression is "confused" and the expression intensity is vs, the weight of AU1 is s, and when the expression is "confused" and the expression intensity is b, the weight of AU1 is m. Similarly, it can be seen that each AU has a weight corresponding to the "confused" expression and different expression intensities.
[0143] Furthermore, based on the AU rules under the "confused" expression, fuzzy rules are used to calculate the corresponding AU weights driven by the expression. At the same time, AUs related to the eyebrow movement state can be selected from several AUs under the "confused" expression as driving units for eyebrow animation. For example, in the corner of the eye and the upper and lower eyelid areas, one or more AUs related to eyebrow movement can be selected, such as: AU1 (InnerBrowRaiser, used to represent the eyebrow raising), AU2 (OuterBrowRaiser, used to represent the eyebrow tail raising), AU4 (BrowLowerer, used to represent the eyebrow lowering), AU5 (UpperLidRaiser, used to represent the upper eyelid raising), AU6 (CheekRaiser, used to represent the slight squinting when the cheek is raised), AU44 (Squint, used to represent squinting), etc. This allows the subsequent eyebrow animation curve corresponding to the eyebrow movement state to be obtained based on the selected driving units of the eyebrow animation and the weights of each driving unit under different expression intensities. Thus, eyebrow animation can be generated based on the eyebrow animation curve.
[0144] In step S105, the AU of the eyebrows associated with the emotional state corresponding to each target time period is interpolated to obtain the eyebrow animation curve of each segment corresponding to the emotional state.
[0145] In this embodiment, after the AU of the eyebrows associated with the emotional state corresponding to each target time period, interpolation calculation can be performed on the AU of the eyebrows associated with the emotional state corresponding to each target time period to obtain the segment eyebrow animation curve corresponding to each emotional state.
[0146] Specifically, after the AU of the eyebrows associated with the emotional state corresponding to each target time period, interpolation calculations can be performed on the AU of the eyebrows associated with the emotional state corresponding to each target time period. Specifically, piecewise cubic Hermite interpolation algorithm or interpolation algorithm can be used, or other interpolation algorithms, such as cubic spline interpolation algorithm, can be used. No specific restrictions are imposed here, in order to obtain the segment eyebrow animation curve corresponding to each emotional state.
[0147] In step S106, the eyebrow animation curves of each segment are spliced together to obtain the target eyebrow animation curve;
[0148] In this embodiment, after obtaining the eyebrow animation curve of each segment, the eyebrow animation curves of each segment can be spliced together to obtain the eyebrow animation curve corresponding to the entire audio to be processed, i.e., the target eyebrow animation curve.
[0149] Specifically, such as Figure 12As shown, after obtaining the eyebrow animation curve for each segment, it is possible to first check if there are any segments for which the corresponding eyebrow animation curve has not yet been generated (e.g., ...). Figure 12 (As indicated, are there any remaining segments?) If so, the remaining segments can continue to be processed to generate eyebrow animation curves until there are no remaining segments. Then, the eyebrow animation curves of each segment can be spliced together (e.g., ...). Figure 12 The desired segment fusion can be achieved by using the concat function for splicing, or by splicing the eyebrow animation curves of each segment according to the order of the target time period corresponding to each segment's eyebrow animation curve. Then, for every two adjacent eyebrow animation curves after splicing, the segment curve corresponding to the splicing point is obtained according to the basic window length, and the segment curve is Gaussian smoothed to obtain the target eyebrow animation curve. No specific restrictions are imposed here.
[0150] In step S107, based on the target eyebrow animation curve, an eyebrow animation corresponding to the audio to be processed is generated.
[0151] Specifically, after obtaining the target eyebrow animation curve, it can be edited and rendered based on the target eyebrow animation curve to generate a simulated eyebrow animation that can be used in a facial expression driving system, that is, the eyebrow animation corresponding to the audio to be processed. It can be understood that the target eyebrow animation curve can also be directly superimposed on the overall expression driver to generate the facial expression of the virtual object.
[0152] In this embodiment, a method for generating eyebrow animation is provided. This method fragments the audio to be processed, dividing it into audio segments with sound, silent segments, expressive segments, and expressionless segments. The intersection of these segments is then used to obtain the emotional state corresponding to each target time period. Based on expression-driven AU rules, different emotional states are further fragmented to more accurately obtain the corresponding eyebrow animation curve for each emotional state. These fragmented eyebrow animation curves are then concatenated to form the target eyebrow animation curve. Based on this target curve, the eyebrow animation corresponding to the audio to be processed is generated. This method not only eliminates the influence of ambiguous environmental factors on the overall facial expression animation and simplifies the automatic generation process of eyebrow animation, but also considers the natural movement of expressionless or silent eyebrows based on expression-driven AU rules. This avoids stiff facial expressions when expressionless or silent, ensuring that the eyebrow animation fits the scene setting of the audio to be processed, thereby enhancing the realism of the virtual object's facial expression-driven animation.
[0153] Optionally, in the above Figure 2 Based on the corresponding embodiments, in another optional embodiment of the eyebrow animation generation method provided in this application, such as... Figure 3 As shown, when the emotional state is expressionless and silent; step S104 determines the AU of the eyebrows associated with the emotional state for each target time period based on the expression-driven AU rule, including: steps S301 to S303; and step S105 includes: step S304.
[0154] In step S301, multiple first random seed points on the segment corresponding to the expressionless and silent state are calculated using a random function;
[0155] In step S302, for each of the multiple first random seed points, the extended segment corresponding to each first random seed point is extracted from the segment corresponding to the expressionless and silent state;
[0156] In step S303, based on the expression-driven AU rule, the default AU of the eyebrow movement state associated with the expressionless state is selected as the AU of the eyebrow.
[0157] In step S304, interpolation calculations are performed based on each first random seed point, the extended segment corresponding to each first random seed point, and each default AU to obtain the eyebrow animation curve corresponding to the expressionless and silent state.
[0158] In this embodiment, when the emotional state corresponding to the current target time period is a blank and silent state, multiple first random seed points on the segment corresponding to the blank and silent state can be calculated by a random function. Then, for each of the multiple first random seed points, an extended segment corresponding to each first random seed point is extracted on the segment corresponding to the blank and silent state. Then, based on the expression-driven AU rule, the default AU of the eyebrow movement state associated with the blank state is selected as the AU of the eyebrow. Interpolation calculation is performed based on each first random seed point, the extended segment corresponding to each first random seed point, and each default AU to obtain the eyebrow animation curve of the segment corresponding to the blank and silent state. This can take into account the eyebrow movement state of the virtual object in the blank state. Combined with the expression-driven AU rule, several eyebrow animations in the relaxed state of blank expression can be randomly generated to avoid the stiffness of the facial expression state of the virtual object in the blank state.
[0159] Understandably, random numbers are obtained through complex mathematical algorithms, and the random seed is the initial value of these random numbers. Generally, random numbers generated by computers are pseudo-random numbers. A pseudo-random number is a number that remains unchanged. The first random seed is the initial value of a random number calculated based on a random function such as the Random function, within a segment corresponding to a state of expressionlessness and silence.
[0160] The extended segment refers to the segment obtained by extending a random length outward from each first random seed point. Specifically, the extended segment can be, for example, a segment from the start timestamp to the first random seed point and a segment from the first random seed point back to the start timestamp, or a segment obtained by extending a random length outward from both ends with the first random seed point as the center, or it can be in other forms, without specific restrictions here.
[0161] Specifically, when the emotional state is a state of being expressionless and silent (such as...) Figure 12 In the state 1) shown, multiple first random seed points are randomly generated on the segment corresponding to the expressionless and silent state using a random function. Then, the segment corresponding to each first random seed point is expanded outward by a random length centered on each first random seed point.
[0162] Furthermore, based on the expression-driven AU rules, default AUs associated with eyebrow movement states in a non-expression state can be selected. Specifically, this can be based on one or more of the AUs related to eyebrows in a predefined natural state, such as AU1, AU2, or AU4. For example, a relaxed state can be represented by AU1+AU2 (e.g., relaxed eyebrows), or another relaxed state can be represented by AU4 (e.g., lowered eyebrows). Small variations can also be added to increase the diversity of eyebrow movement states. Simultaneously, the coefficients corresponding to each default AU under a preset expression intensity are obtained based on the predefined natural state; these are the default AU weights.
[0163] Furthermore, when there are no art resources, a segmented cubic Hermite interpolation algorithm can be used as needed to calculate the animation curves corresponding to one or more of the default AUs (AU1, AU2, or AU4) under the preset random length, i.e., the extended segment. The corresponding animation curves are then weighted based on the weight of each default AU to obtain the eyebrow animation curves for the segment in a silent and expressionless state.
[0164] Optionally, in the above Figure 2 Based on the corresponding embodiments, in another optional embodiment of the eyebrow animation generation method provided in this application, such as... Figure 4 As shown, when the emotional state is an expressionless state; step S104 determines the AU of the eyebrows associated with the emotional state for each target time period based on the expression-driven AU rule, including: steps S401 to S403; and step S105 includes: steps S404 to S405.
[0165] In step S401, multiple second random seed points on the segment corresponding to the expressionless but silent state are calculated using a random function;
[0166] In step S402, for each of the multiple second random seed points, the extended segment corresponding to each second random seed point is extracted from the segment corresponding to the expressionless and silent state;
[0167] In step S403, based on the expression-driven AU rule, the first expression AU of the eyebrow movement state associated with the expression type in the expressionless state is selected as the AU of the eyebrow.
[0168] In step S404, the weight of the first AU corresponding to each first expression AU is obtained through fuzzy rules;
[0169] In step S405, interpolation calculation is performed based on each second random seed point, the extended segment corresponding to each second random seed point, and each first expression AU. The interpolation calculation results are then weighted based on the weight of each first AU to obtain the eyebrow animation curve corresponding to the expressionless state.
[0170] In this embodiment, when the emotional state corresponding to the current target time period is an expressionless state, multiple second random seed points on the segment corresponding to the expressionless state can be calculated using a random function. Then, for each of the multiple second random seed points, an extended segment corresponding to each second random seed point is extracted on the segment corresponding to the expressionless state. Based on the expression-driven AU rule, the first expression AU of the eyebrow movement state associated with the expression type in the expressionless state is selected as the eyebrow AU. Then, through fuzzy rules, the first AU weight corresponding to each first expression AU is obtained. Interpolation calculation is performed based on each second random seed point, the extended segment corresponding to each second random seed point, and each first expression AU. The interpolation calculation result is weighted based on each first AU weight to obtain the eyebrow animation curve of the segment corresponding to the expressionless state. This can take into account the eyebrow movement state of the virtual object in the silent state. Combined with the expression-driven AU rule, several eyebrow animations in the silent and relaxed state can be randomly generated to avoid the stiffness of the virtual object's facial expression state in the silent state.
[0171] The second random seed point is the initial value of a random number calculated based on a random function such as the Random function on the segment corresponding to the expressionless state. The expanded segment refers to the segment obtained by expanding outwards by a random length centered on each second random seed point. The first expression AU is the AU driving unit associated with the eyebrow movement state under the expression type and corresponding expression intensity in the expressionless state. The fuzzy rule is a binary fuzzy relation R defined in X×Y, which can be understood as a relationship where if X is A, then Y is B, i.e., A→B.
[0172] Specifically, when the emotional state is an expressionless state (such as...) Figure 12 When the desired state 2) is obtained, multiple second random seed points are randomly generated on the segment corresponding to the expressionless and silent state using a random function. Then, the segment corresponding to each second random seed point is expanded outward by a random length with each second random seed point as the center.
[0173] Furthermore, based on the expression-driven AU rules, the first expression AU associated with the eyebrow movement state in the expression-driven silent state can be selected. Specifically, based on the expression-driven AU rules, the AU driving unit associated with the eyebrow movement state under the expression type and the expression intensity corresponding to that expression type can be selected. At the same time, the coefficient corresponding to each first expression AU under the expression intensity corresponding to the expression type can be searched in the expression-driven AU rules shown in Table 2 through fuzzy rules, that is, the first AU weight.
[0174] Furthermore, a segmented cubic Hermite interpolation algorithm can be used as needed to calculate the animation curve corresponding to each first expression AU under a preset random length, i.e., an extended segment. The animation curve corresponding to each first expression AU is then weighted based on the weight of each first AU to obtain the eyebrow animation curve corresponding to the expressionless state.
[0175] Optionally, in the above Figure 2 Based on the corresponding embodiments, in another optional embodiment of the eyebrow animation generation method provided in this application, such as... Figure 5 As shown, when the emotional state is a blank expression with vocalization; step S104, based on the expression-driven AU rule, determines the AU of the eyebrows associated with the emotional state for each target time period, including:
[0176] In step S501, multiple third random seed points on the segment corresponding to the expressionless but vocal state are calculated using a random function;
[0177] In step S502, for each of the multiple third random seed points, the extended segment corresponding to each third random seed point is extracted from the segment corresponding to the expressionless but vocal state;
[0178] In step S503, based on the expression-driven AU rule, the default AU of the eyebrow movement state associated with the expressionless state is selected as the AU of the eyebrow.
[0179] In step S504, the sound signal in the segment corresponding to the sound state is extracted;
[0180] In step S505, the sound signal is smoothed to obtain the smoothed value corresponding to the sound signal, and the smoothed value is used as the AU adjustment weight.
[0181] In step S506, interpolation calculation is performed based on each third random seed point, the extended segment corresponding to each third random seed point, and each default AU. The interpolation calculation results are then weighted based on the AU adjustment weights to obtain the eyebrow animation curve corresponding to the expressionless and vocal state.
[0182] In this embodiment, when the emotional state corresponding to the current target time period is a blank expression with sound, multiple third random seed points on the segment corresponding to the blank expression with sound can be calculated using a random function. Then, for each of the multiple third random seed points, an extended segment corresponding to each third random seed point is extracted on the segment corresponding to the blank expression with sound. Based on the expression-driven AU rule, a default AU for the eyebrow movement state associated with the blank expression state is selected. At the same time, the sound signal in the segment corresponding to the sound state is extracted, and the sound signal is smoothed to obtain the smooth value corresponding to the sound signal. The smooth value is used as the AU adjustment weight. Then, interpolation calculation is performed based on each third random seed point, the extended segment corresponding to each third random seed point, and each default AU. The interpolation calculation result is weighted based on the AU adjustment weight to obtain the eyebrow animation curve of the segment corresponding to the blank expression with sound state. This can take into account the influence of audio fluctuations in the sound signal on the eyebrow movement state when speaking in a relaxed state. Furthermore, high-frequency small jitters in the sound signal can be filtered out by smoothing the sound signal, such as by low-pass filtering, to avoid the situation where the eyebrow jitters are caused by relying entirely on the sound signal.
[0183] The third random seed point is the initial value of a random number calculated based on a random function such as the Random function on the segment corresponding to the expressionless yet vocal state. The extended segment refers to a segment obtained by expanding outwards by a random length centered on each third random seed point. The sound signal can specifically represent a sound intensity signal or a sound pitch signal, or other sound signals; no specific restrictions are imposed here.
[0184] Specifically, when the emotional state is a state of being expressionless but vocal (such as...) Figure 12 In the case of state 3), in order to maintain the diversity of the algorithm (for example, different segments under the same audio represent the same emotion and the audio signal is consistent), this embodiment can add a small number of random seed points to obtain the corresponding extended segments to superimpose slight adjustments, so that different effects will be reflected in the specific eyebrow animation, and the effect will also have some random slight changes. Specifically, multiple third random seed points can be randomly generated on the segment corresponding to the expressionless but vocal state through a random function, and then the extended segment corresponding to each third random seed point is obtained by expanding outward by a random length with each third random seed point as the center.
[0185] Furthermore, based on the expression-driven AU rule, the default AU of the eyebrow movement state associated with the expressionless state can be selected. Specifically, it can be one or more of the AUs associated with the eyebrows under the expression-driven AU rule in the predefined natural state, such as AU1, AU2, or AU4. At the same time, the coefficient corresponding to each default AU under the preset expression intensity can be obtained according to the predefined natural state, that is, the default AU weight.
[0186] Furthermore, the changes in the eyebrow animation curve are mainly adjusted by extracting the sound signal from the segment corresponding to the voiced state. Specifically, the sound signal can be obtained using the Praat tool or other audio editors, without specific limitations. The sound signal is obtained from the segment corresponding to the expressionless voiced state. Then, the sound signal is smoothed to obtain the corresponding smooth value. Specifically, a low-pass filter can be used to obtain a continuous smooth distribution of the sound signal to reduce the influence of high-frequency jitter on the sound signal. Then, the continuous smooth distribution of the sound signal is normalized to obtain the normalized value, which is the smooth value. The smooth value is used as the AU adjustment weight.
[0187] Furthermore, a segmented cubic Hermite interpolation algorithm can be used as needed to calculate the animation curves corresponding to one or more of the default AUs (AU1, AU2, or AU4) under the preset random length, i.e., the extended segment. The corresponding animation curves are then weighted based on the weight of each default AU and adjusted according to the AU adjustment weights to obtain the eyebrow animation curves corresponding to the expressionless but vocal state.
[0188] Optionally, in the above Figure 2 Based on the corresponding embodiments, in another optional embodiment of the eyebrow animation generation method provided in this application, such as... Figure 6 As shown, when the emotional state is one of expression and speech; step S104, based on the expression-driven AU rule, determines the AU of the eyebrows associated with the emotional state for each target time period, including:
[0189] In step S601, multiple fourth random seed points on the segment corresponding to the expressive and vocal state are calculated using a random function;
[0190] In step S602, for each of the multiple fourth random seed points, the extended segment corresponding to each fourth random seed point is extracted from the segment corresponding to the expressive and vocal state;
[0191] In step S603, based on the expression-driven AU rule, the second expression AU of the eyebrow movement state associated with the expression type in the expression and voice state is selected as the AU of the eyebrow.
[0192] In step S604, the sound signal in the segment corresponding to the sound state is extracted;
[0193] In step S605, the sound signal is smoothed to obtain the smoothed value corresponding to the sound signal, and the smoothed value is used as the AU adjustment weight.
[0194] In step S606, interpolation calculation is performed based on each fourth random seed point, the extended segment corresponding to each fourth random seed point, and each second expression AU. The interpolation calculation results are then weighted based on the AU adjustment weight to obtain the eyebrow animation curve of the segment corresponding to the expressive and vocal state.
[0195] In this embodiment, when the emotional state corresponding to the current target time period is an expressive and vocal state, multiple fourth random seed points can be calculated on the segment corresponding to the expressive and vocal state using a random function. Then, for each of the multiple fourth random seed points, an extended segment corresponding to each fourth random seed point is extracted on the segment corresponding to the expressive and vocal state. Based on the expression-driven AU rule, a second expression AU for the eyebrow movement state associated with the expression type in the expressive and vocal state is selected. Simultaneously, the sound signal in the segment corresponding to the vocal state is extracted and smoothed to obtain the sound... The signal is smoothed and used as the AU adjustment weight. Then, interpolation is performed based on each fourth random seed point, the corresponding extended segment of each fourth random seed point, and each second expression AU. The interpolation results are weighted based on the AU adjustment weight to obtain the eyebrow animation curve corresponding to the expressive and vocal state. This method can consider the influence of audio fluctuations in the sound signal on the expression when speaking in different emotional states. High-frequency small jitters in the sound signal can be filtered out by smoothing the sound signal, such as by low-pass filtering, to avoid the situation where the eyebrows jitter high-frequency due to complete reliance on the sound signal.
[0196] The fourth random seed point is the initial value of a random number calculated based on a random function such as the Random function on the segment corresponding to the expressionless but vocal state. The expanded segment refers to a segment obtained by expanding outwards by a random length centered on each fourth random seed point. The second expression AU is an AU driving unit associated with the eyebrow movement state under the expression type and corresponding expression intensity in the expressionless and vocal state.
[0197] Specifically, when the emotional state is one of expression and vocalization (such as...) Figure 12In the case of state 4), in order to maintain the diversity of the algorithm (for example, different segments under the same audio represent the same emotion and the audio signal is consistent), this embodiment can add a small number of random seed points to obtain the corresponding extended segments to superimpose slight adjustments, so that different effects will be reflected in the specific eyebrow animation, and the effect will also have some random slight changes. Specifically, multiple fourth random seed points can be randomly generated on the segments corresponding to the expressive and vocal states through a random function, and then the extended segments corresponding to each fourth random seed point are obtained by expanding outwards by a random length with each fourth random seed point as the center.
[0198] Furthermore, based on the expression-driven AU rule, a second expression AU related to the eyebrow movement state in the expression-and-voice state can be selected. Specifically, based on the expression-driven AU rule, the AU driving unit associated with the eyebrow movement state under the expression type and the corresponding expression intensity can be selected. At the same time, through fuzzy rules, the coefficient corresponding to the expression intensity of each second expression AU under the expression type can be searched in the expression-driven AU rule shown in Table 2, that is, the second AU weight.
[0199] Furthermore, the changes in the eyebrow animation curve are mainly adjusted by extracting the sound signal from the segment corresponding to the voiced state. Specifically, the sound signal can be obtained using the Praat tool or other audio editors, without any specific restrictions. The sound signal is then obtained from the segment corresponding to the expressive and voiced state. Then, the sound signal is smoothed to obtain the corresponding smooth value. Specifically, a low-pass filter can be used to obtain a continuous smooth distribution of the sound signal to reduce the impact of high-frequency jitter on the sound signal. Then, the continuous smooth distribution of the sound signal is normalized to obtain the normalized value, which is the smooth value. The smooth value is used as the AU adjustment weight.
[0200] Furthermore, a segmented cubic Hermite interpolation algorithm can be used as needed to calculate the animation curve corresponding to each second expression AU under a preset random length, i.e., an extended segment. The animation curve corresponding to each second expression AU is then weighted based on the weight of each second AU, and adjusted according to the AU adjustment weight to obtain the eyebrow animation curve of the segment corresponding to the expressive and vocal state.
[0201] Optionally, in the above Figure 5 or Figure 6 Based on the corresponding embodiments, in another optional embodiment of the eyebrow animation generation method provided in this application, such as... Figure 7As shown, when the sound signal is a sound intensity signal, step S504 or step S604 extracts the sound signal from the segment corresponding to the sound state, including step S701; step S505 or step S605 includes steps S702 to S703.
[0202] In step S701, the sound intensity signal is extracted from the segment corresponding to the sound state;
[0203] In step S702, the sound intensity signal is filtered by a low-pass filter to obtain a smooth distribution of the sound intensity signal;
[0204] In step S703, the smooth distribution of the sound intensity signal is normalized to obtain the smooth value corresponding to the sound intensity signal, and the smooth value is used as the AU adjustment weight.
[0205] In this embodiment, when the sound signal is a sound intensity signal, the sound intensity signal in the segment corresponding to the sound state can be extracted, and the sound intensity signal can be filtered by a low-pass filter to obtain a smooth distribution of the sound intensity signal. Then, the smooth distribution of the sound intensity signal is normalized to obtain the smooth value corresponding to the sound intensity signal, and the smooth value is used as the AU adjustment weight. The low-pass filter can be used to filter the sound intensity signal to reduce the influence of high-frequency jitter on the sound intensity signal, thereby avoiding the high-frequency jitter of the eyebrows caused by relying entirely on the sound intensity signal, which would affect the presentation of the eyebrow animation.
[0206] In this context, sound intensity refers to the strength of a sound wave. Sound intensity, also known as volume, refers to the output or power of sound energy. For example, in a piece of music, some sections need to be played softly, while others need to be played loudly; this refers to sound intensity. Understandably, excessive sound intensity makes the sound feel very loud.
[0207] Specifically, when the sound signal is a sound intensity signal, it is possible to extract the expressionless vocal state (e.g., Figure 12 The sound signal in the segment corresponding to state 3) shown can be obtained using the Praat tool or other audio editors to acquire the intensity signal; no specific restrictions are imposed here, in order to obtain a state of expressionless speech (such as...). Figure 12 The sound intensity signal in the segment corresponding to state 3) is then smoothed to obtain the smoothed value corresponding to the sound intensity signal. Specifically, a low-pass filter can be used to obtain a continuous smooth distribution of the sound intensity signal to reduce the influence of high-frequency jitter on the sound intensity signal. Then, the continuous smooth distribution of the sound intensity signal is normalized to obtain the normalized value, i.e., the smoothed value, and the smoothed value is used as the AU adjustment weight.
[0208] Furthermore, a segmented cubic Hermite interpolation algorithm can be used as needed to calculate the animation curves corresponding to one or more of the default AUs (AU1, AU2, or AU4) under the preset random length, i.e., the extended segment. The corresponding animation curves are then weighted based on the weight of each default AU and adjusted according to the AU adjustment weights to obtain the eyebrow animation curves corresponding to the expressionless but vocal state.
[0209] Similarly, when the sound signal is a sound intensity signal, expressive and vocal states (such as...) can be extracted. Figure 12 The sound signal in the segment corresponding to state 4) can be obtained using the Praat tool or other audio editors to acquire the intensity signal; no specific restrictions are imposed here, in order to obtain an expressive and vocal state (such as...). Figure 12 The sound intensity signal in the segment corresponding to state 4) is then smoothed to obtain the smoothed value corresponding to the sound intensity signal. Specifically, a low-pass filter can be used to obtain a continuous smooth distribution of the sound intensity signal to reduce the influence of high-frequency jitter on the sound intensity signal. Then, the continuous smooth distribution of the sound intensity signal is normalized to obtain the normalized value, i.e., the smoothed value, and the smoothed value is used as the AU adjustment weight.
[0210] Furthermore, a segmented cubic Hermite interpolation algorithm can be used as needed to calculate the animation curve corresponding to each second expression AU under a preset random length, i.e., an extended segment. The animation curve corresponding to each second expression AU is then weighted based on the weight of each second AU, and adjusted according to the AU adjustment weight to obtain the eyebrow animation curve of the segment corresponding to the expressive and vocal state.
[0211] Optionally, in the above Figure 5 or Figure 6 Based on the corresponding embodiments, in another optional embodiment of the eyebrow animation generation method provided in this application, such as... Figure 8 As shown, when the sound signal is a pitch signal, step S504 or step S604 extracts the sound signal from the segment corresponding to the sound state, including step S801; step S505 or step S605 includes steps S802 to S803.
[0212] In step S801, the pitch signal is extracted from the segment corresponding to the expressive and vocal state;
[0213] In step S802, the pitch signal is filtered by a low-pass filter to obtain a smooth distribution of the pitch signal;
[0214] In step S803, the smooth distribution of the pitch signal is normalized to obtain the smooth value corresponding to the pitch signal, and the smooth value is used as the AU adjustment weight.
[0215] In this embodiment, when the sound signal is a pitch signal, the pitch signal in the segment corresponding to the sound state can be extracted, and the pitch signal can be filtered by a low-pass filter to obtain a smooth distribution of the pitch signal. Then, the smooth distribution of the pitch signal is normalized to obtain the smooth value corresponding to the pitch signal, and the smooth value is used as the AU adjustment weight. The low-pass filter can be used to filter the pitch signal to reduce the influence of high-frequency jitter on the pitch signal, thereby avoiding the high-frequency jitter of the eyebrows caused by relying entirely on the pitch signal, which would affect the presentation of the eyebrow animation.
[0216] Among them, pitch signal refers to the frequency of sound wave. Pitch is the height of sound. The higher the pitch, the higher the frequency. It can be understood that excessively high pitch will make people feel harsh.
[0217] Specifically, when the sound signal is a pitch signal, the expressionless vocal state (e.g.) can be extracted. Figure 12 The audio signal corresponding to state 3) can be obtained using the Praat tool or other audio editors to capture pitch signals; no specific restrictions are imposed here. This is to obtain a voiceless, expressionless state (such as...). Figure 12 The pitch signal in the segment corresponding to state 3) is then smoothed to obtain the smoothed value corresponding to the pitch signal. Specifically, a low-pass filter can be used to obtain a continuous smooth distribution of the pitch signal to reduce the influence of high-frequency jitter on the pitch signal. Then, the continuous smooth distribution of the pitch signal is normalized to obtain the normalized value, i.e., the smoothed value, and the smoothed value is used as the AU adjustment weight.
[0218] Furthermore, a segmented cubic Hermite interpolation algorithm can be used as needed to calculate the animation curves corresponding to one or more of the default AUs (AU1, AU2, or AU4) under the preset random length, i.e., the extended segment. The corresponding animation curves are then weighted based on the weight of each default AU and adjusted according to the AU adjustment weights to obtain the eyebrow animation curves corresponding to the expressionless but vocal state.
[0219] Similarly, when the sound signal is a pitch signal, expressive and vocal states (such as...) can be extracted. Figure 12 The audio signal corresponding to state 4) can be obtained using the Praat tool or other audio editors to capture pitch signals; no specific restrictions are imposed here. The goal is to obtain an expressive and vocalized state (such as...). Figure 12The pitch signal in the segment corresponding to state 4) is then smoothed to obtain the smoothed value corresponding to the pitch signal. Specifically, a low-pass filter can be used to obtain a continuous smooth distribution of the pitch signal to reduce the influence of high-frequency jitter on the pitch signal. Then, the continuous smooth distribution of the pitch signal is normalized to obtain the normalized value, i.e., the smoothed value, and the smoothed value is used as the AU adjustment weight.
[0220] Furthermore, a segmented cubic Hermite interpolation algorithm can be used as needed to calculate the animation curve corresponding to each second expression AU under a preset random length, i.e., an extended segment. The animation curve corresponding to each second expression AU is then weighted based on the weight of each second AU, and adjusted according to the AU adjustment weight to obtain the eyebrow animation curve of the segment corresponding to the expressive and vocal state.
[0221] Optionally, in the above Figure 5 or Figure 6 Based on the corresponding embodiments, in another optional embodiment of the eyebrow animation generation method provided in this application, such as... Figure 9 As shown, when the sound signal is a sound intensity signal and a sound pitch signal, step S504 or step S604 extracts the sound signal from the segment corresponding to the sound state, including step S901; step S505 or step S605 includes steps S902 to S904.
[0222] In step S901, the intensity signal and pitch signal of the segment corresponding to the sound state are extracted;
[0223] In step S902, the intensity signal and the pitch signal are filtered by a low-pass filter to obtain a smooth distribution of the intensity signal and a smooth distribution of the pitch signal.
[0224] In step S903, the smooth distribution of the sound intensity signal is normalized to obtain the smooth value corresponding to the sound intensity signal, and the smooth value is used as the first AU adjustment weight.
[0225] In step S904, the smooth distribution of the pitch signal is normalized to obtain the smooth value corresponding to the pitch signal, and the smooth value is used as the second AU adjustment weight.
[0226] In this embodiment, when processing the sound signal, the distribution of pitch and intensity can be comprehensively considered to further refine the correlation between audio features and AU. Especially when accent recognition is inaccurate, the numerical distribution of pitch and intensity can be used to more accurately confirm the correlation with AU. Therefore, the intensity and pitch signals in the segment corresponding to the sound state can be extracted, and the intensity and pitch signals can be filtered by a low-pass filter to obtain the smooth distribution of the intensity signal and the smooth distribution of the pitch signal. Then, the smooth distribution of the intensity signal is normalized to obtain the smooth value corresponding to the intensity signal, and the smooth value is used as the first AU adjustment weight. Similarly, the smooth distribution of the pitch signal is normalized to obtain the smooth value corresponding to the pitch signal, and the smooth value is used as the second AU adjustment weight, so that the AU of the eyebrow can be better obtained based on the first AU adjustment weight and the second AU adjustment weight.
[0227] Specifically, when the sound signal consists of intensity and pitch signals, it is possible to extract the expressionless vocal state (e.g., Figure 12 The sound signal in the segment corresponding to state 3) can be obtained using the Praat tool or other audio editors to acquire intensity and pitch signals. No specific restrictions are imposed here, in order to obtain a voiceless, expressionless state (e.g., Figure 12 The intensity and pitch signals in the segment corresponding to state 3) are then smoothed to obtain smoothed values for the intensity and pitch signals, respectively. Specifically, a low-pass filter can be used to obtain a continuous smooth distribution of the intensity and pitch signals to reduce the impact of high-frequency jitter on the intensity and pitch signals. Then, the continuous smooth distribution of the intensity signal is normalized to obtain a normalized value, i.e., a smoothed value, which is used as the first AU adjustment weight. The continuous smooth distribution of the pitch signal is also normalized to obtain a normalized value, i.e., a smoothed value, which is used as the second AU adjustment weight.
[0228] Furthermore, a segmented cubic Hermite interpolation algorithm can be used as needed to calculate the animation curves corresponding to one or more of the default AUs (AU1, AU2, or AU4) under the preset random length, i.e., the extended segment. The corresponding animation curves are then weighted based on the weight of each default AU, and adjusted according to the first AU adjustment weight and the second AU adjustment weight to obtain the eyebrow animation curve of the segment corresponding to the expressionless but vocal state.
[0229] Similarly, when the sound signal consists of intensity and pitch signals, expressive and vocal states can be extracted (e.g., Figure 12The sound signal in the segment corresponding to state 4) can be obtained using the Praat tool or other audio editors to acquire intensity and pitch signals. No specific restrictions are imposed here, in order to obtain an expressive and vocal state (such as...). Figure 12 The intensity and pitch signals in the segment corresponding to state 4) are then smoothed to obtain smoothed values for the intensity and pitch signals, respectively. Specifically, a low-pass filter can be used to obtain a continuous smooth distribution of the intensity and pitch signals to reduce the impact of high-frequency jitter on the intensity and pitch signals. Then, the continuous smooth distribution of the intensity signal is normalized to obtain a normalized value, i.e., a smoothed value, which is used as the first AU adjustment weight. Similarly, the continuous smooth distribution of the pitch signal is normalized to obtain a normalized value, i.e., a smoothed value, which is used as the second AU adjustment weight.
[0230] Furthermore, a segmented cubic Hermite interpolation algorithm can be used as needed to calculate the animation curve corresponding to each second expression AU under a preset random length, i.e., an extended segment. The animation curve corresponding to each second expression AU is weighted based on the weight of each second AU, and then adjusted according to the adjustment weights of the first AU and the second AU to obtain the eyebrow animation curve of the segment corresponding to the expressive and vocal state.
[0231] Optionally, in the above Figure 7 Based on the corresponding embodiments, in another optional embodiment of the eyebrow animation generation method provided in this application, such as... Figure 10 As shown, step S703 normalizes the smooth distribution of the sound intensity signal, and after using the smoothed value corresponding to the sound intensity signal as the AU adjustment weight, the method further includes:
[0232] In step S1001, the accented segments and the intensity signals in the accented segments are extracted from the segments corresponding to the sound state.
[0233] In step S1002, the ratio of the intensity signal in the accented segment to the intensity threshold is used as the enhancement coefficient;
[0234] In step S1003, the AU adjustment weight is adjusted based on the enhancement coefficient to obtain the target AU adjustment weight.
[0235] In this embodiment, based on the extracted intensity signal, it is possible to further search whether the audio segment contains an accented segment. The intensity signal in the accented segment can be used to further refine the correlation between audio features and AU. Therefore, the accented segment and the intensity signal in the accented segment can be extracted from the segment corresponding to the audio state. Then, the ratio of the intensity signal in the accented segment to the intensity threshold is used as the enhancement coefficient. The AU adjustment weight is adjusted based on the enhancement coefficient to obtain the target AU adjustment weight, so that the AU of the eyebrow can be better obtained based on the target AU adjustment weight.
[0236] Understandably, in phonetics, stress is the phenomenon of a particular syllable being emphasized in a sequence of connected syllables. Stress can be expressed through force (increased intensity) or through pitch (changes in pitch). Therefore, a stressed segment refers to a segment of spoken text where a particular syllable is emphasized in a sequence of connected syllables. The intensity threshold is set according to practical application requirements, and a typical reference value is 60 dB (the volume of normal conversational sound).
[0237] Specifically, the process involves extracting accented segments and intensity signals from the corresponding audio segments. This can be done using speech recognition tools, such as MFA, or other tools; no specific limitations are imposed here. The process involves aligning the phonemes of the audio to be processed with those of the script (e.g.,...). Figure 12 After aligning the phonemes (as indicated), one can obtain the expressionless vocal state (such as...). Figure 12 In the segment corresponding to state 3), the duration of the stressed phoneme is selected as the stressed segment. Then, the intensity signal in the stressed segment can be obtained using Praat or other audio editors. No specific restrictions are imposed here. The ratio of the intensity signal in the stressed segment to the intensity threshold is used as the enhancement coefficient. The AU adjustment weight is adjusted based on the enhancement coefficient to obtain the target AU adjustment weight.
[0238] Furthermore, a segmented cubic Hermite interpolation algorithm can be used as needed to calculate the animation curves corresponding to one or more of the default AUs (AU1, AU2, or AU4) under the preset random length, i.e., the extended segment. The corresponding animation curves are then weighted based on the weight of each default AU, and the weights are adjusted according to the target AU to obtain the eyebrow animation curves corresponding to the expressionless but vocal state.
[0239] Similarly, extract the accented segments and intensity signals from the segments corresponding to the spoken state. This can be done using speech recognition tools, such as MFA, or other tools; no specific limitations are made here. This is achieved by aligning the phonemes of the audio to be processed with those of the script (e.g., ...). Figure 12After aligning the phonemes (as indicated), you can extract the expressive and vocal states (such as...). Figure 12 In the segment corresponding to state 4), the duration of the stressed phoneme is selected as the stressed segment. Then, the intensity signal in the stressed segment can be obtained using Praat or other audio editors. No specific restrictions are imposed here. The ratio of the intensity signal in the stressed segment to the intensity threshold is used as the enhancement coefficient. The AU adjustment weight is adjusted based on the enhancement coefficient to obtain the target AU adjustment weight.
[0240] Furthermore, a segmented cubic Hermite interpolation algorithm can be used as needed to calculate the animation curve corresponding to each second expression AU under a preset random length, i.e., an extended segment. The animation curve corresponding to each second expression AU is then weighted based on the weight of each second AU, and the weight is adjusted for the target AU to obtain the eyebrow animation curve corresponding to the segment with expression and sound.
[0241] Optionally, in the above Figure 2 Based on the corresponding embodiments, in another optional embodiment of the eyebrow animation generation method provided in this application, such as... Figure 11 As shown, step S106 splices the eyebrow animation curves of each segment to obtain the target eyebrow animation curve, including:
[0242] In step S1101, the eyebrow animation curves of each segment are spliced together according to the order of the target time periods corresponding to each segment of the eyebrow animation curve to obtain the spliced eyebrow animation curve.
[0243] In step S1102, for each pair of adjacent eyebrow animation curve segments in the spliced eyebrow animation curve, the corresponding segment curve at the splicing point is obtained according to the basic window length, and the segment curve is Gaussian smoothed to obtain the target eyebrow animation curve.
[0244] In this embodiment, after obtaining the eyebrow animation curve of each segment, the eyebrow animation curve of each segment can be spliced together according to the chronological order of the target time period corresponding to each segment to obtain the spliced eyebrow animation curve. Then, for every two adjacent eyebrow animation curves in the spliced eyebrow animation curve, the segment curve corresponding to the splicing point can be obtained according to the basic window length, and the segment curve can be Gaussian smoothed to obtain the target eyebrow animation curve. This can handle the transition problem between the segment eyebrow animation curves, making the target eyebrow animation curve continuous and smooth, so that eyebrow animation can be generated better based on the target eyebrow animation curve.
[0245] Specifically, such as Figure 12As shown, after obtaining the eyebrow animation curve for each segment, it is possible to first check if there are any segments for which the corresponding eyebrow animation curve has not yet been generated (e.g., ...). Figure 12 (As indicated, are there any remaining segments?) If so, the remaining segments can continue to be processed to generate eyebrow animation curves until there are no remaining segments. Then, the eyebrow animation curves of each segment can be spliced together (e.g., ...). Figure 12 The segment fusion (as shown) can be achieved by splicing the eyebrow animation curves of each segment according to the order of the target time period corresponding to each segment. Then, for every two adjacent eyebrow animation curves after splicing, the segment curve corresponding to the splicing point is obtained according to the basic window length. Specifically, a window near the basic window length can be selected for the start and end timestamps corresponding to each eyebrow animation curve segment, and the selected segment curve in the window can be Gaussian smoothed to obtain the target eyebrow animation curve.
[0246] The eyebrow animation generation apparatus of this application is described in detail below. Please refer to [link / reference]. Figure 13 , Figure 13 This is a schematic diagram of one embodiment of the eyebrow animation generation device in this application. The eyebrow animation generation device 20 includes:
[0247] This application also provides an apparatus for generating eyebrow animation, comprising:
[0248] The processing unit 201 is used to divide the audio to be processed into audio segments and silent segments, wherein each audio segment corresponds to an audio start and end timestamp, and each silent segment corresponds to a silent start and end timestamp.
[0249] The processing unit 201 is further configured to divide the audio to be processed into expressive segments and non-expressive segments based on the expression tags corresponding to the audio to be processed. The expression tags include expression types and expression start and end timestamps corresponding to each expression type. Each expression segment corresponds to an expression type and an expression start and end timestamp, and a non-expressive segment corresponds to a non-expressive start and end timestamp.
[0250] The acquisition unit 202 is used to acquire the intersection of the audio segment, the silent segment, the facial expression segment, and the expressionless segment based on the audio start and end timestamp, the silent start and end timestamp, the facial expression start and end timestamp, and the expressionless start and end timestamp, so as to obtain the emotional state corresponding to the target time period. The emotional state includes the expressionless and silent state, the facial expression and silent state, the expressionless and audio state, or the facial expression and audio state.
[0251] The determination unit 203 is used to determine the AU of the eyebrows associated with the emotional state for each target time period based on the expression-driven AU rule;
[0252] The processing unit 201 is also used to perform interpolation calculation on the AU of the eyebrows associated with the emotional state corresponding to each target time period, so as to obtain the eyebrow animation curve of each segment corresponding to the emotional state.
[0253] The processing unit 201 is also used to splice the eyebrow animation curves of each segment to obtain the target eyebrow animation curve;
[0254] The processing unit 201 is also used to generate an eyebrow animation corresponding to the audio to be processed based on the target eyebrow animation curve.
[0255] Optionally, in the above Figure 13 Based on the corresponding embodiments, in another embodiment of the eyebrow animation generation apparatus provided in this application, the determining unit 203 can specifically be used for:
[0256] Multiple first random seed points on the segment corresponding to the expressionless and silent state are calculated using a random function;
[0257] For each of the multiple first random seed points, extract the extended segment corresponding to each first random seed point from the segment corresponding to the expressionless and silent state;
[0258] Based on the expression-driven AU rule, the default AU of the eyebrow movement state associated with the expressionless state is selected as the AU of the eyebrow.
[0259] The processing unit 201 can be used to perform interpolation calculations based on each first random seed point, the extended segment corresponding to each first random seed point, and each default AU to obtain the eyebrow animation curve corresponding to the expressionless and silent state.
[0260] Optionally, in the above Figure 13 Based on the corresponding embodiments, in another embodiment of the eyebrow animation generation apparatus provided in this application, the determining unit 203 can specifically be used for:
[0261] Multiple second random seed points on the segment corresponding to the expressionless but silent state are calculated using a random function;
[0262] For each of the multiple second random seed points, extract the extended segment corresponding to each second random seed point from the segment corresponding to the expressionless and silent state;
[0263] Based on the expression-driven AU rule, the first expression AU of the eyebrow movement state associated with the expression type in the expressionless state is selected as the AU of the eyebrow.
[0264] The processing unit 201 can be specifically used for:
[0265] By using fuzzy rules, the weight of the first AU corresponding to each first expression AU is obtained;
[0266] Interpolation calculations are performed based on each second random seed point, the corresponding extended segment of each second random seed point, and each first expression AU. The interpolation calculation results are then weighted based on the weight of each first AU to obtain the eyebrow animation curve corresponding to the expressionless state.
[0267] Optionally, in the above Figure 13 Based on the corresponding embodiments, in another embodiment of the eyebrow animation generation apparatus provided in this application, the determining unit 203 can specifically be used for:
[0268] Multiple third random seed points on the segment corresponding to the expressionless but vocal state are calculated using a random function;
[0269] For each of the multiple third random seed points, extract the extended segment corresponding to each third random seed point from the segment corresponding to the expressionless but vocal state;
[0270] Based on the expression-driven AU rule, the default AU of the eyebrow movement state associated with the expressionless state is selected as the AU of the eyebrow.
[0271] The processing unit 201 can be specifically used for:
[0272] Extract the sound signal from the segment corresponding to the audio state;
[0273] The sound signal is smoothed to obtain the smoothed value, and the smoothed value is used as the AU adjustment weight.
[0274] Interpolation calculations are performed based on each third random seed point, the corresponding extended segment of each third random seed point, and each default AU. The interpolation calculation results are then weighted based on the AU adjustment weights to obtain the eyebrow animation curve corresponding to the expressionless yet vocal state.
[0275] Optionally, in the above Figure 13 Based on the corresponding embodiments, in another embodiment of the eyebrow animation generation apparatus provided in this application, the determining unit 203 can specifically be used for:
[0276] Multiple fourth random seed points on the segment corresponding to the expressive and vocal state are calculated using a random function;
[0277] For each of the multiple fourth random seed points, extract the extended segment corresponding to each fourth random seed point from the segment corresponding to the expressive and vocal state;
[0278] Based on the expression-driven AU rule, the second expression AU of the eyebrow movement state associated with the expression type in the expression and voice state is selected as the AU of the eyebrow.
[0279] The processing unit 201 can be specifically used for:
[0280] Extract the sound signal from the segment corresponding to the audio state;
[0281] The sound signal is smoothed to obtain the smoothed value, and the smoothed value is used as the AU adjustment weight.
[0282] Interpolation calculations are performed based on each fourth random seed point, the corresponding extended segment of each fourth random seed point, and each second expression AU. The interpolation calculation results are then weighted based on the AU adjustment weights to obtain the eyebrow animation curve of the segment corresponding to the expressive and vocal state.
[0283] Optionally, in the above Figure 13 Based on the corresponding embodiments, in another embodiment of the eyebrow animation generation apparatus provided in this application, the determining unit 203 can specifically be used for:
[0284] Extract the sound intensity signal from the segment corresponding to the sound state;
[0285] Specifically, unit 203 can be used for:
[0286] The sound intensity signal is filtered by a low-pass filter to obtain a smooth distribution of the sound intensity signal;
[0287] The smooth distribution of the sound intensity signal is normalized to obtain the smooth value corresponding to the sound intensity signal, and the smooth value is used as the AU adjustment weight.
[0288] Optionally, in the above Figure 13 Based on the corresponding embodiments, in another embodiment of the eyebrow animation generation apparatus provided in this application, the determining unit 203 can specifically be used for:
[0289] Extract the pitch signal from the segment corresponding to the audio state;
[0290] Specifically, unit 203 can be used for:
[0291] The pitch signal is filtered by a low-pass filter to obtain a smooth distribution of the pitch signal;
[0292] The smooth distribution of the pitch signal is normalized to obtain the smooth value corresponding to the pitch signal, and the smooth value is used as the AU adjustment weight.
[0293] Optionally, in the above Figure 13Based on the corresponding embodiments, in another embodiment of the eyebrow animation generation apparatus provided in this application, the determining unit 203 can specifically be used for:
[0294] Extract the intensity and pitch signals from the segments corresponding to the audio state;
[0295] Specifically, unit 203 can be used for:
[0296] By using a low-pass filter, the intensity signal and the pitch signal are filtered separately to obtain a smooth distribution of the intensity signal and a smooth distribution of the pitch signal.
[0297] The smooth distribution of the sound intensity signal is normalized to obtain the smooth value corresponding to the sound intensity signal, and the smooth value is used as the adjustment weight of the first AU.
[0298] The smooth distribution of the pitch signal is normalized to obtain the smooth value corresponding to the pitch signal, and the smooth value is used as the adjustment weight of the second AU.
[0299] Optionally, in the above Figure 13 Based on the corresponding embodiments, in another embodiment of the eyebrow animation generation device provided in this application,
[0300] The acquisition unit 202 is also used to extract the accented segment and the intensity signal in the accented segment from the segment corresponding to the sound state.
[0301] The processing unit 201 is also used to use the ratio of the intensity signal in the accented segment to the intensity threshold as an enhancement coefficient;
[0302] The processing unit 201 is also used to adjust the AU adjustment weight based on the enhancement coefficient to obtain the target AU adjustment weight.
[0303] Optionally, in the above Figure 13 Based on the corresponding embodiments, in another embodiment of the eyebrow animation generation apparatus provided in this application, the processing unit 201 may specifically be used for:
[0304] According to the order of the target time periods corresponding to each segment of the eyebrow animation curve, the eyebrow animation curves of each segment are spliced together to obtain the spliced eyebrow animation curve.
[0305] For each pair of adjacent eyebrow animation curves in the spliced eyebrow animation curve, the corresponding segment curve at the splicing point is obtained according to the basic window length, and the segment curve is Gaussian smoothed to obtain the target eyebrow animation curve.
[0306] This application also provides a schematic diagram of another computer device, such as... Figure 14 As shown, Figure 14This is a schematic diagram of a computer device structure provided in an embodiment of this application. The computer device 300 can vary significantly due to different configurations or performance. It may include one or more central processing units (CPUs) 310 (e.g., one or more processors) and a memory 320, and one or more storage media 330 (e.g., one or more mass storage devices) for storing application programs 331 or data 332. The memory 320 and storage media 330 can be temporary or persistent storage. The program stored in the storage media 330 may include one or more modules (not shown in the diagram), each module including a series of instruction operations on the computer device 300. Furthermore, the CPU 310 may be configured to communicate with the storage media 330 and execute the series of instruction operations in the storage media 330 on the computer device 300.
[0307] Computer device 300 may also include one or more power supplies 340, one or more wired or wireless network interfaces 350, one or more input / output interfaces 360, and / or one or more operating systems 333, such as Windows Server. TM MacOSX TM Unix TM Linux TM FreeBSD TM etc.
[0308] The aforementioned computer device 300 is also used to perform, for example Figures 2 to 11 The steps in the corresponding embodiments.
[0309] Another aspect of this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements... Figures 2 to 11 The steps in the method described in the illustrated embodiment.
[0310] Another aspect of this application provides a computer program product comprising a computer program, which, when executed by a processor, implements as follows: Figures 2 to 11 The steps in the method described in the illustrated embodiment.
[0311] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0312] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.
[0313] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0314] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0315] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
Claims
1. A method for generating eyebrow animation, characterized in that, include: The audio to be processed is divided into audio segments and silent segments, wherein each audio segment corresponds to an audio start and end timestamp, and each silent segment corresponds to a silent start and end timestamp. Based on the expression tags corresponding to the audio to be processed, the audio to be processed is divided into expression segments and expressionless segments. The expression tags include expression types and expression start and end timestamps corresponding to each expression type. Each expression segment corresponds to an expression type and an expression start and end timestamp. Each expressionless segment corresponds to an expressionless start and end timestamp. Based on the start and end timestamps of the audio segment, the start and end timestamps of the silent segment, the start and end timestamps of the facial expression segment, and the start and end timestamps of the expressionless segment, the intersection of the audio segment, the silent segment, the facial expression segment, and the expressionless segment is obtained to obtain the emotional state corresponding to each target time period. The emotional state includes a silent expressionless state, a silent expressionless state, a voiceless expressionless state, or a voiceless expressionless state. Based on the expression-driven AU rule, the AU of the eyebrows associated with the emotional state corresponding to each target time period is determined. Specifically, when the emotional state is the expressionless and silent state or the expressionless and vocal state, the AU of the eyebrows is the default AU of the eyebrow movement state associated with the expressionless state; when the emotional state is the expressionful and silent state, the AU of the eyebrows is the first expression AU of the eyebrow movement state associated with the expression type in the expressionful and silent state; and when the emotional state is the expressionful and vocal state, the AU of the eyebrows is the second expression AU of the eyebrow movement state associated with the expression type in the expressionful and vocal state. Interpolation calculations are performed on the eyebrow AU associated with the emotional state corresponding to each target time period to obtain the segment eyebrow animation curve corresponding to each emotional state. Specifically, when the emotional state is the expressionless and silent state, the segment eyebrow animation curve is determined based on the default AU and a first random seed point on the segment corresponding to the expressionless and silent state; when the emotional state is the expressive and silent state, the segment eyebrow animation curve is determined based on the first expression AU, the first AU weight corresponding to the first expression AU, and a second random seed point on the segment corresponding to the expressive and silent state; when the emotional state is the expressionless and vocal state, the segment eyebrow animation curve is determined based on the default AU, the AU adjustment weight corresponding to the sound signal in the vocal state segment, and a third random seed point on the segment corresponding to the expressionless and vocal state; and when the emotional state is the expressive and vocal state, the segment eyebrow animation curve is determined based on the second expression AU, the AU adjustment weight corresponding to the sound signal in the vocal state segment, and a fourth random seed point on the segment corresponding to the expressive and vocal state. The eyebrow animation curves of each segment are spliced together to obtain the target eyebrow animation curve; Based on the target eyebrow animation curve, generate the eyebrow animation corresponding to the audio to be processed.
2. The method according to claim 1, characterized in that, When the emotional state is the expressionless and silent state; the step of determining the AU of the eyebrows associated with the emotional state for each target time period based on the expression-driven AU rule includes: Multiple first random seed points on the segment corresponding to the expressionless and silent state are calculated using a random function; For each of the multiple first random seed points, extract the extended segment corresponding to each first random seed point from the segment corresponding to the expressionless and silent state; The expression-driven AU rule selects the default AU of the eyebrow movement state associated with the expressionless state as the AU of the eyebrow. The step of interpolating the AU of the eyebrows associated with the emotional state corresponding to each target time period to obtain the eyebrow animation curve for each segment corresponding to the emotional state includes: Interpolation calculations are performed based on each of the first random seed points, the extended segments corresponding to each of the first random seed points, and each of the default AUs to obtain the eyebrow animation curve corresponding to the expressionless and silent state.
3. The method according to claim 1, characterized in that, When the emotional state is the expressive but silent state; the step of determining the AU of the eyebrows associated with the emotional state for each target time period based on the expression-driven AU rule includes: Multiple second random seed points on the segment corresponding to the expressive but silent state are calculated using a random function; For each of the multiple second random seed points, extract the extended segment corresponding to each second random seed point from the segment corresponding to the expressionless silent state; The expression-driven AU rule selects the first expression AU of the eyebrow movement state associated with the expression type in the expressionless state as the AU of the eyebrow. The step of interpolating the AU of the eyebrows associated with the emotional state corresponding to each target time period to obtain the eyebrow animation curve for each segment corresponding to the emotional state includes: By using fuzzy rules, the weight of the first AU corresponding to each of the first expressions is obtained; Interpolation calculations are performed based on each second random seed point, the corresponding extended segment of each second random seed point, and each first expression AU. The interpolation calculation results are then weighted based on the weight of each first AU to obtain the eyebrow animation curve corresponding to the expressionless state.
4. The method according to claim 1, characterized in that, When the emotional state is the expressionless but vocal state; the step of determining the AU of the eyebrows associated with the emotional state for each target time period based on the expression-driven AU rule includes: Multiple third random seed points on the segment corresponding to the expressionless yet vocal state are calculated using a random function; For each of the multiple third random seed points, extract the extended segment corresponding to each third random seed point from the segment corresponding to the expressionless but vocal state; The expression-driven AU rule selects the default AU of the eyebrow movement state associated with the expressionless state as the AU of the eyebrow. The step of interpolating the AU of the eyebrows associated with the emotional state corresponding to each target time period to obtain the eyebrow animation curve for each segment corresponding to the emotional state includes: Extract the sound signal from the segment corresponding to the audio state; The sound signal is smoothed to obtain a smoothed value, and the smoothed value is used as the AU adjustment weight. Interpolation calculations are performed based on each of the third random seed points, the corresponding extended segments of each of the third random seed points, and each of the default AUs. The interpolation calculation results are then weighted based on the AU adjustment weights to obtain the eyebrow animation curve corresponding to the expressionless yet vocal state.
5. The method according to claim 1, characterized in that, When the emotional state is the expressive and vocal state; the step of determining the AU of the eyebrows associated with the emotional state for each target time period based on the expression-driven AU rule includes: Multiple fourth random seed points on the segment corresponding to the expressive and vocal state are calculated using a random function; For each of the multiple fourth random seed points, extract the extended segment corresponding to each fourth random seed point from the segment corresponding to the expressive and vocal state; The expression-driven AU rule selects the second expression AU of the eyebrow movement state associated with the expression type in the expressive and vocal state as the AU of the eyebrow. The step of interpolating the AU of the eyebrows associated with the emotional state corresponding to each target time period to obtain the eyebrow animation curve for each segment corresponding to the emotional state includes: Extract the sound signal from the segment corresponding to the audio state; The sound signal is smoothed to obtain a smoothed value, and the smoothed value is used as the AU adjustment weight. Interpolation calculations are performed based on each of the fourth random seed points, the extended segments corresponding to each of the fourth random seed points, and each of the second expression AUs. The interpolation calculation results are then weighted based on the AU adjustment weights to obtain the eyebrow animation curve of the segment corresponding to the expressive and vocal state.
6. The method according to any one of claims 4 or 5, characterized in that, When the sound signal is a sound intensity signal, the extraction of the sound signal from the segment corresponding to the sound state includes: Extract the sound intensity signal from the segment corresponding to the sound state; The step of smoothing the sound signal to obtain a smoothed value corresponding to the sound signal, and using the smoothed value as the AU adjustment weight, includes: The sound intensity signal is filtered by a low-pass filter to obtain a smooth distribution of the sound intensity signal; The smooth distribution of the sound intensity signal is normalized to obtain the smooth value corresponding to the sound intensity signal, and the smooth value is used as the AU adjustment weight.
7. The method according to any one of claims 4 or 5, characterized in that, When the sound signal is a pitch signal, the extraction of the sound signal from the segment corresponding to the sound state includes: Extract the pitch signal from the segment corresponding to the sound state; The step of smoothing the sound signal to obtain a smoothed value corresponding to the sound signal, and using the smoothed value as the AU adjustment weight, includes: The pitch signal is filtered by a low-pass filter to obtain a smooth distribution of the pitch signal; The smooth distribution of the pitch signal is normalized to obtain the smooth value corresponding to the pitch signal, and the smooth value is used as the AU adjustment weight.
8. The method according to any one of claims 4 or 5, characterized in that, When the sound signal is a sound intensity signal and a sound pitch signal, the extraction of the sound signal from the segment corresponding to the sound state includes: Extract the intensity and pitch signals from the segment corresponding to the sound state; The step of smoothing the sound signal to obtain a smoothed value corresponding to the sound signal, and using the smoothed value as the AU adjustment weight, includes: The intensity signal and the pitch signal are filtered by a low-pass filter to obtain a smooth distribution of the intensity signal and a smooth distribution of the pitch signal. The smooth distribution of the sound intensity signal is normalized to obtain the smooth value corresponding to the sound intensity signal, and the smooth value is used as the first AU adjustment weight. The smooth distribution of the pitch signal is normalized to obtain the smooth value corresponding to the pitch signal, and the smooth value is used as the second AU adjustment weight.
9. The method according to claim 6, characterized in that, After normalizing the smooth distribution of the sound intensity signal, obtaining the smoothed value corresponding to the sound intensity signal, and using the smoothed value as the AU adjustment weight, the method further includes: Extract the accented segments and the intensity signals of the accented segments from the segments corresponding to the audio state; The ratio of the intensity signal in the stressed segment to the intensity threshold is used as the enhancement coefficient; The AU adjustment weight is adjusted based on the enhancement coefficient to obtain the target AU adjustment weight.
10. The method according to claim 1, characterized in that, The step of stitching together the eyebrow animation curves of each segment to obtain the target eyebrow animation curve includes: According to the chronological order of the target time periods corresponding to each segment of the eyebrow animation curve, the eyebrow animation curves of each segment are spliced together to obtain the spliced eyebrow animation curve. For each pair of adjacent eyebrow animation curves in the spliced eyebrow animation curve, the corresponding segment curve at the splicing point is obtained according to the basic window length, and the segment curve is Gaussian smoothed to obtain the target eyebrow animation curve.
11. A device for generating eyebrow animation, characterized in that, include: The processing unit is used to divide the audio to be processed into audio segments and silent segments, wherein each audio segment corresponds to an audio start and end timestamp, and each silent segment corresponds to a silent start and end timestamp. The processing unit is further configured to divide the audio to be processed into expressive segments and expressionless segments based on the expression tags corresponding to the audio to be processed. The expression tags include expression types and expression start and end timestamps corresponding to each expression type. Each expression segment corresponds to an expression type and an expression start and end timestamp. The expressionless segment corresponds to an expressionless start and end timestamp. The acquisition unit is used to acquire the intersection of the audio segment, the silent segment, the facial expression segment, and the expressionless segment based on the audio start and end timestamp, the silent start and end timestamp, the facial expression start and end timestamp, and the expressionless segment to obtain the emotional state corresponding to the target time period. The emotional state includes an expressionless and silent state, an expressionless and silent state, an expressionless and audio state, or an expressionless and audio state. A determining unit is configured to determine the eyebrow AU associated with the emotional state for each target time period based on expression-driven AU rules. Specifically, when the emotional state is a blank, silent state or a blank, vocal state, the eyebrow AU is the default AU of the eyebrow movement state associated with the blank state; when the emotional state is an expressive, silent state, the eyebrow AU is the first expression AU of the eyebrow movement state associated with the expression type in the expressive, silent state; and when the emotional state is an expressive, vocal state, the eyebrow AU is the second expression AU of the eyebrow movement state associated with the expression type in the expressive, vocal state. The processing unit is further configured to perform interpolation calculations on the AU of the eyebrows associated with the emotional state corresponding to each target time period to obtain the segment eyebrow animation curve corresponding to each emotional state. Specifically, when the emotional state is the expressionless and silent state, the segment eyebrow animation curve is determined based on the default AU and a first random seed point on the segment corresponding to the expressionless and silent state; when the emotional state is the expressive and silent state, the segment eyebrow animation curve is determined based on the first expression AU, a first AU weight corresponding to the first expression AU, and a second random seed point on the segment corresponding to the expressive and silent state; when the emotional state is the expressionless and vocal state, the segment eyebrow animation curve is determined based on the default AU, an AU adjustment weight corresponding to the sound signal in the vocal state segment, and a third random seed point on the segment corresponding to the expressionless and vocal state; and when the emotional state is the expressive and vocal state, the segment eyebrow animation curve is determined based on the second expression AU, an AU adjustment weight corresponding to the sound signal in the vocal state segment, and a fourth random seed point on the segment corresponding to the expressive and vocal state. The processing unit is also used to splice the eyebrow animation curves of each segment to obtain the target eyebrow animation curve; The processing unit is also used to generate an eyebrow animation corresponding to the audio to be processed based on the target eyebrow animation curve.
12. The apparatus according to claim 11, characterized in that, When the emotional state is the expressionless and silent state; the determining unit is specifically used for: Multiple first random seed points on the segment corresponding to the expressionless and silent state are calculated using a random function; For each of the multiple first random seed points, extract the extended segment corresponding to each first random seed point from the segment corresponding to the expressionless and silent state; The expression-driven AU rule selects the default AU of the eyebrow movement state associated with the expressionless state as the AU of the eyebrow. The processing unit is specifically used to: perform interpolation calculations based on each first random seed point, the extended segment corresponding to each first random seed point, and each default AU to obtain the segment eyebrow animation curve corresponding to the expressionless and silent state.
13. The apparatus according to claim 11, characterized in that, When the emotional state is the expressive but silent state; the determining unit is specifically used for: Multiple second random seed points on the segment corresponding to the expressive but silent state are calculated using a random function; For each of the multiple second random seed points, extract the extended segment corresponding to each second random seed point from the segment corresponding to the expressionless silent state; The expression-driven AU rule selects the first expression AU of the eyebrow movement state associated with the expression type in the expressionless state as the AU of the eyebrow. The processing unit is specifically used for: By using fuzzy rules, the weight of the first AU corresponding to each of the first expressions is obtained; Interpolation calculations are performed based on each second random seed point, the corresponding extended segment of each second random seed point, and each first expression AU. The interpolation calculation results are then weighted based on the weight of each first AU to obtain the eyebrow animation curve corresponding to the expressionless state.
14. The apparatus according to claim 11, characterized in that, When the emotional state is the expressionless but vocal state; the determining unit is specifically used for: Multiple third random seed points on the segment corresponding to the expressionless yet vocal state are calculated using a random function; For each of the multiple third random seed points, extract the extended segment corresponding to each third random seed point from the segment corresponding to the expressionless but vocal state; The expression-driven AU rule selects the default AU of the eyebrow movement state associated with the expressionless state as the AU of the eyebrow. The processing unit is specifically used for: Extract the sound signal from the segment corresponding to the audio state; The sound signal is smoothed to obtain a smoothed value, and the smoothed value is used as the AU adjustment weight. Interpolation calculations are performed based on each of the third random seed points, the corresponding extended segments of each of the third random seed points, and each of the default AUs. The interpolation calculation results are then weighted based on the AU adjustment weights to obtain the eyebrow animation curve corresponding to the expressionless yet vocal state.
15. The apparatus according to claim 11, characterized in that, When the emotional state is the expressive and vocal state; the determining unit is specifically used for: Multiple fourth random seed points on the segment corresponding to the expressive and vocal state are calculated using a random function; For each of the multiple fourth random seed points, extract the extended segment corresponding to each fourth random seed point from the segment corresponding to the expressive and vocal state; The expression-driven AU rule selects the second expression AU of the eyebrow movement state associated with the expression type in the expressive and vocal state as the AU of the eyebrow. The processing unit is specifically used for: Extract the sound signal from the segment corresponding to the audio state; The sound signal is smoothed to obtain a smoothed value, and the smoothed value is used as the AU adjustment weight. Interpolation calculations are performed based on each of the fourth random seed points, the extended segments corresponding to each of the fourth random seed points, and each of the second expression AUs. The interpolation calculation results are then weighted based on the AU adjustment weights to obtain the eyebrow animation curve of the segment corresponding to the expressive and vocal state.
16. A computer device comprising a memory, a processor, and a bus system, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 10; The bus system is used to connect the memory and the processor to enable communication between the memory and the processor.
17. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 10.
18. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 10.