A method of generating an eye expression

By analyzing the energy and fundamental frequency sequence of the text-to-speech system, identifying the speech rhythm center, and adjusting the eye expressions of the virtual character based on weight adjustments, the problem of inconsistency between the virtual character's eye expressions and speech output was solved, thereby improving user trust and the naturalness of the interaction.

CN122244241APending Publication Date: 2026-06-19BEIJING KEYI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING KEYI TECH CO LTD
Filing Date
2026-03-16
Publication Date
2026-06-19

AI Technical Summary

Technical Problem

In existing technologies, the eye expressions of virtual characters are not coordinated with their voice output, resulting in inconsistent voice and emotion, which reduces user trust and the naturalness of the interaction.

Method used

By analyzing the energy and fundamental frequency sequence of the text-to-speech system, the speech rhythm center is identified, and the eye expressions of the virtual character are adjusted based on weight adjustment. The naturalness of the eye animation is enhanced by combining text content and emotion recognition model.

Benefits of technology

It achieves precise coordination between the virtual character's eye expressions and voice output, enhancing user trust and the naturalness of the interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122244241A_ABST
    Figure CN122244241A_ABST
Patent Text Reader

Abstract

This application provides a method for generating eye expressions, which accurately and effectively generates eye expressions, thereby improving user trust and the naturalness of interaction. In this embodiment, the electronic device determines the rhythm center based on the dB sequence of the TTS (Text-to-Speech) and then determines whether to determine the adjustment weight based on the dB sequence or the pitch sequence based on whether the pitch difference in the pitch sequence within the rhythm center's region meets a preset threshold. This adjustment weight is then used to adjust the initial display state of the virtual character's eyes in the animation, thereby generating eye expressions that match the speech rhythm. This ensures that the generated eye expressions correspond to the rhythmic stresses, energy shifts, and intonation transitions in the speech, thus enabling accurate and effective eye expression generation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a method for generating eye expressions. Background Technology

[0002] With the rapid development of artificial intelligence, natural language processing, and human-computer interaction technologies, text-to-speech (TTS) systems have been widely applied in scenarios such as intelligent voice assistants, in-vehicle navigation, virtual digital humans, online education, and game character dialogue. Users' expectations for voice interaction are no longer limited to simply being able to "hear clearly," but rather they seek a natural, "emotional," and "human-like" expressive experience. Therefore, how to coordinate the facial expressions and movements of virtual characters with their voice output has become a key issue in enhancing the immersive experience of interaction.

[0003] In existing technologies, TTS-driven facial animation methods typically employ phoneme synchronization or prosodic mapping to match mouth shape with pronunciation. For example, they control lip opening and closing shapes by identifying the phoneme category corresponding to the current pronunciation (such as / a / , / i / , / u / , etc.); or they adjust the amplitude of movements corresponding to the intensity of the voice using energy and fundamental frequency information. However, these methods mostly focus on mouth modeling and have relatively weak support for non-voice-driven micro-expression control, especially advanced emotional expression features such as eye contact, eyebrow movements, and eyelid movements. If the virtual character merely plays speech mechanically without corresponding eye feedback, it can easily create a disconnect between voice and emotion, reducing user trust and the naturalness of the interaction. Summary of the Invention

[0004] This application provides an eye expression method, apparatus, and electronic device for accurately and effectively generating eye expressions, thereby improving user trust and natural interaction.

[0005] This application provides a method for generating eye expressions, the method comprising: Obtain the energy intensity (decibel, dB) sequence and time-frequency fundamental frequency (pitch) sequence of the text-to-speech (TTS) data; Based on the variation characteristics of the dB sequence, at least one rhythm center that satisfies the preset rhythm conditions is determined; the pitch difference of the pitch sequence in the neighborhood of each rhythm center is determined; if the pitch difference satisfies the preset threshold, the adjustment weight is determined based on the dB sequence; otherwise, the adjustment weight is determined based on the pitch sequence. For each rhythm center, in the event window corresponding to that rhythm center, the initial display state of the eye is adjusted based on the corresponding determined adjustment weight.

[0006] In one possible implementation, determining at least one rhythm center that satisfies a preset rhythm condition based on the variation characteristics of the dB sequence includes: If the dB values ​​of more than a preset number of target sampling points in the dB sequence are all greater than a preset value, then the sampling point corresponding to the maximum dB change is determined as the rhythm center based on the change in dB value between each target sampling point and its corresponding adjacent sampling point.

[0007] In one possible implementation, adjusting the initial eye display state based on a determined adjustment weight within the event window corresponding to the rhythm center includes: The vertical scaling ratio of the eye animation is determined based on the adjustment weight, wherein the larger the adjustment weight, the larger the vertical scaling ratio; the offset value of the corresponding eye shape is calculated according to the vertical scaling ratio, and the initial display state of the eye is adjusted in the event window corresponding to the rhythm center based on the vertical scaling ratio and the offset value.

[0008] In one possible implementation, before determining the vertical scaling ratio of the eye animation based on the adjusted weights, the method further includes: Obtain the text content corresponding to the TTS. If the text content contains any preset symbol, including question mark, exclamation mark, and tilde, then determine the symbol modification factor stored for the preset symbol; wherein the symbol modification factor is greater than 1. The vertical scaling ratio is updated based on the product of the symbol modification factor and the vertical scaling ratio. In the event window corresponding to the rhythm center, the initial eye display state is adjusted based on the vertical scaling ratio and the offset value, including: Within the target window where the event window corresponding to the rhythm center overlaps with the symbol window of a preset number of adjacent characters, the initial eye display state is adjusted based on the updated vertical scaling ratio and the offset value.

[0009] In one possible implementation, it also includes: The text content corresponding to the TTS is obtained. For each rhythm center, the sub-text content corresponding to the event window where the rhythm center is located and the information of the sampling point corresponding to the rhythm center are input into the emotion recognition model to obtain the emotion type and emotion intensity output by the emotion recognition model. Based on the emotion type and the emotion intensity, an emotion baseline factor is determined; The vertical scaling ratio is updated based on the product of the emotion baseline factor and the vertical scaling ratio. The adjustment of the initial eye display state based on the vertical scaling ratio and the offset value includes: The initial display state of the eye is adjusted based on the updated vertical scaling ratio and the offset value.

[0010] In one possible implementation, before determining the vertical scaling ratio of the eye animation based on the adjusted weights, the method further includes: Based on a preset time length, the time of each moment within the event window corresponding to the rhythm center is normalized to determine the normalized time. The normalized time is processed by a slow-in / slow-out function to determine the rhythm factor; The vertical scaling ratio is updated based on the product of the rhythm factor and the vertical scaling ratio. The adjustment of the initial eye display state based on the vertical scaling ratio and the offset value includes: The initial display state of the eye is adjusted based on the updated vertical scaling ratio and the offset value.

[0011] In one possible implementation, the step of processing the normalized time using a slow-in / slow-out function to determine the rhythm factor includes: The rhythm factor is determined using the following formula:

[0012] in, g(u) is the rhythm factor determined at time u, and a, b, and c are all preset values.

[0013] In one possible implementation, before determining the pitch difference in the neighborhood of each rhythm center, the method further includes: The dB at each rhythm center is converted into a linear amplitude, and each linear amplitude is normalized based on a preset maximum linear amplitude to obtain each target linear amplitude; rhythm centers with target linear amplitudes greater than a preset amplitude threshold are determined to have rhythm events. Based on the rhythm center where rhythm events exist, the subsequent step of determining the pitch difference of the pitch sequence in the neighborhood of each rhythm center is performed.

[0014] In one possible implementation, adjusting the initial eye display state based on a corresponding determined adjustment weight includes: According to the order of each rhythm center in the sequence of each rhythm center, polarity is alternately assigned to each rhythm event in turn; wherein the polarities of two adjacent rhythm events are opposite, and the polarity is used to identify magnification or reduction; The initial eye display state is adjusted based on the adjusted weights and the values ​​stored for the corresponding polarities.

[0015] In one possible implementation, before adjusting the initial eye display state based on the corresponding determined adjustment weights, the method further includes: Based on the voice playback speed, the actual number of playback frames and duration of multiple voice segments in the TTS are estimated; and the playback time period corresponding to each voice segment is determined in real time by combining the playback progress, the actual number of playback frames, and the duration; and the event window corresponding to each rhythm center is determined based on the playback time period of each voice segment.

[0016] In this embodiment, the electronic device determines the rhythm center based on the dB sequence of the TTS, and then determines whether to determine the adjustment weight based on the dB sequence or the pitch sequence based on whether the pitch difference in the pitch sequence within the region of the rhythm center meets a preset threshold. Based on the adjustment weight, the initial display state of the virtual character's eyes in the animation is adjusted, thereby generating eye expressions that match the rhythm of the speech. This ensures that the generated eye expressions match the rhythmic stress, energy changes, and intonation transitions in the speech, thus enabling accurate and effective eye expression generation. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 A schematic diagram illustrating the process of an eye expression generation method provided in this application embodiment; Figure 2 A schematic diagram illustrating a process for determining rhythmic events, provided in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of an eye expression generation device provided in an embodiment of this application; Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0020] The application scenarios described in this application are for the purpose of more clearly illustrating the technical solutions of this application, and do not constitute a limitation on the technical solutions provided in this application. Those skilled in the art will understand that with the emergence of new application scenarios, the technical solutions provided in this application are also applicable to similar technical problems. In the description of this application, unless otherwise stated, "multiple" means two or more.

[0021] In order to accurately and effectively generate eye expressions, this application provides a method for generating eye expressions.

[0022] The eye expression generation method includes: an electronic device acquiring a dB sequence and a pitch sequence of TTS; determining at least one rhythm center that meets preset rhythm conditions based on the variation characteristics of the dB sequence; determining the pitch difference in the neighborhood of each rhythm center; if the pitch difference meets a preset threshold, determining an adjustment weight based on the dB sequence; otherwise, determining an adjustment weight based on the pitch sequence; and for each rhythm center, adjusting the initial eye display state in the event window corresponding to that rhythm center based on the determined adjustment weight.

[0023] Figure 1 This application provides a schematic diagram of a method for generating eye expressions, which includes the following steps: S101: Obtain the dB sequence and pitch sequence of the TTS.

[0024] The eye expression generation method provided in this application is applied to an electronic device, which can be a smart device such as a PC or a server.

[0025] In this embodiment, the electronic device can first acquire TTS (Text-to-Speech). In one possible implementation, the TTS engine does not output audio streams as whole sentences or full texts, but rather uses a mechanism of pushing audio streams by semantic segments. Each speech segment of the TTS is pushed to the electronic device as an independent data unit. This design not only meets the low-latency requirements of real-time interactive scenarios but also provides a fine-grained time control basis for subsequent rhythm analysis and facial expression generation. Each speech segment chunk (k) It contains two core prosodic feature sequences: dB sequence and dB sequence. (k)[n] and pitch sequence (k) [n], where k represents the k-th speech segment, and n is the frame index within that speech segment, ranging from 0 to N. k -1, while N k This represents the total number of valid audio frames contained in the k-th speech segment. Here, a "frame" refers to a temporally localized data unit obtained through short-time analysis of the speech signal; typically, each frame corresponds to 10-25 milliseconds of audio content. Therefore, dB (k) [n] reflects the acoustic energy of the nth frame in the kth speech segment, usually expressed in decibels (dB). Essentially, it is the logarithmic expression of the square of the signal amplitude in that frame, used to measure the loudness of the sound; while (k) [n] represents the pitch of the nth frame in the kth speech segment, which is the perceived frequency of the vibration of the speech body. After normalization, it can be used for consistency modeling across speakers and speech rates.

[0026] To improve the accuracy of rhythm center detection, it is essential to eliminate the influence of irrelevant interference factors, especially false low-energy regions that may be introduced by leading silence segments. Although these silence frames do not carry valid speech information, they are still included in the rhythm analysis process without processing, leading to misjudgment or missed detection. Therefore, embodiments of this application define a silence determination mechanism. In one possible implementation, if the dB of a frame is less than a first threshold and the pitch is less than a second threshold, then the frame can be determined to be a silence frame.

[0027] In one possible implementation, the silence frame can be determined using the following formula:

[0028] in, and These are preset energy thresholds and preset fundamental frequency silence thresholds, used to comprehensively determine whether a frame is in a silent state. Only when the energy is extremely low and there are no significant periodic speech characteristics is it identified as a silent frame. Subsequently, the electronic device can remove consecutive leading silent frames, retaining only the effective speech portion after the first non-silent frame, forming a denoised segment. ,in, , This indicates the number of valid frames remaining after removing silence. This preprocessing step ensures that all subsequent rhythm analysis is based on real speech content, improving the robustness and consistency of key event recognition.

[0029] Considering that important events in speech rhythm (such as stress and emphasis) often span multiple speech segments, simply cutting the audio stream by segment boundaries may cause rhythm peaks to fall precisely between two segments, resulting in splitting and affecting detection integrity. To address this issue, embodiments of this application introduce a cross-segment overlapping splicing mechanism to establish temporal continuity between adjacent speech segments. Specifically, when processing the k-th speech segment, the electronic device retains the last L frames of the previously processed speech segment as the overlapping area. Then, the tail data is merged with the complete content of the current speech segment into a new analysis window through a splicing operation: The '||' operator represents sequence concatenation, not numerical addition. It doesn't simply sum the two pitches or intensities; instead, it places the last L frames of the previous segment before the current segment, creating an extended window with temporal continuity. This mechanism effectively avoids the break in key rhythmic signals caused by segmentation, allowing the rhythmic center to be accurately identified within a more complete context.

[0030] To further achieve collaborative control and event synchronization, this application embodiment constructs a unified voice timeline key frame identifier (Frame Identifier, FrameID) as the time reference for various dynamic behaviors throughout the electronic device. The electronic device internally maintains a global frame counter G, initially set to 0, to track the total number of all processed valid voice frames. For the current segment... The absolute time position of the m-th frame in the time frame is determined by the following formula: This means that each frame is assigned a globally unique identifier, which reflects its relative position within the overall audio stream and supports cross-module referencing. For example, when a rhythmic center is detected at FrameID=1256, the animation engine can precisely trigger eye movements based on this ID, while external callback interfaces can also use this to report event logs or synchronize other visual feedback. After processing the current segment, the global counter is updated. This provides a sequential number for the next audio segment. Thus, FrameID becomes a bridge connecting TTS output, rhythm analysis, facial expression generation, and external systems, achieving true audiovisual synchronization and event traceability.

[0031] S102: Based on the variation characteristics of the dB sequence, determine at least one rhythm center that satisfies the preset rhythm conditions; determine the pitch difference of the pitch sequence in the neighborhood of each rhythm center; if the pitch difference satisfies the preset threshold, determine the adjustment weight based on the dB sequence; otherwise, determine the adjustment weight based on the pitch sequence.

[0032] To achieve natural coordination between virtual character eye expressions and speech content, electronic devices can first perform dynamic rhythm analysis based on the dB of TTS (Text-to-Speech) to identify multiple key locations with significant prosodic changes as rhythm centers. These rhythm centers are not simply high-energy points, but rather locations exhibiting significant energy transitions within a local context, such as rapid energy increases, prominent peaks, or sustained levels above the background. These features collectively constitute the change characteristics. In one possible implementation, the electronic device can filter candidate locations by setting preset rhythm conditions. These conditions include, but are not limited to: an energy change rate exceeding a threshold between adjacent frames, the existence of local maxima, and the absence of closely competing peaks before and after the location. In another possible implementation, only frames that simultaneously meet multiple dynamic criteria can be identified as valid rhythm centers, thereby avoiding misjudging noise, breathing sounds, or ordinary intonation fluctuations as emphasis points.

[0033] To differentiate the intensity of linguistic intent carried by different rhythm centers, the electronic device uses pitch sequence as the basis for judgment. In one possible implementation, the fundamental frequency variation difference can be calculated within the neighborhood of each determined rhythm center, for example, within 3-5 frames before and after it; that is, the cumulative change between the current frame and neighboring frames. If this variation exceeds a preset intonation sensitivity threshold, the pitch sequence variation can be used as the basis for adjusting the weight; otherwise, the dB loudness variation can be used to adjust the weight. In one possible implementation, if the variation exceeds the preset intonation sensitivity threshold, the rhythm center can be considered to correspond to a significant intonation shift, belonging to the emphasized part of the language expression; otherwise, it is classified as a natural fluctuation in normal speech. Based on this analysis result, the electronic device can assign an event type to each rhythm center. This event type is an eye-animation triggered event type, which includes two categories: emphasized and normal. If the event type is emphasized, the pitch sequence variation is used as the weight to highlight the expressiveness brought by intonation fluctuations; if the event type is normal, the dB loudness variation is used as the weight to maintain a natural and smooth rhythm response.

[0034] In one possible implementation, the electronic device can calculate the pitch change rate near the rhythm center and normalize the pitch at the rhythm center. If the determined pitch change rate is greater than a preset value and the normalized pitch is greater than a preset threshold, the event type corresponding to the rhythm center can be determined to be emphasized; otherwise, the event type corresponding to the rhythm center can be determined to be normal.

[0035] For example, the rate of pitch change near the rhythm center can be calculated using the following formula:

[0036] in, For the determined pitch change rate, For the pitch at the center of the rhythm, The pitch of the previous frame corresponding to the rhythm center.

[0037] The pitch at the rhythm center can be normalized using the following formula:

[0038] in, The preset normalization factor is used to map the original pitch to the interval [0,1]. In one possible implementation, the maximum pitch value of a large number of samples can be counted, and the reciprocal of the maximum pitch value can be determined. .

[0039] In this application, the embodiments support the identification of multiple rhythm centers within the same speech tone in TTS and the independent determination of their event types, thereby adapting to the multi-emphasis expression needs under complex sentence structures. For example, in a statement containing two emphatic words: "This is not just a difficulty, but a huge challenge!", the electronic device can detect rhythmic abrupt changes at "huge" and "challenge" respectively, and determine that both are emphatic events based on the pitch change trend, thereby driving two independent but coordinated eye responses.

[0040] In one possible implementation, if the event type is ordinary, the adjustment weights can be determined based on the dB sequence, wherein the adjustment weights are determined based on the dB sequence within the event window corresponding to the rhythm center. In another possible implementation, the average value of each dB within the event window can be determined, and the adjustment weights stored for the average value can be determined. Alternatively, a sub-dB sequence composed of each dB within the event window according to the corresponding time order can be input into a large language model to obtain the adjustment weights output by the large language model. If the event type is emphasis-based, the adjustment weights can be determined based on the pitch sequence, wherein the adjustment weights are determined based on the pitch sequence within the event window corresponding to the rhythm center. In one possible implementation, the average value of each pitch within the event window can be determined, and the adjustment weights stored for the average value can be determined. Alternatively, a sub-pitch sequence composed of each pitch within the event window according to the corresponding time order can be input into a large language model to obtain the adjustment weights output by the large language model.

[0041] S103: For each rhythm center, in the event window corresponding to the rhythm center, adjust the initial eye display state based on the corresponding determined adjustment weight.

[0042] For each identified rhythm center, the electronic device can construct a temporally localized scope associated with that rhythm center, also known as an event window. The event window defines the range of eye animation influence triggered by the rhythm corresponding to that rhythm center. In one possible implementation, the event window can use the rhythm center as a time anchor, extending forward a certain number of frames as a preparatory phase and backward a certain number of frames as a recovery phase. The overall duration is typically set between 300ms and 600ms, and the specific length can be dynamically adjusted according to speech rate, emotional intensity, or device refresh rate. Within this window, the virtual character's eye state no longer remains static or transitions at a uniform speed, but instead undergoes semantically meaningful dynamic deformations based on the event type determined by the rhythm center. This design avoids the mechanical feel of traditional animation, where eyes are constantly open or blinking randomly, achieving precise alignment between vocal emphasis and visual expression.

[0043] In one possible implementation, an event window can be constructed within a speech segment based on the rhythm center:

[0044] in, and These are the preset frame numbers.

[0045] And it can be mapped to a global primary key:

[0046] in, This represents the number of frames in TTS where the rhythm center is located.

[0047] Within the event window, the initial eye display state of the virtual character in the animation is adjusted based on defined adjustment weights. This adjustment process does not interrupt other ongoing expressions, but rather overlays rhythm-driven instantaneous eye changes on top of them, achieving a fusion expression of multimodal emotions.

[0048] This application's embodiments utilize cross-speech segment overlapping and splicing, along with the FrameID primary key, to ensure the continuous existence of rhythmic events on the global timeline, avoiding inter-speech breaks and resets. By determining the rhythm center through dB sequences and combining normalized pitch and pitch trend joint detection, details such as stress, plosives, and rising intonation can be stably identified, making eye dynamics closer to the rhythm of human speech.

[0049] In this embodiment, the electronic device determines the rhythm center based on the dB sequence of the TTS, and then determines whether to determine the adjustment weight based on the dB sequence or the pitch sequence based on whether the pitch difference in the pitch sequence within the region of the rhythm center meets a preset threshold. Based on the adjustment weight, the initial display state of the virtual character's eyes in the animation is adjusted, thereby generating eye expressions that match the rhythm of the speech. This ensures that the generated eye expressions match the rhythmic stress, energy changes, and intonation transitions in the speech, thus enabling accurate and effective eye expression generation.

[0050] To determine the rhythm center, based on the above embodiments, in this embodiment of the application, determining at least one rhythm center that satisfies a preset rhythm condition based on the variation characteristics of the dB sequence includes: If the dB values ​​of more than a preset number of target sampling points in the dB sequence are all greater than a preset value, then the sampling point corresponding to the maximum dB change is determined as the rhythm center based on the change in dB value between each target sampling point and its corresponding adjacent sampling point.

[0051] To achieve coordination between voice and virtual character eye expressions, this application's embodiments not only rely on the absolute magnitude of dB values ​​but also emphasize their dynamic change trends. In one possible implementation, the dB sequence can be differentially processed to calculate the change in dB values ​​between adjacent sampling points. Here, sampling points can also be referred to as frames, and the change can be called the rate of change or slope. This slope reflects the degree of "sudden intensification" of the sound, capturing the initial emphasis in language, such as stressed words or the onset of interrogative sentences, more effectively than simply using peak values. The electronic device then searches for the position corresponding to the maximum value among all positive values ​​and identifies this position as the main rhythmic center within the current analysis window. This method effectively avoids misinterpreting long high notes as rhythmic points, ensuring that what is identified is a truly "explosive" semantic emphasis.

[0052] To improve detection robustness, the electronic device can introduce a dual-determination mechanism: the rhythm center search process is only initiated when the energy values ​​of multiple consecutive target sampling points are all higher than a preset threshold, representing non-background noise, and at least one of these points shows a significant upward trend. This design eliminates interference from transient noise or breathing sounds, ensuring that each triggering event has sufficient linguistic significance. Once the rhythm center is identified, the electronic device constructs a temporally localized event window centered on it and classifies the event type based on the pitch sequence.

[0053] To accurately and effectively generate eye expressions, based on the above embodiments, in this embodiment, the adjustment of the initial eye display state in the event window corresponding to the rhythm center, based on a corresponding determined adjustment weight, includes: The vertical scaling ratio of the eye animation is determined based on the adjustment weight, wherein the larger the adjustment weight, the larger the vertical scaling ratio; the corresponding eye shape offset value is calculated according to the vertical scaling ratio, and the initial display state of the eye is adjusted in the event window corresponding to the rhythm center based on the vertical scaling ratio and the offset value.

[0054] To achieve a high degree of coordination between the virtual character's eye movements and the rhythm of its speech, the electronic device can determine a corresponding vertical scaling ratio based on the adjustment weights assigned to the identified event types. This vertical scaling ratio controls the range of change in the degree of eyelid opening and closing. In one possible implementation, the vertical scaling ratio can be determined by multiplying the adjustment weight by an initial ratio; a larger adjustment weight results in a larger determined vertical scaling ratio. In another possible implementation, emphasized events will trigger a higher vertical scaling ratio, while ordinary events will only cause slight deformations, maintaining overall facial expression stability. The electronic device also calculates a corresponding offset value for the eye shape based on this vertical scaling ratio. In one possible implementation, the offset value for the eye shape can be determined using the following formula:

[0055] in, The offset value of the determined F-th frame, The preset weight values, The vertical scaling ratio determined for frame F-1.

[0056] In one possible implementation, the offset value not only includes static displacement, but can also introduce a dynamic hysteresis mechanism. That is, the offset value of the current frame is affected by the scaling state of the previous moment, simulating the inertia and recovery delay characteristics in real muscle movement, making the movement more natural and smooth.

[0057] At the visual level, each rhythmic event achieves facial expression changes by adjusting the geometric shape of the eye contour. Initially, the eyes are controlled by a set of N points. Where N is the number of control points, and i is the control point number. Let x be the x-coordinate of the i-th control point. Let be the ordinate of the i-th control point. The electronic device can describe the outline shape of the animated character's eye through N control points, where N is the center point of the control points: Within the event window, the electronic device generates the corresponding vertical scaling factor based on the current event type. and offset Used to update the ordinates of each control point: ,in, Let be the updated ordinate of the i-th control point. Let be the ordinate of the center point of N control points. This is the ordinate of the i-th control point before the update. For the determined vertical scaling ratio, The offset values ​​are defined, where the horizontal coordinate remains unchanged to maintain the basic structure of the eye shape, while the vertical deformation achieves dynamic effects such as opening and squinting the eyes.

[0058] In one possible implementation, when the time windows of multiple rhythmic events overlap, the electronic device does not simply superimpose or cover them. Instead, it employs a time-nearest neighbor priority or intensity-weighted fusion strategy to smoothly integrate the vertical scaling ratios and offset values ​​of different event outputs. For example, if two emphasis-type events are close together, they are merged into a single eye-opening action with a longer duration and moderate amplitude; if one is strong and the other weak, the former is prioritized and the latter is used for fine-tuning. This approach effectively prevents problems such as eye jitter, discontinuity, or excessive deformation caused by frequent triggering, ensuring the overall coherence and visual comfort of the animation. Finally, the updated set of control points forms a new eye contour and is rendered onto the virtual character model in real time, completing the closed-loop drive from voice rhythm to visual expression.

[0059] To accurately and effectively generate eye expressions, based on the above embodiments, in this embodiment, before determining the vertical scaling ratio of the eye animation based on the adjusted weights, the method further includes: Obtain the text content corresponding to the TTS. If the text content contains any preset symbol, including question mark, exclamation mark, and tilde, then determine the symbol modification factor stored for the preset symbol; wherein the symbol modification factor is greater than 1. The vertical scaling ratio is updated based on the product of the symbol modification factor and the vertical scaling ratio. In the event window corresponding to the rhythm center, the initial eye display state is adjusted based on the vertical scaling ratio and the offset value, including: Within the target window where the event window corresponding to the rhythm center overlaps with the symbol window of a preset number of adjacent characters, the initial eye display state is adjusted based on the updated vertical scaling ratio and the offset value.

[0060] To further enhance the emotional expressiveness of virtual characters, electronic devices not only rely on speech prosody features but also incorporate text content. In one possible implementation, before determining the vertical scaling ratio of the character's eye animation based on event type, the electronic device can first acquire the original text content corresponding to TTS and perform symbolic semantic analysis on that text.

[0061] The system detects whether the text contains any preset symbols, which are emotional cues such as question marks (?), exclamation marks (!), and tildes (5 characters). After detecting the presence of a preset symbol, a time region associated with that symbol, known as a symbol time window, can be constructed. This symbol time window is used to limit the scope of influence of subsequent animation enhancements and is usually aligned with the playback period of the speech segment containing the symbol.

[0062] For each preset symbol, the electronic device is pre-configured with a corresponding symbol modification factor, which is greater than 1 and is used to amplify the intensity of the basic animation response.

[0063] In one possible implementation, the sign modification factor can be determined using the following formula:

[0064] in, For the determined sign modifier, The value saved for this preset symbol. This is the symbol window corresponding to the preset symbol, where F is the number of the corresponding frame.

[0065] After determining the symbol modification factor, the product of the symbol modification factor and the vertical scaling ratio can be determined. Based on this product, the determined vertical scaling ratio is updated. Furthermore, within the target window where the symbol time window corresponding to the preset symbol and the event window corresponding to the rhythm center overlap, the initial display state of the virtual character's eyes in the animation is adjusted based on the updated vertical scaling ratio and offset value.

[0066] In one possible implementation, the updated vertical scaling ratio is determined as follows:

[0067] in, This is the updated vertical scaling ratio. This is the vertical scaling ratio before the update. It is a symbolic modifier.

[0068] Finally, the electronic device rendering uses vertical scaling and offset values ​​applied to each control point to adjust the initial display state of the virtual character's eyes in the animation.

[0069] To accurately and effectively generate eye expressions, based on the above embodiments, this application embodiment further includes: The text content corresponding to the TTS is obtained. For each rhythm center, the sub-text content corresponding to the event window where the rhythm center is located and the information of the sampling point corresponding to the rhythm center are input into the emotion recognition model to obtain the emotion type and emotion intensity output by the emotion recognition model. Based on the emotion type and the emotion intensity, an emotion baseline factor is determined; The vertical scaling ratio is updated based on the product of the emotion baseline factor and the vertical scaling ratio. The adjustment of the initial eye display state based on the vertical scaling ratio and the offset value includes: The initial display state of the eye is adjusted based on the updated vertical scaling ratio and the offset value.

[0070] This application further introduces an emotion enhancement mechanism to achieve more nuanced and emotionally nuanced facial expression responses. In one possible implementation, the electronic device can acquire the original text content corresponding to TTS and, for each identified rhythm center, extract the sub-text covered by the event window containing that rhythm center as a local semantic context. The event window is a time interval constructed using the rhythm center as a time anchor (e.g., 300ms before and after). The sub-text corresponding to this event window can be accurately mapped using a speech-text alignment algorithm. Simultaneously, the electronic device can also extract information about the sampling points corresponding to the rhythm center. In one possible implementation, it can extract the corresponding frame number of the sampling point, as well as the acoustic feature information corresponding to the rhythm center, including but not limited to: peak energy value, pitch change rate, duration, etc. Furthermore, it can also acquire the signal indicating the end of the dialogue.

[0071] The above subtext and extracted information are input into a pre-trained emotion recognition model. The emotion recognition model can output two key results: emotion type and the corresponding emotion intensity. The emotion type includes surprise, doubt, anger, joy, sadness, etc. In one possible implementation, the emotion intensity can be a continuous value between 0 and 1, representing the emotional significance of the segment.

[0072] Electronic devices can find the basic moderating coefficient corresponding to a certain emotion type based on a preset emotion mapping table. For example, anger corresponds to 1.4 and happiness corresponds to 1.2. The coefficient is then multiplied by the intensity of the emotion output by the emotion recognition model to obtain a dynamically weighted emotion score. The electronic device can then add this score to a preset baseline value to obtain the final emotion baseline factor.

[0073] In one possible implementation, the emotional baseline factor can be determined using the following formula:

[0074] in, For the baseline factor of sentiment, This is a preset value. The numerical value saved for this emotion type. The intensity of the emotion corresponding to this type of emotion output by the emotion recognition model.

[0075] After determining the emotion baseline factor, the product of the emotion baseline factor and the vertical scaling ratio can be determined. The determined vertical scaling ratio is updated based on this product, and the initial display state of the virtual character's eyes in the animation is adjusted based on the updated vertical scaling ratio and the offset value.

[0076] To accurately and effectively generate eye expressions, based on the above embodiments, in this embodiment, before determining the vertical scaling ratio of the eye animation based on the adjusted weights, the method further includes: Based on a preset time length, the time of each moment within the event window corresponding to the rhythm center is normalized to determine the normalized time. The normalized time is processed by a slow-in / slow-out function to determine the rhythm factor; The vertical scaling ratio is updated based on the product of the rhythm factor and the vertical scaling ratio. The adjustment of the initial eye display state based on the vertical scaling ratio and the offset value includes: The initial display state of the eye is adjusted based on the updated vertical scaling ratio and the offset value.

[0077] In animation control scenarios, to achieve more natural and rhythmic character eye animation, this application introduces a dynamic adjustment mechanism based on time rhythm in its embodiments. In one possible implementation, the timestamps of each frame within the event window centered on the rhythm center are normalized and mapped to a standardized time interval of [0,1] to obtain normalized time values ​​for subsequent rhythmic control.

[0078] For the j-th event, we can first determine the duration of the event window for event j: ,in, For the determined duration, This is the end time of the event window. Let F be the start time of the event window. For any frame within the event window, let its global time be F. Then the local time t of that frame within the current event window can be expressed as: , Where t is the determined local time, For the time of this frame, This is the start time of the event window. The time can be normalized using the following formula:

[0079] in, Let be the normalized time, and t be a defined local time. This is the duration of the event window for this event. The normalized time value reflects the relative position of the current frame within the entire event window (i.e., the time progress).

[0080] Electronic devices can employ ease-in / ease-out functions, such as Bézier curves or S-curve smoothing functions, to nonlinearly transform the normalized time, generating a rhythm factor that reflects the characteristics of changes in motion acceleration. This rhythm factor can simulate the dynamic characteristics of real biological movements—slow start, accelerated middle, and decelerated finish—thereby enhancing the visual smoothness and emotional expressiveness of animation.

[0081] After determining the vertical scaling ratio of the character's eye animation, and before adjusting the initial eye display state of the virtual character based on the vertical scaling ratio and the preset offset value, the electronic device further introduces rhythm control, multiplying the currently calculated rhythm factor with the original vertical scaling ratio to dynamically update the vertical scaling ratio.

[0082] This update mechanism allows the degree of eye deformation to change dynamically as the event progresses. For example, it amplifies the eye contraction or stretching effect at emotional climaxes and softens the transition at the beginning and end, thereby enhancing the expressive tension and rhythmic consistency of the animation.

[0083] Ultimately, based on the updated vertical scaling ratio and offset value, the electronic device adjusts the display parameters of the virtual character's initial eyes, such as geometry, texture coordinates, or skeletal weights, to achieve fine-grained control over the eye animation state. This method not only improves the accuracy of animation time perception but also achieves a high degree of synergy between emotional rhythm and visual expression, making it suitable for various virtual character interaction scenarios, including expression-driven, voice-synchronized, and emotion-feedback scenarios.

[0084] To accurately and effectively generate eye expressions, based on the above embodiments, in this embodiment, the step of processing the normalized time using an ease-in / ease-out function to determine a rhythm factor includes: The rhythm factor is determined using the following formula:

[0085] in, g(u) is the rhythm factor determined at time u, and a, b, and c are all preset values.

[0086] The final output rhythm factor is a weighted coefficient that evolves over time, reflecting the dynamic importance of the current moment within the event cycle. This rhythm factor will be used to multiply with eye animation parameters (such as vertical scaling) to achieve time-aware dynamic adjustment.

[0087] To achieve precise alignment between the rhythm center and the event window in voice-driven animation, this application proposes an adaptive time modeling method that combines text semantic features with actual playback dynamics. In one possible implementation, the electronic device can estimate the actual number of playback frames and duration of multiple voice segments based on the voice playback speed and text symbol semantics during the TTS generation process. Combined with the real-time playback progress, it can dynamically estimate and correct the actual number of playback frames and duration of each voice segment, thereby determining the time interval corresponding to each voice unit and locating key emotional nodes, i.e., the rhythm center and its associated event window.

[0088] In one possible implementation, a window can be estimated in the following way:

[0089] in, For the estimated window, The starting frame number of this event window. W represents the predicted number of frames to play for this segment based on factors such as text length and speech rate, where W is the preset buffer duration. This is the end frame number of the event window.

[0090] Correction based on actual cumulative frames during playback:

[0091] in, , For actual cumulative frames, The starting frame number of this event window. This is the predicted number of frames to be played for this segment based on factors such as text length and speaking speed.

[0092] In one possible implementation, the symbol window can also be determined in this way, and when the symbol time window is reached, corresponding adjustments are made based on the symbol modification factor.

[0093] To accurately and effectively generate eye expressions, based on the above embodiments, in this embodiment of the application, before determining the pitch difference of the pitch sequence in the neighborhood of each rhythm center, the method further includes: The dB at each rhythm center is converted into a linear amplitude, and each linear amplitude is normalized based on a preset maximum linear amplitude to obtain each target linear amplitude; rhythm centers with target linear amplitudes greater than a preset amplitude threshold are determined to have rhythm events. Based on the rhythm center where rhythm events exist, the subsequent step of determining the pitch difference of the pitch sequence in the neighborhood of each rhythm center is performed.

[0094] To ensure the rhythm center effectively drives the generation of virtual character eye animation, the electronic device not only identifies its location but also determines whether that location has sufficient semantic salience to constitute a genuine rhythmic event. In one possible implementation, after determining the rhythm center, the dB value at the rhythm center can be converted into a physical quantity in the linear amplitude domain to more realistically reflect the energy level of the sound. This conversion is achieved through the following formula:

[0095] in, The converted linear amplitude, It represents the dB value at the center of the rhythm.

[0096] This method restores the decibel value on a logarithmic scale to a linear amplitude proportional to the sound pressure level, thereby eliminating the influence of nonlinearity in human hearing perception and making the subsequent animation response more consistent with actual sound energy changes.

[0097] The electronic device normalizes this linear amplitude. In one possible implementation, normalization can be performed using the following formula:

[0098] in, A preset normalization factor is used to map the original linear amplitude to the [0,1] interval. In one possible implementation, the maximum linear amplitude of a large number of samples can be statistically analyzed, and the reciprocal of the maximum linear amplitude can be determined. .

[0099] The intensity of the rhythmic event can be determined based on the normalized linear amplitude. In one possible implementation, the normalized linear amplitude is directly used as the intensity representation of the rhythmic event. This intensity value not only reflects the volume of the sound but also incorporates its abrupt changes, thus accurately reflecting the degree of emphasis in the language. To further improve robustness and filter noise interference, the electronic device sets a preset intensity threshold. A rhythmic center is considered a valid rhythmic event only if the normalized linear amplitude exceeds the intensity threshold. This mechanism avoids unnecessary facial expression changes being triggered by brief background noise, breathing sounds, or minor speech fluctuations.

[0100] Once a rhythm event is confirmed, the electronic device can use the time position of the rhythm's center as a reference point, extending forward by a certain time, such as 100ms, and backward by a duration, such as 500ms, to construct an event window of fixed or dynamic length. This window is used to define the duration of the eye animation triggered by this rhythm and serves as the scope for subsequent parameter interpolation and state updates.

[0101] Figure 2 A schematic diagram of a process for determining a rhythmic event provided in an embodiment of this application includes the following steps: S201: Based on the variation characteristics of the dB sequence, determine the rhythm center that meets the preset rhythm conditions.

[0102] S202: Convert the dB at the rhythm center into a linear amplitude.

[0103] S203: Based on the preset maximum linear amplitude, normalize the linear amplitude to obtain the target linear amplitude.

[0104] S204: If the target linear amplitude is greater than the preset amplitude threshold, then it is determined that there is a rhythm event at the rhythm center.

[0105] If the target linear amplitude is greater than the preset amplitude threshold, it is determined that there is a rhythm event at the rhythm center, and an event window of a preset time length is constructed with the rhythm center as the center.

[0106] To accurately and effectively generate eye expressions, based on the above embodiments, in this embodiment, adjusting the initial display state of the eyes based on corresponding determined adjustment weights includes: According to the order of each rhythm center in the sequence of each rhythm center, polarity is alternately assigned to each rhythm event in turn; wherein the polarities of two adjacent rhythm events are opposite, and the polarity is used to identify magnification or reduction; The initial eye display state is adjusted based on the adjusted weights and the values ​​stored for the corresponding polarities.

[0107] After identifying multiple rhythm centers that meet preset rhythmic conditions, to create rhythmic alternation, the electronic device does not process each event in isolation. Instead, it introduces a dynamic grade assignment mechanism based on temporal relationships to enhance the natural transitions and rhythm between consecutive eye movements. In one possible implementation, all rhythm centers identified as valid rhythmic events are sorted according to their corresponding temporal order in TTS (Time-of-Sight) to obtain a sequence of rhythm centers, and each is assigned a grade identifier that alternates between two states: for example, a positive grade represents an expanding action, such as widening the eyes, while a negative grade represents a contracting action, such as slightly squinting or retracting the eyes. The grades of two adjacent rhythmic events always remain opposite.

[0108] The method provided in this application can prevent virtual characters from falling into a state of "expression lock" or visual fatigue due to repeatedly performing the same directional actions (such as continuously opening their eyes wide) when multiple emphatic statements appear consecutively. By introducing a graded alternation mechanism, a set of continuous rhythmic events is transformed into a sequence of eye deformations with a tension-relaxation rhythm, simulating the natural eye expression change process of humans in real conversations of "opening eyes - relaxing - refocusing," significantly improving the sense of layering and anthropomorphism of emotional expression.

[0109] In one possible implementation, the corresponding polarity can be determined using the following formula:

[0110] Where j represents the sequence corresponding to the rhythm center. This represents the polarity of the j-th rhythm center.

[0111] In one possible implementation, the event can be represented by the following structure:

[0112] in, For the structure of the j-th event, Information about the start time of the j-th event. Information about the end time of the j-th event. Let the polarity of the j-th event be... Let j be the event type of the j-th event. The intensity is determined based on the normalized linear amplitude.

[0113] All events can be arranged in chronological order to form a rhythmic event chain.

[0114] Electronic devices can participate in the subsequent animation generation process based on the determined level as a key parameter. After constructing the event window corresponding to each rhythm event, the electronic device can combine the determined event type (such as "emphasis type" corresponding to high amplitude and "normal type" corresponding to low amplitude) with the currently assigned level to jointly determine the adjustment method of the virtual character's eye contour: In one possible implementation, if it is a positive level, the control point is driven to stretch upward based on the vertical scaling ratio of the current event type to achieve the effect of opening or widening the eyes; if it is a negative level, the control point is driven to stretch downward based on the vertical scaling ratio of the current event type to achieve the effect of shrinking. A downward offset can be added to guide the eyelids to close slightly or converge towards the center, showing a converging or pensive expression.

[0115] In one possible implementation, the vertical scaling ratio can be determined using the following formula:

[0116] in, This is the vertical scaling ratio. The scaling factor corresponding to the given time u. The intensity corresponding to the j-th rhythm center is determined based on the normalized linear amplitude. Let j be the order of the rhythm center.

[0117] To accurately and effectively generate eye expressions, based on the above embodiments, in this embodiment of the application, before adjusting the initial display state of the eyes based on the corresponding determined adjustment weights, the method further includes: Based on the voice playback speed, the actual number of playback frames and duration of multiple voice segments in the TTS are estimated; and the playback time period corresponding to each voice segment is determined in real time by combining the playback progress, the actual number of playback frames, and the duration; and the event window corresponding to each rhythm center is determined based on the playback time period of each voice segment.

[0118] The detailed process of this embodiment is as follows: Upon arrival of a speech segment, silence is removed, and multiple speech segments are overlapped and spliced. A FrameID is generated for each frame in the speech segment. Multi-feature rhythm detection is performed in real-time to obtain multiple events, generate the event structure, and write it into the event chain. The rendering thread queries the event chain by FrameID, directly reads the pre-calculated event level, event type, and intensity, and calculates the vertical scaling ratio and offset value of the current frame. If the symbolic sentiment window is hit, a symbolic modifier factor is overlaid. If the latest sentiment callback is received, a sentiment baseline factor is overlaid. After fusion, the control points are updated and rendered. During offline animation playback, the control point output is overridden, but the event chain and sentiment baseline continue to advance in the background, smoothly transitioning back to real-time driving after the animation ends.

[0119] In this embodiment, the event chain is pre-written into the future window when the fragment arrives, and the rendering thread directly uses the driving parameters to generate synchronized expressions with almost zero latency. Symbolic emotions are dynamically corrected through the FrameID time window to solve the problem of early / late expression, ensuring consistency between semantic emotions and acoustic expression. Emotional baseline and rhythmic micro-motions are uniformly integrated. Emotional callbacks serve as the slow variable baseline, and rhythmic events serve as the fast variable micro-motions; all three are superimposed on the same FrameID to avoid mutual overwriting or conflict. During offline animation coverage, the event chain progresses in the background, and automatically and seamlessly connects to real-time driving after completion, ensuring overall continuity and naturalness. Rhythm detection and control point transformation are both lightweight computations that can run in real-time on the edge, making them suitable for robots, mobile devices, or embedded devices.

[0120] With the popularization of streaming TTS, real-time dialogue systems, and virtual digital human technologies, robots or smart terminals with screens need to generate eye expressions (such as eyelid opening and closing, micro-movements, and slight shifts) in real time, synchronized with the speech rhythm, to enhance naturalness and biomimicry. However, existing solutions typically suffer from the following technical problems: 1. Difficult to adapt to real-time output of streaming TTS Many expression-driven systems assume that speech is generated all at once or that rhythm analysis is performed after the entire sentence is ready. In streaming TTS scenarios, audio is broken down into segments of varying lengths, and the energy and timestamps between each segment may not be continuous. If "segment-independent analysis" is performed, rhythmic discontinuities can easily form at segment boundaries, leading to a desynchronization between eye expressions and speech.

[0121] 2. The rhythmic driving characteristics are singular, lacking multidimensional acoustic information.

[0122] Existing systems mostly rely on instantaneous energy or average dB to drive eye animation, ignoring multi-dimensional acoustic features such as dB change slope, pitch change, and short bursts. The resulting rhythmic patterns are coarse and monotonous, making it difficult to depict details such as stress, bursts, and inflections in human speech.

[0123] 3. Poor rhythmic continuity in long sentences or multiple segments of speech.

[0124] During long text readings and continuous dialogues, the TTS engine generates audio segments in multiple streams. If the rhythm analysis window is reset for each segment, the temporal continuity of rhythmic events will be interrupted, causing inter-segment resets or animation restarts, thus disrupting the natural and coherent rhythmic fluctuations.

[0125] 4. The emotional tone of the text symbols is difficult to align with the rhythm of the speech.

[0126] Symbols such as “?”, “!”, and “~” reflect tone of question, emphasis, and prolongation, but their actual expression is determined by the audio timeline. Traditional methods directly trigger emoticons during the text parsing stage without considering the actual playback time of streaming TTS, often resulting in problems such as question mark emoticons appearing prematurely and exclamation mark emoticons appearing delayed.

[0127] 5. The upper-level emotion modeling and the lower-level rhythm animation lack a unified timeline and interface.

[0128] Emotion models typically output emotion labels / intensities at a low frequency; while expression-driven modules run independently on the audio side. The two lack a unified time key and a clear interface, making it difficult to achieve unified expression control with "emotion as the baseline + rhythm as the micro-movement".

[0129] 6. Offline high-quality emoji assets are difficult to integrate smoothly with real-time rhythm-driven processing.

[0130] Traditional systems often overwrite real-time drivers when playing offline eye expression animations, causing the rhythm to stop during playback and abruptly change afterward, disrupting the overall naturalness.

[0131] Therefore, there is an urgent need for a method that can: adapt to streaming TTS segment output; use a unified speech timeline as the primary key; combine multiple acoustic features for rhythm detection; and support the alignment and fusion of text symbols and upper-level emotion models on the same timeline. This application addresses the problems of rhythm breaks, single energy drive, asynchrony between symbolic emotion and speech, and disconnect between emotion models and facial expression drivers in streaming TTS scenarios. This application proposes a method that uses the speech timeline primary key FrameID as a unified index, and through cross-segment audio continuity, multi-feature rhythm detection, and rhythm event chain pre-computation, fuses text symbolic emotion and upper-level emotion baselines. This allows the rendering thread to obtain future rhythm-driven parameters in advance, thereby performing continuous geometric deformation and behavior scheduling on eye contour control points, achieving low-latency, cross-segment continuous, and strictly synchronized real-time generation of eye expressions with speech.

[0132] Figure 3 This is a schematic diagram of the structure of an eye expression device provided in an embodiment of this application. The device includes: The acquisition module 301 is used to acquire the dB sequence and pitch sequence of the TTS; The processing module 302 is configured to determine at least one rhythm center that satisfies a preset rhythm condition based on the variation characteristics of the dB sequence; determine the pitch difference of the pitch sequence in the neighborhood of each rhythm center; if the pitch difference satisfies a preset threshold, determine an adjustment weight based on the dB sequence; otherwise, determine an adjustment weight based on the pitch sequence; and for each rhythm center, adjust the initial eye display state in the event window corresponding to that rhythm center based on the determined adjustment weight.

[0133] In one possible implementation, the processing module 302 is specifically used to determine the sampling point corresponding to the maximum dB change as the rhythm center if the dB values ​​of more than a preset number of consecutive target sampling points in the dB sequence are all greater than a preset value, based on the change in dB value between each target sampling point and its corresponding adjacent sampling point.

[0134] In one possible implementation, the processing module 302 is specifically used to determine the vertical scaling ratio of the eye animation based on the adjustment weight, wherein the larger the adjustment weight, the larger the vertical scaling ratio; calculate the offset value of the corresponding eye shape according to the vertical scaling ratio, and adjust the initial display state of the eye based on the vertical scaling ratio and the offset value in the event window corresponding to the rhythm center.

[0135] In one possible implementation, the processing module 302 is further configured to obtain the text content corresponding to the TTS; if the text content contains any preset symbol, wherein the preset symbol includes a question mark, an exclamation mark, or a tilde, then determine the symbol modification factor stored for the preset symbol; wherein the symbol modification factor is greater than 1; and update the vertical scaling ratio according to the product of the symbol modification factor and the vertical scaling ratio. The processing module 302 is specifically used to adjust the initial eye display state based on the updated vertical scaling ratio and the offset value within the target window where the event window corresponding to the rhythm center overlaps with the symbol window of a preset number of adjacent characters.

[0136] In one possible implementation, the processing module 302 is further configured to acquire the text content corresponding to the TTS, and for each rhythm center, input the sub-text content corresponding to the event window where the rhythm center is located and the information of the sampling point corresponding to the rhythm center into the emotion recognition model to acquire the emotion type and emotion intensity output by the emotion recognition model; determine the emotion baseline factor based on the emotion type and the emotion intensity; and update the vertical scaling ratio according to the product of the emotion baseline factor and the vertical scaling ratio. The processing module 302 is specifically used to adjust the initial display state of the eye based on the updated vertical scaling ratio and the offset value.

[0137] In one possible implementation, the processing module 302 is further configured to normalize the time of each moment within the event window corresponding to the rhythm center based on a preset time length, and determine the normalized time; process the normalized time through a easing-in / easing-out function to determine the rhythm factor; and update the vertical scaling ratio according to the product of the rhythm factor and the vertical scaling ratio. The processing module 302 is specifically used to adjust the initial display state of the eye based on the updated vertical scaling ratio and the offset value.

[0138] In one possible implementation, the processing module 302 is specifically used to determine the rhythm factor using the following formula:

[0139] in, g(u) is the rhythm factor determined at time u, and a, b, and c are all preset values.

[0140] In one possible implementation, the processing module 302 is further configured to convert dB at each rhythm center into a linear amplitude, and normalize each linear amplitude based on a preset maximum linear amplitude to obtain each target linear amplitude; determine that rhythm centers with target linear amplitudes greater than a preset amplitude threshold have rhythm events; and based on rhythm centers with rhythm events, perform the subsequent step of determining the pitch difference of the pitch sequence in the neighborhood of each rhythm center.

[0141] In one possible implementation, the processing module 302 is further configured to sequentially assign polarity to each rhythm event according to the order of each rhythm center in the sequence composed of each rhythm center; wherein the polarities of two adjacent rhythm events are opposite, and the polarity is used to identify magnification or reduction; and to adjust the initial display state of the eye based on the adjustment weight and the value stored for the polarity.

[0142] In one possible implementation, the processing module 302 is further configured to estimate the actual number of playback frames and duration of multiple audio segments of the TTS based on the audio playback speed; and determine the playback time period corresponding to each audio segment in real time by combining the playback progress, the actual number of playback frames and the duration; and determine the event window corresponding to each rhythm center based on the playback time period of each audio segment.

[0143] Figure 4 This application provides a schematic diagram of an electronic device structure based on an embodiment of the present application. In addition to the above embodiments, this application also provides an electronic device, such as... Figure 4As shown, it includes: processor 401, communication interface 402, memory 403 and communication bus 404, wherein processor 401, communication interface 402 and memory 403 communicate with each other through communication bus 404. The memory 403 stores a computer program, which, when executed by the processor 401, causes the processor 401 to perform the following steps: Obtain the dB sequence and pitch sequence of the TTS; Based on the variation characteristics of the dB sequence, at least one rhythm center that satisfies the preset rhythm conditions is determined; the pitch difference of the pitch sequence in the neighborhood of each rhythm center is determined; if the pitch difference satisfies the preset threshold, the adjustment weight is determined based on the dB sequence; otherwise, the adjustment weight is determined based on the pitch sequence. For each rhythm center, in the event window corresponding to that rhythm center, the initial display state of the eye is adjusted based on the corresponding determined adjustment weight.

[0144] In one possible implementation, determining at least one rhythm center that satisfies a preset rhythm condition based on the variation characteristics of the dB sequence includes: If the dB values ​​of more than a preset number of target sampling points in the dB sequence are all greater than a preset value, then the sampling point corresponding to the maximum dB change is determined as the rhythm center based on the change in dB value between each target sampling point and its corresponding adjacent sampling point.

[0145] In one possible implementation, adjusting the initial eye display state based on a determined adjustment weight within the event window corresponding to the rhythm center includes: The vertical scaling ratio of the eye animation is determined based on the adjustment weight, wherein the larger the adjustment weight, the larger the vertical scaling ratio; the offset value of the corresponding eye shape is calculated according to the vertical scaling ratio, and the initial display state of the eye is adjusted in the event window corresponding to the rhythm center based on the vertical scaling ratio and the offset value.

[0146] In one possible implementation, before determining the vertical scaling ratio of the eye animation based on the adjusted weights, the method further includes: Obtain the text content corresponding to the TTS. If the text content contains any preset symbol, including question mark, exclamation mark, and tilde, then determine the symbol modification factor stored for the preset symbol; wherein the symbol modification factor is greater than 1. The vertical scaling ratio is updated based on the product of the symbol modification factor and the vertical scaling ratio. In the event window corresponding to the rhythm center, the initial eye display state is adjusted based on the vertical scaling ratio and the offset value, including: Within the target window where the event window corresponding to the rhythm center overlaps with the symbol window of a preset number of adjacent characters, the initial eye display state is adjusted based on the updated vertical scaling ratio and the offset value.

[0147] In one possible implementation, it also includes: The text content corresponding to the TTS is obtained. For each rhythm center, the sub-text content corresponding to the event window where the rhythm center is located and the information of the sampling point corresponding to the rhythm center are input into the emotion recognition model to obtain the emotion type and emotion intensity output by the emotion recognition model. Based on the emotion type and the emotion intensity, an emotion baseline factor is determined; The vertical scaling ratio is updated based on the product of the emotion baseline factor and the vertical scaling ratio. The adjustment of the initial eye display state based on the vertical scaling ratio and the offset value includes: The initial display state of the eye is adjusted based on the updated vertical scaling ratio and the offset value.

[0148] In one possible implementation, before determining the vertical scaling ratio of the eye animation based on the adjusted weights, the method further includes: Based on a preset time length, the time of each moment within the event window corresponding to the rhythm center is normalized to determine the normalized time. The normalized time is processed by a slow-in / slow-out function to determine the rhythm factor; The vertical scaling ratio is updated based on the product of the rhythm factor and the vertical scaling ratio. The adjustment of the initial eye display state based on the vertical scaling ratio and the offset value includes: The initial display state of the eye is adjusted based on the updated vertical scaling ratio and the offset value.

[0149] In one possible implementation, the step of processing the normalized time using a slow-in / slow-out function to determine the rhythm factor includes: The rhythm factor is determined using the following formula:

[0150] in, g(u) is the rhythm factor determined at time u, and a, b, and c are all preset values.

[0151] In one possible implementation, before determining the pitch difference in the neighborhood of each rhythm center, the method further includes: The dB at each rhythm center is converted into a linear amplitude, and each linear amplitude is normalized based on a preset maximum linear amplitude to obtain each target linear amplitude; rhythm centers with target linear amplitudes greater than a preset amplitude threshold are determined to have rhythm events. Based on the rhythm center where rhythm events exist, the subsequent step of determining the pitch difference of the pitch sequence in the neighborhood of each rhythm center is performed.

[0152] In one possible implementation, adjusting the initial eye display state based on a corresponding determined adjustment weight includes: According to the order of each rhythm center in the sequence of each rhythm center, polarity is alternately assigned to each rhythm event in turn; wherein the polarities of two adjacent rhythm events are opposite, and the polarity is used to identify magnification or reduction; The initial eye display state is adjusted based on the adjusted weights and the values ​​stored for the corresponding polarities.

[0153] In one possible implementation, before adjusting the initial eye display state based on the corresponding determined adjustment weights, the method further includes: Based on the voice playback speed, the actual number of playback frames and duration of multiple voice segments in the TTS are estimated; and the playback time period corresponding to each voice segment is determined in real time by combining the playback progress, the actual number of playback frames, and the duration; and the event window corresponding to each rhythm center is determined based on the playback time period of each voice segment.

[0154] The communication bus mentioned in the above server can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.

[0155] The communication interface is used for communication between the aforementioned electronic devices and other devices.

[0156] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0157] The processors mentioned above can be general-purpose processors, including central processing units, network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits, field-programmable gate arrays or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0158] Based on the above embodiments, this application also provides a computer-readable storage medium storing a computer program executable by an electronic device. When the program is run on the electronic device, it causes the electronic device to perform the steps described above.

[0159] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0160] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0161] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0162] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0163] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A method for generating eye expressions, characterized in that, The method includes: Obtain the energy intensity dB sequence and time-frequency fundamental frequency pitch sequence of text-to-speech (TTS); Based on the variation characteristics of the dB sequence, at least one rhythm center that satisfies the preset rhythm conditions is determined; the pitch difference of the pitch sequence in the neighborhood of each rhythm center is determined; if the pitch difference satisfies the preset threshold, the adjustment weight is determined based on the dB sequence; otherwise, the adjustment weight is determined based on the pitch sequence. For each rhythm center, in the event window corresponding to that rhythm center, the initial display state of the eye is adjusted based on the corresponding determined adjustment weight.

2. The method according to claim 1, characterized in that, The step of determining at least one rhythm center that satisfies a preset rhythm condition based on the variation characteristics of the dB sequence includes: If the dB values ​​of more than a preset number of target sampling points in the dB sequence are all greater than a preset value, then the sampling point corresponding to the maximum dB change is determined as the rhythm center based on the change in dB value between each target sampling point and its corresponding adjacent sampling point.

3. The method according to claim 1, characterized in that, In the event window corresponding to the rhythm center, the initial eye display state is adjusted based on the corresponding determined adjustment weight, including: The vertical scaling ratio of the eye animation is determined based on the adjustment weight, wherein the larger the adjustment weight, the larger the vertical scaling ratio; the offset value of the corresponding eye shape is calculated according to the vertical scaling ratio, and the initial display state of the eye is adjusted in the event window corresponding to the rhythm center based on the vertical scaling ratio and the offset value.

4. The method according to claim 3, characterized in that, Before determining the vertical scaling ratio of the eye animation based on the adjusted weights, the method further includes: Obtain the text content corresponding to the TTS. If the text content contains any preset symbol, including question mark, exclamation mark, and tilde, then determine the symbol modification factor stored for the preset symbol; wherein the symbol modification factor is greater than 1. The vertical scaling ratio is updated based on the product of the symbol modification factor and the vertical scaling ratio. In the event window corresponding to the rhythm center, the initial eye display state is adjusted based on the vertical scaling ratio and the offset value, including: Within the target window where the event window corresponding to the rhythm center overlaps with the symbol window of a preset number of adjacent characters, the initial eye display state is adjusted based on the updated vertical scaling ratio and the offset value.

5. The method according to claim 3, characterized in that, Also includes: The text content corresponding to the TTS is obtained. For each rhythm center, the sub-text content corresponding to the event window where the rhythm center is located and the information of the sampling point corresponding to the rhythm center are input into the emotion recognition model to obtain the emotion type and emotion intensity output by the emotion recognition model. Based on the emotion type and the emotion intensity, an emotion baseline factor is determined; The vertical scaling ratio is updated based on the product of the emotion baseline factor and the vertical scaling ratio. The adjustment of the initial eye display state based on the vertical scaling ratio and the offset value includes: The initial display state of the eye is adjusted based on the updated vertical scaling ratio and the offset value.

6. The method according to claim 3, characterized in that, Before determining the vertical scaling ratio of the eye animation based on the adjusted weights, the method further includes: Based on a preset time length, the time of each moment within the event window corresponding to the rhythm center is normalized to determine the normalized time. The normalized time is processed by a slow-in / slow-out function to determine the rhythm factor; The vertical scaling ratio is updated based on the product of the rhythm factor and the vertical scaling ratio. The adjustment of the initial eye display state based on the vertical scaling ratio and the offset value includes: The initial display state of the eye is adjusted based on the updated vertical scaling ratio and the offset value.

7. The method according to claim 6, characterized in that, The process of processing the normalized time using a slow-in / slow-out function to determine the rhythm factor includes: The rhythm factor is determined using the following formula: in, g(u) is the rhythm factor determined at time u, and a, b, and c are all preset values.

8. The method according to claim 1, characterized in that, Before determining the pitch difference in the neighborhood of each rhythm center, the method further includes: The dB at each rhythm center is converted into a linear amplitude, and each linear amplitude is normalized based on a preset maximum linear amplitude to obtain each target linear amplitude; rhythm centers with target linear amplitudes greater than a preset amplitude threshold are determined to have rhythm events. Based on the rhythm center where rhythm events exist, the subsequent step of determining the pitch difference of the pitch sequence in the neighborhood of each rhythm center is performed.

9. The method according to claim 1, characterized in that, The adjustment of the initial eye display state based on the corresponding determined adjustment weights includes: According to the order of each rhythm center in the sequence of each rhythm center, polarity is alternately assigned to each rhythm event in turn; wherein the polarities of two adjacent rhythm events are opposite, and the polarity is used to identify magnification or reduction; The initial eye display state is adjusted based on the adjusted weights and the values ​​stored for the corresponding polarities.

10. The method according to any one of claims 1-9, characterized in that, Before adjusting the initial eye display state based on the corresponding determined adjustment weights, the method further includes: Based on the voice playback speed, the actual number of playback frames and duration of multiple voice segments in the TTS are estimated; and the playback time period corresponding to each voice segment is determined in real time by combining the playback progress, the actual number of playback frames, and the duration; and the event window corresponding to each rhythm center is determined based on the playback time period of each voice segment.