Digital human application method and electronic device

By calculating the emotional value of a song to generate digital human voice and facial animation, the problem of boredom during the waiting period for a performance is solved, the singer's sense of participation and the expressiveness of the song are improved, and the entertainment experience of digital audiovisual is enhanced.

CN119648957BActive Publication Date: 2025-11-25FUJIAN STAR NET EVIDEO INFORMATION SYST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411435799.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-15
Publication Date
2025-11-25
Estimated Expiration
2044-10-15

AI Technical Summary

Technical Problem

In digital audiovisual venues, singers are prone to losing their passion and becoming bored during the waiting period before their performance, which affects the overall expressiveness and appeal of the song.

Method used

By calculating the emotional value of a song and matching it with a pre-generated emotional text template, a digital human voice and facial animation that matches the singer's timbre and facial features are generated and played during the waiting period, enriching the visual and auditory effects and enhancing the singer's sense of participation and immersion.

Benefits of technology

It improved the singer's performance experience, enhanced the overall expressiveness and appeal of the song, and avoided the decline in entertainment experience caused by long waiting times.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119648957B_ABST
    Figure CN119648957B_ABST
Patent Text Reader

Abstract

The application discloses a digital person application method and electronic equipment, mainly through calculating the emotional value of a song, randomly selecting a script in a pre-generated emotional script template; then generating a singer script voice corresponding to the emotional value according to the singer's timbre; meanwhile, a digital person is shaped according to the singer's photo, which has a facial expression corresponding to the emotional value; during the waiting singing period, the digital person is driven to play the singer script voice, and the facial animation corresponding to the emotion is played at the same time; not only the visual presentation and auditory effect are enriched, but also the communication of the script is more lively and interesting, and the singer's participation and immersion are enhanced; the emotional color in the voice can be more intuitively expressed, the singer is fully conveyed the emotion of the song, the singer better understands and feels the emotion behind the song, the overall performance and appeal of the song singing are improved, and the singer's long waiting time is avoided, and their entertainment experience is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of digital audiovisual technology, and in particular to a digital human application method and electronic device. Background Technology

[0002] In digital audiovisual venues, singers can fulfill their singing needs through on-demand systems. With the rapid development and continuous innovation of on-demand technology and systems, singers' experience of singing as an entertainment activity is undergoing unprecedented enrichment and diversification, bringing them a richer and more diverse singing entertainment experience, and profoundly influencing their expectations and pursuits of singing entertainment.

[0003] The basic framework of a song consists of an intro, verse, chorus, bridge, interlude, and outro. The intro, interlude, and outro are typically purely instrumental sections showcasing solo or ensemble techniques, serving as waiting periods for the singer to begin or continue. In existing streaming systems, the singer must complete these waiting periods before the performance can begin or continue. This waiting time can cause the singer's emotional expression to gradually dissipate, hindering the full delivery of the song's emotions and significantly diminishing its overall expressiveness and impact. Furthermore, prolonged waiting times can be monotonous and uninteresting, easily causing impatience and reducing the singer's enjoyment. Summary of the Invention

[0004] Therefore, a digital human application method is needed to solve the problems that the singer's emotions may gradually dissipate during the waiting process of a song, and that long waiting times are too monotonous and boring, lacking interest.

[0005] To achieve the above objectives, the present invention provides a digital human application method, which includes the following steps:

[0006] Calculate the emotional value of the song to be played;

[0007] Match the corresponding emotional copy template from multiple pre-generated emotional copy templates based on the song's emotional value;

[0008] Based on the original singer's timbre and the matched emotional text template, the corresponding singer's voice is generated in the audio model, and the emotion of the singer's voice is adjusted according to the emotional value to generate the singer's text voice.

[0009] The time period for playing the singer's lyrics is marked based on the song's anchor point information;

[0010] A digital human simulating the singer's image is generated based on the singer's facial features, and the digital human's facial animation is matched with the song's emotional value.

[0011] The digital human is driven to play the singer's text during the marked time period, while the digital human's face displays facial animations corresponding to the emotions.

[0012] Furthermore, in the step of calculating the emotional value of the song to be played, the emotional value of the song is calculated using a music emotional value model or a lyrics emotional value model.

[0013] Furthermore, the lyrics sentiment value model includes a text preprocessing module, a text feature extraction module, a sentiment analysis module, and a sentiment scoring module;

[0014] The specific steps for calculating the emotional value of a song using the lyrics emotional value model are as follows:

[0015] The text preprocessing module performs word segmentation and stop word removal on the song lyrics and outputs the preprocessed text.

[0016] The text feature extraction module extracts the feature words from the text output by the text preprocessing module.

[0017] The sentiment analysis module analyzes the sentiment value corresponding to each feature word.

[0018] The emotion scoring module calculates the sum of the emotion values ​​corresponding to all feature words, which is then used as the emotion value of the song.

[0019] Furthermore, the step of matching the corresponding emotional copywriting template among the pre-generated multiple emotional copywriting templates based on the emotional value of the song specifically includes: each pre-generated emotional copywriting template is associated with an emotional value range, the emotional value range into which the emotional value of the song falls is determined, and a matching relationship is established between the song and the emotional copywriting template associated with the corresponding emotional value range.

[0020] Furthermore, after matching the corresponding emotional copy template among the pre-generated multiple emotional copy templates based on the emotional value of the song, the process further includes: combining the emotional copy template with the user information of the song to be played, or the location information of the place where the song to be played is located, or a greeting, to form an updated emotional copy template.

[0021] The step of generating the corresponding singer's voice in the audio model based on the original singer's timbre and the matched emotional text template is as follows: generating the corresponding singer's voice in the audio model based on the original singer's timbre and the updated emotional text template.

[0022] Furthermore, the step of marking the time period for playing the singer's text voice based on the anchor information of the song specifically involves: obtaining the anchor information of the song by automatically recognizing the vocals in the song, obtaining the waiting time period of the song based on the anchor information, and marking one of the waiting time periods as the time period for playing the singer's text voice.

[0023] Furthermore, the step of marking the time period for playing the singer's text voice based on the song's anchor point information includes determining whether the duration of the time period for playing the singer's text voice is greater than the duration of the singer's text voice; if it is greater, then the singer's text voice is played; otherwise, it is not played.

[0024] Furthermore, before generating the digital human simulating the singer's image based on the original singer's facial feature information, the process further includes a step of generating the original singer's facial feature information, specifically:

[0025] Use OpenCV to load photos of the original singer of a song and convert them to grayscale;

[0026] The face detector detects faces in grayscale images and obtains the coordinates of facial feature information.

[0027] The facial features are classified based on the coordinates of the detected facial feature information, and the facial feature information of the original singer of the song is generated.

[0028] Furthermore, adjusting the singer's voice emotion based on the emotion value includes: the total emotion value of the song is 1; when the emotion value of the song is lower than 0.5, it is determined to be a sad song, and at least one of the following singer voice parameters is reduced: speech rate, pitch, or volume; when the emotion value of the song is higher than 0.8, it is determined to be a cheerful song, and at least one of the following singer voice parameters is increased: speech rate, pitch, or volume.

[0029] Furthermore, when the digital human is driven to play the singer's text during the marked time period, the song playback volume is simultaneously reduced, and the emotional text is displayed on the screen in the form of animation effects. When the singer's text playback is complete, the song playback volume is restored.

[0030] An electronic device includes a memory and a processor, wherein the memory stores a computer program, and the computer program, when executed by the processor, implements the steps of the above-described digital human application method.

[0031] Unlike existing technologies, the above-mentioned technical solution mainly calculates the emotional value of a song and selects a matching emotional text template from pre-generated templates. It then generates a singer's voice text corresponding to that emotional value based on the singer's timbre. Simultaneously, it creates a digital human based on the singer's photo, with facial expressions corresponding to that emotional value. During the waiting period before the performance, the digital human plays the singer's voice text while simultaneously displaying corresponding facial animations. This digital human presentation of the singer's voice text not only enriches the visual and auditory experience but also makes the text more vivid and engaging, enhancing the singer's sense of participation and immersion. The combination of the digital human and the singer's voice text allows for a more intuitive expression of the emotional nuances in the voice, fully conveying the song's emotions to the singer, engaging the singer's emotions, helping them better understand and feel the emotions behind the song, and ultimately improving the overall expressiveness and impact of the performance. This also avoids long waiting times for the singer and reduces their entertainment experience. Attached Figure Description

[0032] Figure 1 This is a flowchart illustrating the digital human application method described in this invention. Detailed Implementation

[0033] To explain in detail the technical content, structural features, objectives, and effects of the technical solution, the following description is provided in conjunction with specific embodiments and accompanying drawings.

[0034] In this document, the term "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The term "embodiment" appearing in various places throughout the specification does not necessarily refer to the same embodiment, nor does it specifically limit its independence or connection with other embodiments. In principle, in this application, as long as there are no technical contradictions or conflicts, the technical features mentioned in each embodiment can be combined in any way to form corresponding implementable technical solutions.

[0035] Unless otherwise defined, the technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the use of related terms herein is merely for the purpose of describing particular embodiments and is not intended to limit this application.

[0036] In the description of this application, the term "and / or" is used to describe the logical relationship between objects, indicating that three relationships can exist. For example, A and / or B means: A exists, B exists, and A and B exist simultaneously. Additionally, the character " / " in this document generally indicates that the preceding and following objects have an "or" logical relationship.

[0037] In this application, terms such as “first” and “second” are used only to distinguish one entity or operation from another, and do not necessarily require or imply any actual quantity, hierarchy or order relationship between these entities or operations.

[0038] Unless otherwise specified, the use of terms such as “comprising,” “including,” “having,” or other similar expressions in this application is intended to cover non-exclusive inclusion, which does not exclude the presence of additional elements in a process, method, or product that includes the stated elements, such that a process, method, or product that includes a list of elements may include not only those defined elements but also other elements not expressly listed, or elements inherent to such a process, method, or product.

[0039] Similar to the interpretation in the Patent Examination Guidelines, in this application, expressions such as "greater than," "less than," and "exceeding" are understood to exclude the stated number; expressions such as "above," "below," and "within" are understood to include the stated number. Furthermore, in the description of the embodiments in this application, "multiple" means two or more (including two), and similar expressions related to "multiple" are also interpreted in this way, such as "multiple groups" and "multiple times," unless otherwise explicitly specified.

[0040] In the description of the embodiments of this application, the space-related expressions used, such as "center," "longitudinal," "lateral," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "vertical," "top," "bottom," "inner," "outer," "clockwise," "counterclockwise," "axial," "radial," and "circumferential," indicate the orientation or positional relationship based on the orientation or positional relationship shown in the specific embodiments or drawings. They are only for the purpose of describing the specific embodiments of this application or for the reader's understanding, and do not indicate or imply that the device or component referred to must have a specific position, a specific orientation, or be constructed or operated in a specific orientation. Therefore, they should not be construed as limitations on the embodiments of this application.

[0041] Unless otherwise expressly specified or limited, the terms "installation," "connection," "linking," "fixing," and "setting," as used in the description of the embodiments of this application, should be interpreted broadly. For example, "connection" can be a fixed connection, a detachable connection, or an integral setting; it can be a mechanical connection, an electrical connection, or a communication connection; it can be a direct connection or an indirect connection through an intermediate medium; it can be the internal connection of two components or the interaction between two components. For those skilled in the art to which this application pertains, the specific meaning of the above terms in the embodiments of this application can be understood according to the specific circumstances.

[0042] See Figure 1 As shown, this invention provides a digital human application method, which mainly calculates the emotional value of a song and selects a matching emotional text template from pre-generated emotional text templates; then, it generates a singer's text voice corresponding to the emotional value based on the singer's timbre; simultaneously, it creates a digital human based on the singer's photo, which has facial expressions corresponding to the emotional value; during the waiting period before the performance, it drives the digital human to play the singer's text voice while simultaneously playing facial animations corresponding to the emotions; by displaying the singer's text voice through the digital human, not only is the visual presentation and auditory effect enriched, but the text is also conveyed more vividly and interestingly, enhancing the singer's sense of participation and immersion; the combination of the digital human and the singer's text voice can more intuitively express the emotional color in the voice, fully convey the song's emotions to the singer, evoke the singer's emotions, help the singer better understand and feel the emotions behind the song, perform the song better, improve the overall expressiveness and appeal of the song performance, and avoid long waiting times for the singer, thus reducing their entertainment experience.

[0043] To further illustrate the invention's digital human application method, the following implementation method includes the following steps:

[0044] S1. Calculate the emotional value of the song to be played;

[0045] S2. Match the corresponding emotional copy template from multiple pre-generated emotional copy templates based on the song's emotional value;

[0046] S3. Based on the original singer's timbre and the matched emotional text template, generate the corresponding singer's voice in the audio model, and adjust the emotion of the singer's voice according to the emotional value to generate the singer's text voice.

[0047] S4. Mark the time period for playing the singer's text audio based on the song's anchor point information;

[0048] S5. Generate a digital human simulating the singer's image based on the facial features of the original singer of the song, and match the digital human's facial animation according to the emotional value of the song.

[0049] S6. During the marked time period for playing the singer's script, drive the digital human to play the singer's script, while simultaneously displaying facial animations corresponding to the emotions on the digital human's face.

[0050] The emotional value of the above-mentioned songs can be calculated in real time or pre-calculated. In some embodiments, the emotional value of the songs to be sung in the singing list is obtained by extracting the music features of the songs, such as the intensity, rhythm speed, melody, harmony, etc. of the music. Specifically, the emotional value of the songs can be calculated by training a suitable music emotional value model. In some embodiments, the emotional value of the songs in the singing list is obtained by extracting the lyric features of the songs, such as verbs, adjectives, nouns, etc. in the lyrics. Specifically, the emotional value of the songs can be calculated by training a suitable lyric emotional value model, that is, using the lyric emotional value model to calculate the emotional value of the songs. Of course, after calculation, the songs can also be classified and labeled according to their different emotional values for direct call next time.

[0051] In a specific embodiment, a lyric emotional value model is used to calculate the emotional value of the songs. The lyric emotional value model can be a natural language processing model (such as snowNLP), which includes a text preprocessing module, a text feature extraction module, an emotional analysis module, an emotional scoring module, etc.

[0052] The text preprocessing module performs word segmentation and stop word removal on the lyrics of the songs. The word segmentation can use word segmentation tools (such as jieba分词器 (Chinese), NLTK (English), etc.) to cut the lyric text into a large number of single words or phrases. By presetting a stop word list, this list usually contains some words that frequently appear in the text but are meaningless for emotional analysis (such as "de", "shi", "he", etc.). After word segmentation, each word in the text is traversed. If the word appears in the stop word list, it is removed from the text.

[0053] The text feature extraction module extracts the feature words output by the text preprocessing module. The feature words are mainly the text word types, such as verbs, adjectives, nouns, etc.

[0054] The emotional analysis module analyzes the emotional value corresponding to each feature word. The emotional analysis module contains a large number of feature words and their corresponding emotional values, that is, each feature word can map to a corresponding emotional value. Machine learning or deep learning algorithms (such as Naive Bayes, Support Vector Machine, Convolutional Neural Network, etc.) can be used to train the emotional analysis module. During the training process, the model will learn the association between the feature words and the emotional values. After the emotional analysis module is trained, the feature words are input into the emotional analysis module, that is, the feature words are mapped to the corresponding emotional values.

[0055] The sentiment scoring module calculates the sentiment values ​​corresponding to all feature words, which are then used as the song's sentiment value. Specifically, the feature words output by the text feature extraction module are matched with the sentiment analysis module to obtain the corresponding sentiment values. The sentiment value of the entire lyrics is then calculated, resulting in the song's sentiment value. The calculation method can be a simple summation or weighted average of the feature word sentiment values, or a more complex model output. The calculated song's sentiment value is then used as output for subsequent analysis or application.

[0056] The aforementioned emotional value is a value within a preset emotional value range, which is divided into several ranges, each with an associated emotional text template. That is, each pre-generated emotional text template is associated with an emotional value range. During matching, the range into which the song's emotional value falls is determined, and a matching relationship is established between the song and the corresponding emotional text template associated with that range. Of course, there are also cases where the calculated emotional value exceeds the preset emotional value range. In this case, emotional words can be added to the song's lyrics, and the song's emotional value can be recalculated. Specifically, the step of matching the song's emotional value with the corresponding emotional text template from among the pre-generated emotional text templates also includes determining whether the emotional score conforms to the emotional value range of the pre-generated emotional text templates; if it does, the song is matched with the corresponding emotional text template based on the emotional value; otherwise, emotional words are added to the song's lyrics, and the song's emotional value is recalculated. The calculated emotional value of the song can be saved directly for subsequent performances, or it can be calculated in real-time each time it is performed.

[0057] The emotions represented by the aforementioned emotional value range can gradually increase or gradually decrease. For example, if the emotional value range is set to (0, 1), the larger the emotional value, the more positive the emotion, with a maximum of 1 indicating the most cheerful song; the smaller the emotional value, the more negative the emotion, with a minimum of 0 indicating the most sorrowful song. The pre-generated emotional text templates are pre-generated templates, each associated with an emotional value range. Each template contains several texts, and the texts within each template express the same or similar emotional values. When a song is matched with a particular emotional text template, any text from the matching template is selected for the song.

[0058] The following embodiments use four emotional copywriting templates as examples to further illustrate the present invention. They are: emotional copywriting template 1, corresponding to an emotional value of 0-0.2 for the song; emotional copywriting template 2, corresponding to an emotional value of 0.2-0.5 for the song; emotional copywriting template 3, corresponding to an emotional value of 0.5-0.8 for the song; and emotional copywriting template 4, corresponding to an emotional value of 0.8-1.0 for the song.

[0059] Furthermore, for ease of understanding, the following specific embodiments are provided:

[0060] Emotional copywriting template 1: Corresponding to the song's emotional value of 0-0.2:

[0061] When the emotional value of a song is in the range of 0-0.2, it indicates that the song is very sad and the singer is in great pain. In this case, the accompanying text needs to be more heartfelt and sincere.

[0062] Example of copywriting:

[0063] I can feel your pain right now; it's truly unbearable. But please remember, you are not alone, and I will always be by your side.

[0064] Emotional copywriting template 2: Corresponding song emotional value 0.2-0.5:

[0065] If the current song is judged to be somewhat sad and the singer is also in a slightly depressed mood, then the corresponding text should be encouraging.

[0066] Example copywriting: I know you may be feeling a little sad right now, but please believe that this is only temporary. You are strong and will definitely get through this.

[0067] Emotional copywriting template 3: Corresponding song emotional value 0.5-0.8:

[0068] If the song is judged to be light and cheerful, then the copywriting should use warm language.

[0069] Example of copywriting:

[0070] Seeing your smile makes me incredibly happy. Keep up this good mood and let happiness become a constant in your life!

[0071] Emotional copywriting template 4: Corresponding song emotional value 0.8-1.0:

[0072] Judging from the song's upbeat tone, the copywriting style leans towards a cheerful tone.

[0073] Example of copywriting:

[0074] Your happiness is infectious; seeing you so joyful makes me feel truly happy too. Let's celebrate this joy together and enjoy this wonderful time.

[0075] In some embodiments, after matching the corresponding emotional text template among a plurality of pre-generated emotional text templates based on the emotional value of the song, the method further includes: combining the emotional text template with the on-demand user information of the song to be played, or the location information of the location where the song to be played is located, or a greeting, to form an updated emotional text template.

[0076] The step of generating the corresponding singer's voice in the audio model based on the original singer's timbre and the matched emotional text template is as follows: generating the corresponding singer's voice in the audio model based on the original singer's timbre and the updated emotional text template.

[0077] The above-mentioned method generates corresponding singer voice based on the original singer of a song and the corresponding emotional copywriting template, randomly matching the singer's audio model. It obtains and adjusts the singer's voice emotion based on the emotional value to generate the singer's copywriting voice. The singer's audio model is a pre-trained model. Training audio data can include audio data of the singer speaking normally, such as interviews or live performances, ensuring the audio sources cover different tones and emotional expressions to improve the model's generalization ability. The collected training audio data is preprocessed (including format conversion, noise reduction and denoising, audio segmentation, feature extraction, etc.), and deep learning models such as GPT-Sovits and Sovits-SVC are selected for training. The preprocessed training audio data is used as input, and the model output is calculated through forward propagation. Then, the difference between the output and the true label is calculated using a loss function, and the model parameters are updated through backpropagation. During training, hyperparameters such as learning rate, batch size, and number of iterations are continuously adjusted to optimize model performance. This ensures that the singer's audio model generated after training can generate the corresponding singer's voice when inputting copywriting. Next, adjust the singer's voice emotion based on the song's emotional value to generate the final singer's script voiceover. The singer's voice emotion can be adjusted through speech rate, pitch, and volume. For example, a sadder voice can be adjusted by lowering the pitch, volume, and speech rate. Conversely, a cheerful mood can be adjusted accordingly. For instance, with an emotional value of 0-0.2, reduce speech rate by 4%, pitch by 2%, and volume by 2%; with an emotional value of 0.2-0.5, reduce speech rate by 2%, pitch by 1%, and volume by 2%; with an emotional value of 0.5-0.8, no adjustment is needed; and with an emotional value of 0.8-1.0, increase speech rate by 2%, pitch by 1%, and volume by 1%.

[0078] The aforementioned method of marking the time period for playing the singer's text / audio based on the song's anchor point information specifically involves: obtaining the song's anchor point information through automatic recognition of vocals; obtaining the song's waiting performance time period based on the anchor point information; and marking one of these waiting performance time periods as the time period for playing the singer's text / audio. The aforementioned waiting performance time period refers to the part of the original song without vocals, such as the intro, interlude, and outro. This can be achieved using online or local audio editing software, or by using a trained deep learning model. The anchor point information defines the time points in the song's duration where vocals are present, thus reversing this process to obtain the time points in the song's duration without vocals, and defining the time periods without vocals as the waiting performance time period. One of these waiting performance time periods is selected as the time period for playing the singer's text / audio and marked, so that when the song reaches the marked time period, the singer's text / audio is played. If the selected waiting performance time period is shorter than the singer's text / audio duration, no marking is required to avoid the singer's text / audio interfering with the singer's performance. Specifically, marking the time period for playing the singer's text message based on the song's anchor point information includes determining whether the duration of the time period for playing the singer's text message is greater than the duration of the singer's text message; if it is greater, then the singer's text message is played; otherwise, it is not played.

[0079] Before generating a digital avatar simulating the singer's image based on the original singer's facial features, the process includes a step of generating the original singer's facial features, specifically:

[0080] Use OpenCV to load photos of the original singer of a song and convert them to grayscale;

[0081] The face detector detects faces in grayscale images and obtains the coordinates of facial feature information.

[0082] The facial features are classified based on the coordinates of the detected facial feature information, and the facial feature information of the original singer of the song is generated.

[0083] The original singer photos mentioned above can be obtained from a local song library, online song information searches, or search engines, such as using "singer's name" + "image" as search keywords. High-resolution, clear, and unobstructed frontal and / or side-view photos of the singer are preferred to better extract facial feature information, resulting in a more realistic digital avatar. The facial feature information coordinates include multiple points for the eyes, nose, mouth, and contours, used to classify facial features and generate the singer's facial feature information, including face shape, lip shape, eye shape, nose shape, eyebrow shape, hairstyle, and facial accessories. For example, if the contour coordinates are close to a circle, it is defined as a round face, and the face shape in the facial feature information is round.

[0084] The above-mentioned method generates a digital avatar simulating the singer's image based on the original singer's facial features, and matches the digital avatar's facial animation to the song's emotional value. A digital avatar image feature library can be pre-established, mainly including the digital avatar image and facial feature information applied to it, including face shape, lip shape, eye shape, nose shape, eyebrow shape, hairstyle, facial accessories, and facial expression animation. The image feature library can obtain facial feature resource files and corresponding feature information from online 3D model libraries (such as Unity's FBX model resource files). For example, face shapes include: round face, square face, oval face; lip shapes include: wide mouth, narrow lips; eye shapes include standard eyes, long and narrow eyes, round eyes (mainly defined by eye height), etc.; and facial expression animations are pre-made for each emotional value range (each emotional text template).

[0085] For example, four different animated expressions can be created in advance for the four different scores, including:

[0086] Emotional caption template 1: Corresponds to an emotional value of 0-0.2, with a sad expression;

[0087] Emotional copywriting template 2: Corresponding to an emotional value of 0.2-0.5, with a slightly sad expression;

[0088] Emotional copywriting template 3: Corresponds to an emotional value of 0.5-0.8, with a slightly cheerful expression;

[0089] Emotional copywriting template 4: Corresponds to an emotional value of 0.8-1.0, with a cheerful emoji.

[0090] The aforementioned system, which drives the digital human to play the singer's text during a designated time period, simultaneously lowers the song's volume while displaying the emotional text on the screen with animated effects. Once the text playback is complete, the song volume returns to normal. This automatic adjustment of the song volume during text playback avoids conflict between the text and background music, ensuring users can clearly hear the content without manual volume adjustment, thus improving ease of use.

[0091] This invention also provides an electronic device including a memory and a processor. The memory stores a computer program, which, when executed by the processor, implements the steps of the aforementioned method. The memory includes, but is not limited to, magnetic disks, magnetic tapes, magnetic cards, floppy disks, flash memory, optical disks, optical cards, read-only memory (ROM), random access memory (RAM), erasable programmable ROM (EPROM), and electrically erasable programmable ROM (EEPROM), etc. It also includes other biological, physical, or chemical structures capable of performing similar or equivalent functions to the aforementioned memories, such as DNA, RNA, proteins, etc., which possess information storage capabilities. In specific embodiments, the memory may be one of the aforementioned memory types or a combination thereof. In different embodiments, the computer program involved in the embodiments may be centrally stored in a single memory or distributed across multiple memories. The memory containing a computer device readable storage medium may be non-volatile memory or random access memory. These computer device readable memories may be built into the device or connected to the device involved in the embodiments as an external device or part of an external device. In some embodiments, the memory is deployed locally; in other embodiments, the memory may be deployed remotely from the processor, such as via RF circuitry or external ports and a communication network, which may be the Internet, one or more intranets, local area networks (LANs), wide area networks (WLANs), storage area networks (SANs), or suitable combinations thereof, as long as it enables computer devices to access the memory. Furthermore, the computer programs involved in the embodiments may be stored in plaintext / ciphertext form, or may be designed as training data, integrated and recombined through model training, and implicitly stored in the parameter states of deep neural networks or other machine learning models.The processor described in the embodiments of this application can be implemented by hardware, firmware, software, or a combination thereof. It can be a circuit, one or more of an application-specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field-programmable gate array (FPGA), a central processing unit (CPU), a controller, a microcontroller, or a microprocessor. It also includes other physical, biological, or chemical structures that can implement the same or equivalent functions as the processors listed above, such as biological neurons, quantum computing units, DNA computing units, etc., so that the processor can execute some or all of the steps in the computer program or method involved in the various embodiments of this application, or any combination of the steps mentioned therein.

[0092] It should be noted that although the above embodiments have been described herein, this does not limit the scope of patent protection of the present invention. Therefore, any changes and modifications made to the embodiments described herein based on the innovative concept of the present invention, or equivalent structural or procedural transformations made using the content of the present invention's specification and drawings, directly or indirectly applying the above technical solutions to other related technical fields, are all included within the scope of patent protection of the present invention.

Claims

1. A method for applying digital humans, characterized in that, Includes the following steps: Calculate the emotional value of the song to be played; Match the corresponding emotional copy template from multiple pre-generated emotional copy templates based on the song's emotional value; Based on the original singer's timbre and the matched emotional text template, the corresponding singer's voice is generated in the audio model, and the emotion of the singer's voice is adjusted according to the emotional value to generate the singer's text voice. The time period for playing the singer's lyrics is marked based on the song's anchor point information; A digital human simulating the singer's image is generated based on the singer's facial features, and the digital human's facial animation is matched with the song's emotional value. The digital human is driven to play the singer's text during the marked time period, while the digital human's face displays facial animations corresponding to the emotions.

2. The digital human application method according to claim 1, characterized in that, In the step of calculating the emotional value of the song to be played, the emotional value of the song is calculated using a music emotional value model or a lyrics emotional value model.

3. The digital human application method according to claim 2, characterized in that, The lyrics sentiment value model includes a text preprocessing module, a text feature extraction module, a sentiment analysis module, and a sentiment scoring module; The specific steps for calculating the emotional value of a song using the lyrics emotional value model are as follows: The text preprocessing module performs word segmentation and stop word removal on the song lyrics and outputs the preprocessed text. The text feature extraction module extracts the feature words from the text output by the text preprocessing module. The sentiment analysis module analyzes the sentiment value corresponding to each feature word. The emotion scoring module calculates the sum of the emotion values ​​corresponding to all feature words, which is then used as the emotion value of the song.

4. The digital human application method according to claim 1 or 3, characterized in that, The step of matching the corresponding emotional copy template among multiple pre-generated emotional copy templates based on the emotional value of the song specifically includes: each pre-generated emotional copy template is associated with an emotional value range, the emotional value range into which the emotional value of the song falls is determined, and a matching relationship is established between the song and the emotional copy template associated with the corresponding emotional value range.

5. The digital human application method according to claim 1, characterized in that, After matching the corresponding emotional copy template among the pre-generated multiple emotional copy templates based on the emotional value of the song, the method further includes: combining the emotional copy template with the user information of the song to be played, or the location information of the place where the song to be played is located, or a greeting, to form an updated emotional copy template. The step of generating the corresponding singer's voice in the audio model based on the original singer's timbre and the matched emotional text template is as follows: generating the corresponding singer's voice in the audio model based on the original singer's timbre and the updated emotional text template.

6. The digital human application method according to claim 1, characterized in that, The specific steps for marking the time period for playing the singer's text voice based on the anchor information of the song are as follows: the anchor information of the song is obtained by automatically recognizing the human voice in the song, the waiting time period of the song is obtained based on the anchor information of the song, and one of the waiting time periods is marked as the time period for playing the singer's text voice.

7. The digital human application method according to claim 1, characterized in that, The step of marking the time period for playing the singer's text voice based on the anchor point information of the song includes determining whether the duration of the time period for playing the singer's text voice is greater than the duration of the singer's text voice; if it is greater, then the singer's text voice is played; otherwise, it is not played.

8. The digital human application method according to claim 1, characterized in that, Before generating a digital avatar simulating the singer's image based on the original singer's facial features, the process includes a step of generating the original singer's facial features, specifically: Use OpenCV to load photos of the original singer of a song and convert them to grayscale; The face detector detects faces in grayscale images and obtains the coordinates of facial feature information. The facial features are classified based on the coordinates of the detected facial feature information, and the facial feature information of the original singer of the song is generated.

9. The digital human application method according to claim 1, characterized in that, The process of adjusting the singer's voice based on the emotional value includes: the total emotional value of the song is 1; when the emotional value of the song is below 0.5, it is determined to be a sad song, and at least one of the following singer voice parameters is reduced: speech rate, pitch, or volume; when the emotional value of the song is above 0.8, it is determined to be a cheerful song, and at least one of the following singer voice parameters is increased: speech rate, pitch, or volume.

10. The digital human application method according to claim 1, characterized in that, When the digital human is driven to play the singer's text during the designated time period, the song playback volume is simultaneously reduced, and the emotional text is displayed on the screen in the form of animation effects. When the singer's text playback is complete, the song playback volume is restored.

11. An electronic device, characterized in that: It includes a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, it implements the steps of the digital human application method as described in any one of claims 1-10.

Citation Information

Patent Citations

  • Three-dimensional virtual character construction method and system on the basis of human face simulation

    CN105139450A

  • Interaction method based on audio and virtual image and terminal

    CN112402952A