Method for synchronising lip movements and acoustic output of content by an avatar
By adapting the pronunciation of content to a predefined lip movement sequence, the method achieves fast and synchronized lip movements and acoustic output, addressing the inefficiencies of conventional methods and enabling uninterrupted communication between users and avatars.
Patent Information
- Application Number
- PCT/EP2024/078544
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-28
- Filing Date
- 2024-10-10
- Publication Date
- 2025-06-05
AI Technical Summary
Conventional methods for synchronizing lip movements and acoustic output of content by an avatar are time-consuming and require high computing power, making them unsuitable for uninterrupted communication.
Adapting the pronunciation of generated content to a predefined or stored lip movement sequence, allowing for fast and timely acoustic output with synchronous lip movements, thereby reducing processing time to less than 1 second.
Enables uninterrupted and fluid communication between a user and the avatar by synchronizing lip movements and acoustic output quickly and efficiently, eliminating waiting times between user requests and avatar responses.
Smart Images

Figure EP2024078544_05062025_PF_FP_ABST
Abstract
Description
[0001] Method for synchronizing lip movements and acoustic output of a content by an avatar
[0002] The present invention relates to a method for synchronizing lip movements and acoustic output of a content by an avatar according to the preamble of claim 1, a data processing device, a computer program and a computer-readable data carrier.
[0003] In conventional methods, content to be spoken is first generated, and the pronunciation of the content is calculated, simulated, and / or created. Subsequently, the avatar's lip movements can be calculated based on the pronunciation of the content.
[0004] For example, DE 10 2007 042 583 B4 discloses a method for communication between a natural person and an artificial language system. First, a content to be spoken is generated. The correct phonetic pronunciation of words is stored in a language system. Based on the correct phonetic pronunciation of the words, the avatar's lip movements can be adapted to the correct pronunciation of the words.
[0005] Adapting lip movements to the correct pronunciation of words is generally very time-consuming and requires high computing power from a computer or data processing device. Adapting lip movements to the content to be pronounced typically requires a long processing time, up to several minutes. Such methods are therefore not suitable for providing uninterrupted communication between a user and the avatar.
[0006] The object of the present invention is to provide a method for synchronizing lip movements and acoustic output of content by an avatar that is improved compared to the prior art, wherein a fast and timely acoustic output of content with synchronous visually output lip movements by the avatar is enabled or supported, wherein particularly simple, user-friendly and / or intuitive communication is created or supported. The object underlying the present invention is achieved by the method according to claim 1, the data processing device according to claim 21, the computer program according to claim 22 or the computer-readable data carrier according to claim 23.
[0007] The present invention relates to a method, in particular a computer-implemented method, for synchronizing lip movements and acoustic output of a content by an avatar.
[0008] In the context of the present invention, the term “avatar” refers to digital beings with an anthropomorphic appearance that are controlled by humans or software and have the ability to interact.
[0009] Preferably, the proposed method, in particular individual or all method steps of the proposed method, is or are carried out (partially) automatically or automatically by means of a data processing device, in particular by means of corresponding means for data processing and controlling the device, such as a data processing device or the like.
[0010] In the proposed method for synchronizing lip movements and acoustic output of a content by the avatar, a content to be spoken is first generated, which is to be acoustically output by the avatar.
[0011] The proposed method is characterized by adapting the pronunciation of the previously generated content to a predefined or stored lip movement sequence of the avatar. It has been shown that adapting the pronunciation to a predefined lip movement sequence can be done faster and with less computing power than the conventional prior art adaptation of lip movement to the phonetic pronunciation of the content. It is then possible for the avatar to output generated content with synchronous lip movements in response to a user request within a particularly short time, in particular less than 1 second, preferably less than 0.5 seconds. In this way, uninterrupted and fluid communication can take place between the user and the avatar.In particular, there are no waiting times between a user's request and the avatar's response, which enables at least essentially uninterrupted and / or direct communication with the avatar.
[0012] The invention is based on the finding that only approximately 30% of all words can be recognized based on the lip movements that occur when pronouncing them. Furthermore, only approximately 15 (speech) sounds or phonemes can be unequivocally assigned to a defined lip position or viseme.
[0013] For the purposes of the present invention, the term "phoneme" refers to the smallest linguistic unit that distinguishes meaning, or the abstract class of all sounds that have the same meaning-distinguishing function in a spoken language. For the purposes of the present invention, the term "viseme" refers to the mouth position and / or lip position and / or lip movement during the pronunciation of a phoneme or sound.
[0014] It is therefore possible to assign a multitude of sounds (phonemes) to a single lip position (viseme). At the same time, multiple words can be assigned to an identical lip movement without the user noticing a lack of synchronization between spoken content and the avatar's lip movement.
[0015] To adapt the pronunciation of the content to the given lip movement sequence, the accentuation of individual sounds, syllables, letters, and / or words in the content can be adjusted to the given lip movement sequence. For example, a syllable of a word can be emphasized particularly strongly to synchronize with a passage or section of the lip movement sequence. In particular, it is possible to adjust the accentuation of individual phonemes so that they match a viseme of the lip movement sequence.
[0016] It is possible to change or adapt the pronunciation of individual sounds or phonemes. For example, the letter “R” can be rolled to varying degrees during pronunciation. To adapt to the predefined lip movement sequence, the pronunciation of the letter “R”, in particular how strongly the letter is rolled, can be varied in order to achieve optimal agreement with a viseme or lip position and / or lip movement. Alternatively or additionally, the output speed of at least one sound, one syllable, one letter and / or one word can be adapted to the predefined lip movement sequence, at least in certain passages. For example, a first syllable of a word can be output faster than the second syllable of the word in order to achieve a synchronous pronunciation with a passage of the lip movement sequence.
[0017] When the output speed of tones, sounds, and / or spoken text changes, the pitch or frequency of the output is usually also changed. To prevent a change in the pitch or frequency during the output, a pitch or frequency control is preferably provided, which is designed to keep the pitch or frequency of the output constant regardless of the output speed. In this way, a particularly natural-sounding sound or speech image of the avatar can be achieved.
[0018] To further or better adapt the pronunciation of the content to the given lip movement sequence, at least one pause can be inserted between sounds, syllables, and / or words. For example, a (short) pause can be inserted between syllables of a word or between two consecutive words to align the pronunciation with a passage of the lip movement sequence.
[0019] Alternatively or additionally, a pause between sounds, words, and / or syllables can be adapted to the given lip movement sequence. For example, a pause between consecutive words of the content can be shortened or lengthened as needed to maintain synchronization between the corresponding passage of the lip movement sequence and the pronunciation of the words.
[0020] Furthermore, it is preferably provided that defined sounds, syllables, letters and / or words are assigned to a defined lip movement and / or defined lip position or a viseme. In this way, predetermined components of the content to be pronounced are assigned to defined lip movements and / or defined lip positions in order to create correspondences between the pronunciation of the content and the assigned lip movements or lip positions. Particularly preferably, the output speed, emphasis and / or pauses between the defined sounds, syllables, letters and / or words are adapted to the predetermined lip movement sequence. Thus, the output speed, emphasis and / or pauses are adapted, in particular, of components of the content that are located between the sounds, syllables, letters and / or words previously assigned to a predetermined lip movement or lip position.In this way, individual lip positions and / or lip movements can first be assigned a sound, syllable, letter, and / or word based on appropriate assignment criteria. Subsequently and / or simultaneously, the intermediate components can then be adapted to the lip movement sequence.
[0021] To increase the synchronization between the pronunciation of the content and the predefined lip movement sequence, the predefined lip movement sequence can be selected from a plurality of predefined lip movement sequences. It is then possible to select a suitable or most suitable lip movement sequence, as described in detail below.
[0022] In particular, the selection of the predefined lip movement sequence can be made depending on the length and / or number of words and / or the number of vowels. For example, by analyzing the content to be pronounced, the number of vowels in the content can be determined. Based on the aforementioned analysis, a most suitable lip movement sequence can then be selected from the plurality of lip movement sequences.
[0023] Alternatively or additionally, the selection of the predefined lip movement sequence can also be dependent on the language of the content. The term "language" in the context of the present invention is to be understood as a human (national) language and / or a dialect or the like. For example, a lip movement sequence may be more suitable for synchronization with content in one language than for synchronization with content in another language. It is then possible for the avatar to output or pronounce content in a variety of languages. The method can then be used for many different languages and / or when switching languages within a communication. The content to be pronounced is generated within the framework of the method according to the invention. In a preferred embodiment, the content is selected and / or generated from a plurality of predefined or stored content modules.It is then possible to use only content-controlled or approved content modules in order to avoid the avatar outputting objectively false information.
[0024] The number of predefined content blocks is preferably greater than the number of predefined lip movement sequences. Since the pronunciation of the content is adapted to a predefined lip movement sequence, multiple content blocks can be synchronized with the same lip movement sequence. This allows a significantly higher number of content blocks to be synchronized with a selected number of predefined lip movement sequences.
[0025] The number of predefined content modules is preferably more than twice as large, further preferably more than four times as large, further preferably more than ten times as large, further preferably more than one hundred times as large, further preferably more than one thousand times as large as the number of predefined lip movement sequences.
[0026] The predetermined lip movement sequences are preferably divided into at least two categories or groups. Here, and preferably, the predetermined lip movement sequences are divided into a group of short sequences and a group of long sequences. In particular, the short sequences have a length of 1.5 s to 5 s, preferably from 1.5 s to 4 s, more preferably from 2 s to 3 s. The long sequences can have a length of 5 s to 15 s, preferably from 5 s to 13 s, preferably from 6 to 12 s.
[0027] It is also possible to provide more than two groups for classifying the lip movement sequences.
[0028] The content is preferably an answer and / or reaction to an input content and / or part of an answer and / or reaction to an input content or comprises an answer and / or a reaction and / or part of an answer and / or reaction to an input content. The term “input content” is understood within the scope of the invention to mean information that can be presented acoustically and / or visually and / or is machine-readable. The input content and / or the information can, for example, be spoken or written text, tones, sounds, gestures or the like. Here and preferably, the input content is in particular a question, request, statement and / or request that is articulated acoustically by the user.
[0029] Thus, the method is particularly suitable for outputting a reaction to input content, especially a question entered or asked by a user. Since the content modules are preferably predefined, it can be ensured that only objectively correct content is output as a reaction or answer to input content.
[0030] The user's input content is preferably entered in speech form and converted into input text and / or binary code. The user's input content is then available as written or machine-readable input text or binary code. The input text or binary code can then be easily processed, analyzed, and / or evaluated.
[0031] Alternatively, the input content can be entered in text form. The input text can be converted into binary code. Further processing, analysis, and / or evaluation can then take place.
[0032] The input content is then analyzed for words, expressions, and / or synonyms, particularly using a Large Language Model (LLM). This allows the input content or text to be analyzed, evaluated, and / or logically understood.
[0033] Based on the analysis carried out, a predefined content module or part(s) thereof can be selected. It is also possible to select several predefined content modules. The content to be spoken can then be generated using the selected content modules, for example by stringing together and / or nesting predefined content modules. The predefined content modules are preferably divided into at least two thematic categories. Based on the input content or the analysis of the input content or input text, a category can be selected for selecting a predefined content module from this category. For example, the input text can be assigned to a category of content modules that best matches the question. The best-fitting predefined content module or input text can then be found in the correspondingly selected category.several suitable predefined content modules can be selected.
[0034] For example, assignment criteria can be specified that enable the assignment of input content, components of the input content, analyses of the input content, and / or components of the analyses of the input content to one and / or more categories of predefined content modules. Using the assignment criteria, an input content can thus be assigned to at least one, in particular exactly one, category.
[0035] Furthermore, assignment criteria can be specified in order to assign at least one predefined content module within the selected category to an input content or input text, particularly after selecting a category.
[0036] It is preferably provided that at least one method step is carried out by means of a computer.
[0037] According to a further aspect of the present invention, a data processing device comprising means for implementing the proposed method is proposed. Reference may be made to all statements regarding the proposed method in this regard. In particular, corresponding advantages are achieved.
[0038] According to a further aspect of the present invention, a computer program comprising instructions which, when executed by a computer, cause the computer to execute a proposed method is proposed. Reference may be made to all statements relating to the proposed method in this regard. In particular, corresponding advantages are achieved. According to a further aspect of the present invention, a data carrier, in particular a non-volatile, computer-readable data carrier, on which the proposed computer program is stored is proposed. Reference may be made to all statements relating to the proposed computer program in this regard. In particular, corresponding advantages are achieved.
[0039] The aforementioned aspects, features and method steps as well as the aspects, features and method steps of the present invention resulting from the claims and the following description can in principle be implemented independently of one another, but also in any desired combination and / or sequence.
[0040] Further aspects, advantages, features, properties, and advantageous developments of the present invention will become apparent from the claims and the following description of preferred embodiments with reference to the figures. They show, in a schematic representation, not to scale:
[0041] Fig. 1 is a schematic view of a proposed data processing device designed to synchronize lip movements and the acoustic output of a content by an avatar, and
[0042] Fig. 2 shows a schematic flow diagram of a proposed method for synchronizing lip movements and acoustic output of a content by an avatar or individual method steps of the proposed method.
[0043] In the figures, some of which are not to scale and are merely schematic, the same reference symbols are used for identical, identical or similar parts and components, whereby corresponding or comparable properties or advantages are achieved, even if repetition is omitted.
[0044] Fig. 1 shows a schematic view of a data processing device 1, in particular a computer. The data processing device 1 can also be a tablet, mobile phone, or the like. The device 1 preferably has a visual output device, in particular a screen 2, for displaying or reproducing an avatar 3. The screen 2 is designed in particular for reproducing moving images, in particular of the avatar 3. By means of the screen 2, movements of the avatar 3, and in particular movements of the avatar 3 while speaking, can be displayed, preferably smoothly.
[0045] Furthermore, the device 1 has an input device 4 for inputting an input content 5. The input device 4 can be an integral part of the device 1 or be designed as a separate input device.
[0046] The input device 4 can, for example, have a keyboard 4A, a mouse 4B, a touchpad / trackpad 4C, a camera 4D, a microphone 4E and / or a touchscreen for entering the input content 5.
[0047] For example, the microphone 4E can be integrated into the camera 4D. Alternatively or additionally, it is possible for the screen 2 to be designed as a touchscreen or for a touchscreen to be integrated into the screen 2.
[0048] A user 6 can enter input content 5 (Fig. 2) via the input device 4. The input content 5 can in particular be an acoustic input, in particular via the microphone 4E. For example, the user 6 can ask a question acoustically or acoustically request certain information. Alternatively or additionally, the input content 5 can also be entered via the keyboard 4A, the mouse 4B, the touchpad / trackpad 4C, the camera 4D and / or the touchscreen. A combined acoustic and motor input via the microphone 4E and keyboard 4A, mouse 4B, touchpad 4C and / or touchscreen is also conceivable. Here and preferably, the input content 5 is a text, in particular a spoken text.
[0049] The device 1, in particular a data processing device 7 of the device 1, can analyze the input content 5 and, based on the analysis, generate a content 8 to be pronounced, in particular text. Within the context of the present invention, a "data processing device" is understood to mean a device for automatically processing data, in particular program code. The data processing device 7 can, in particular, have a processor, memory, an internal and / or external database, and / or a communication device for connecting to the Internet and / or other computers.
[0050] In the context of the present invention, the term "content" refers to information that can be presented acoustically and / or visually. For example, the content 8 can be spoken text, sounds, noises, or the like.
[0051] The device 1 preferably further comprises an acoustic output device 9 for acoustically outputting the content 8. The acoustic output device 9 may, for example, be a loudspeaker that is an integral component of the device 1 or is formed as a separate part / device of the device 1 and / or is combined with the input device 4.
[0052] The previously generated content 8 can be acoustically output by the avatar 3 via the output device 9. At the same time, the visual output can be provided by the avatar 3 via the screen 2. In this way, communication between the user 6 and the avatar 3 can be enabled.
[0053] To support the acoustic output of the content 8, the device 1 is preferably designed to synchronize the lip movements of the avatar 3 and the acoustic output of the content 8. Within the context of the present invention, the term "lip movement" refers to the lip movement and preferably the facial expression and / or facial expression of the avatar 3 when pronouncing the content 8.
[0054] The user 6 thus hears the acoustic output by the avatar 3 and sees on the screen 2 a lip movement of the avatar 3 that is synchronous with the acoustic output. In this way, the understanding of the user 6 can be increased and the communication between the user 6 and the avatar 3 can be improved.
[0055] In particular, it is provided that the pronunciation of the content 9 is adapted to a predetermined lip movement sequence 14 of the avatar 3. Within the context of the present invention, the term "lip movement sequence" refers to a lip movement lasting over a predetermined period of time. Fig. 2 shows a schematic flow diagram of a proposed method for synchronizing lip movements and acoustic output of a content 9 by the avatar 3 or individual method steps of the proposed method.
[0056] The method is preferably multi-stage or multi-step. In particular, the method comprises several process steps, whereby the individual process steps can in principle be carried out independently of one another and in any order, unless otherwise stated below.
[0057] The proposed method is preferably carried out by means of the device 1 for data processing, in particular by means of the input device 4, the screen 2 and / or the output device 8.
[0058] The data processing device 1 is preferably designed to carry out the method described herein or individual or all method steps.
[0059] Preferably, the instructions or the algorithm for executing the proposed method or individual method steps of the proposed method are stored electronically in a (data) memory of the device 1, in particular the data processing device 7.
[0060] However, it is also possible that one or more process steps are carried out by means of an (external) device or an (external) device and / or that individual or more commands for executing the process or individual process steps are stored there.
[0061] The proposed method preferably includes one or more method steps and / or a program to synchronize lip movements and the acoustic output of a content 9 by the avatar 3.
[0062] The proposed method is characterized in that the pronunciation of the content 9 is adapted to a predetermined lip movement sequence 14 of the avatar 3. In this way, a particularly fast and short-term synchronization between the lip movements and the acoustic output of a content 9 by the avatar 3 can be achieved. It is then possible to perform an acoustic output with synchronous lip movements by the avatar 3 without a significant time delay and with particularly little computing power.
[0063] The invention is based on the finding that only 30% of all words can be recognized based on the lip movements that occur when pronouncing these words. Furthermore, only 15 (speech) sounds (phonemes) can be unequivocally assigned to a defined lip position (viseme). It is therefore possible to assign a large number of sounds (phonemes) to a lip position (viseme). At the same time, a large number of words can also be assigned to a single lip movement without the user being able to detect a lack of synchronization between a spoken content 9 and the lip movement.
[0064] It is then possible, in particular, to adapt the pronunciation of different content 9 to an identical lip movement sequence 14 of the avatar 3, as will be described in more detail below. In this way, a significantly larger number of content 9 can be output with a limited number of lip movement sequences 14, with the spoken content 9 being output synchronously with the lip movement sequences.
[0065] The method is preferably initiated by the user 6, for example, by opening a corresponding program and / or a corresponding website and / or entering a corresponding command. For example, the user 6 can open a corresponding program or enter a corresponding command using the input device 4, preferably in a first method step / process A1.
[0066] The first procedural step / process A1 can include or include a greeting by avatar 3, in particular a greeting adapted to the respective situation. For example, the greeting by avatar 3 can be dependent on the time of day and / or year. Other factors influencing the greeting are also possible.
[0067] It is also possible for the method to be initiated by the user 6 pronouncing a predefined command, for example, by pronouncing "Hello Avatar." Optionally, in a further / second method step A2, the user 6 can enter an input content 5, in particular as spoken text, via the input device 4. The input content 5 can be a question, a request for information, and / or a statement. Here, and preferably, the input content 5 is a question.
[0068] The second method step A2 preferably comprises the detection of the input content 5, in particular acoustic input content. Detection is carried out in particular by means of the input device 4.
[0069] During the input of the input content 5, preprocessing can be performed to improve the quality of the capture and / or increase the recognition accuracy of the input content 5. For example, background noise and interference can be eliminated. Alternatively or additionally, certain frequencies can be filtered to remove irrelevant frequencies. Additionally or alternatively, it is also possible to standardize the volume level of the sound. It is also possible to perform only some or all of the aforementioned preprocessing steps.
[0070] The input content 5 is preferably recorded and / or stored in order to perform processing as described below.
[0071] In a further, particularly parallel, third method step A3, the avatar 3 is preferably displayed on the screen 2. The avatar 3 is preferably shown in a moving state, wherein a naturally listening movement of the avatar 3 is depicted. The term "naturally listening movement" in the context of the present invention refers to a movement of a natural person listening to a user 6. In this way, the avatar 3 imitates a natural action or a natural movement of a real person. In particular, the avatar 3 does not perform any speaking or speech-like lip movements during the third method step A3.
[0072] The movement of the avatar 3 can be adapted to the language used by the user 6. Thus, the movement performed by the avatar 3 can be different for one language used by the user 6 than for another language used by the user 6. In this way, linguistic and / or cultural characteristics or differences can be imitated by the avatar 3. In this way, communication and / or understanding by the user 6 can be improved overall. At the same time, an atmosphere perceived as pleasant for the user 6 can be created.
[0073] It is also possible to adapt the movements performed by the Avatar 3 depending on other influencing factors, such as the time of day and / or year.
[0074] In a further / fourth method step A4, the input or input content 5 of the user 6 is preferably analyzed.
[0075] Thus, after the input or input content 5 has been captured and / or after the input or input content 5 has been saved (second method step A2) - in addition to or alternatively to the preprocessing in method step A2 - processing can be carried out to improve the quality of the captured input content 5 and to increase the recognition accuracy of the captured input content 5. For example, background noise and interference can be eliminated. Alternatively or additionally, certain frequencies can be filtered to eliminate irrelevant frequencies. It is also possible, additionally or alternatively, to standardize the volume level of the sound. It is also possible to perform only some or all of the aforementioned processing steps, for example, sequentially, in parallel, and / or iteratively.
[0076] Furthermore, in the fourth method step A4, phonetic matches and patterns can be searched for. The matches and patterns can then be compared with patterns stored in a database to identify at least individual components of the input content 5. The database can be an internal and / or external database, in particular a cloud-based database. The database can be embodied as a memory, in particular a hard disk. The database can be a hard disk component of the device 1 or a component of another computer.
[0077] In the fourth method step A4, the input content 5, in particular spoken input content, can be converted into text and / or code, in particular machine-readable text. The conversion can be performed using a voice-to-text application. The voice-to-text application can, in particular, be a speech recognition system that transcribes spoken text. It is also possible for the fourth method step A4 to be designed to support speech recognition and / or improve the accuracy and adaptability of processing using machine learning and / or artificial intelligence.
[0078] Preferably, in a further / fifth method step A5, the logical content or meaning of the particularly transcribed input content 5 or the input text is analyzed. A Large Language Module (LLM) is preferably used for this purpose. The term "Large Language Module" refers to a generative language model for texts or text-based content. Large Language Modules use artificial intelligence, artificial neural networks, and / or deep learning to process, understand, and / or, if necessary, generate natural language. Corresponding language models are used to understand complex texts, questions, and / or instructions.
[0079] Here, and preferably, the input content 5 or the input text is analyzed using the Large Language Model based on expressions, synonyms, words, sentence components, and / or sentences. With the help of the Large Language Model, the logical content or meaning of the input content 5 of the user 6 can thus be captured or understood.
[0080] Subsequently, a reaction or response to the input content 5 of the user 6 is preferably generated. For this purpose, the device 1 or the program or the external database has several stored predefined content modules 10. The predefined content modules 10 can be answers and / or information on various topics.
[0081] The content modules 10 can be information or reactions that have been verified, for example, by humans and are therefore objectively correct. In this way, only information or reactions based on objectively correct content modules 10 can be output. This prevents fictitious and / or objectively incorrect information or reactions from being output by the avatar 3 or the device 1.
[0082] Preferably, the predefined content modules 10 are divided into at least two categories 11. Figure 2 shows, by way of example, three categories 11 with predefined content modules 10. However, it is also possible to provide more than three different predefined categories. In particular, each category has several, in particular more than two, content modules 10.
[0083] The predefined content modules 10 can be divided into categories 11 based on topic. Thus, one category 11 can contain predefined content modules 10 related to one topic, while another category 11 can contain predefined content modules 10 related to a different topic.
[0084] It is also possible that a content module 10 is assigned to two categories 11 or is contained in two different categories 11.
[0085] In the sixth method step A6, one of the predefined categories 11 is preferably selected, as shown in Fig. 2 by the dashed line 12. In particular, the category 11 is selected based on the analysis of the input content 5 or the analysis of the input text according to the fifth method step A5. The category 11 can preferably be based on the analysis by the Large Language Model. Thus, to answer the input of the user 6 or in response to the input content 5, preferably only content modules 10 from the correspondingly selected category 11 are used.
[0086] Subsequently, in a further / seventh method step A7, a predefined content module 10 or several predefined content modules 10 can be selected or identified from the selected category 11, as shown in Fig. 2 by the dashed line 13. The selection of the predefined content module 10 can preferably be made based on the analysis of the input content 5 or based on the analysis by the Large Language Model according to the fifth method step A5.
[0087] The content modules 10 and the categories 11 can be stored or saved in a database. The database can be designed in particular as a memory chip and / or hard disk. The database can be part of the device 1 or as an internal database. It is also possible for the database to be an external database. The device 1 can then access the database, for example, via the Internet. In order to be able to answer complex questions or to respond adequately to complex input content 5, method steps A6 and A7 can be carried out iteratively, as shown in Fig. 2. In this way, predefined content modules 10 from different categories 11 can also be used, for example if the input content 5 contains multiple questions or if access to multiple categories 11 is required to answer the input content 5 or as a reaction to the input content 5.
[0088] In a further / eighth method step A8, the selected predefined content modules 10 or the selected predefined content module 10 can be processed into a content 9 to be spoken, or the content 9 to be spoken is formulated from the selected predefined content module 10 or the selected predefined content modules 10. The content 9 to be spoken can be a response to the input content 5 or a reaction to the input content 5. Alternatively or additionally, the content 9 to be spoken can comprise part of a response to the input content 5 or part of a reaction to the input content 5.
[0089] If several predefined content modules 10 are selected, these can be further processed individually or combined or arranged in a sequence adapted to the input content 5 to form a logical information unit in order to formulate the content 9 to be spoken.
[0090] Subsequently, in a further method step, here preferably in method steps A9 and A10, a predetermined lip movement sequence 14 can be selected from a plurality of predetermined lip movement sequences 14.
[0091] Here and preferably, the predetermined lip movement sequences 14 are divided into at least two groups 15 of lip movement sequences 14.
[0092] Preferably, the predetermined lip movement sequences 14 are divided into a group of short sequences 14A and a group of long sequences 14B. The short sequences 14A can, in particular, represent lip movements of half-sentences, subordinate clauses, sentence segments, or short sentences. Here, and preferably, the short sequences 14A have a length of 1.5 s to 5 s, preferably 1.5 s to 4 s, more preferably 2 s to 3 s.
[0093] The long sequences 14B can, in particular, be the representation of lip movements of entire sentences and / or main clauses. In particular, the long sequences 14B can have a length of 5 s to 15 s, preferably 5 s to 13 s, more preferably 6 s to 12 s.
[0094] Thus, two groups 15 of lip movement sequences 14 are preferably formed here, with the lip movement sequences 14 being divided into one of the two groups 15 based on their duration. It is also possible to form or provide more than two groups 15.
[0095] In the ninth method step A9, a group 15 of lip movement sequences 14 is preferably selected, as shown in Fig. 2 by the dashed line 16.
[0096] In particular, it is provided that the selection of the group 15 of predefined lip movement sequences 14 is dependent on the length and / or number of words, the number of vowels, and / or the language of the content 9 to be pronounced. For example, a different group 15 can be selected for a sentence to be pronounced with a few words than for a sentence to be pronounced with a larger number of words. Alternatively, for two sentences with the same number of words, a different group 15 can be selected based on the different number of vowels.
[0097] It is also possible that a different group 15 is selected for a complete sentence or a main clause than for a subordinate clause.
[0098] After selecting the group 15, at least one predetermined lip movement sequence 14 can be selected in a further / tenth method step A10 depending on the length and / or number of words, the number of vowels and / or the language of the content 9 to be pronounced, as shown by the dashed line 17 in Fig. 2.
[0099] For example, a different lip movement sequence 14 can be selected for a sentence to be pronounced with a few words than for a sentence to be pronounced with a larger number of words. Alternatively, for two sentences with the same number of words, a different lip movement sequence 14 can be selected based on a different number of vowels.
[0100] For complex and / or long content 9 to be pronounced, method steps A9 and A10 can be run iteratively to select multiple lip movement sequences 14 if necessary. The selected lip movement sequences can then be combined into a single lip movement sequence 14.
[0101] It is preferably provided that the number of predefined content modules 10 is greater than the number of predefined lip movement sequences 14. In particular, the number of predefined content modules 10 is more than twice as large, preferably more than four times as large, more preferably more than ten times as large, more preferably more than one hundred times as large, more preferably more than one thousand times as large as the number of predefined lip movement sequences 14. It is thus possible, with a small number of predefined lip movement sequences 14, to output a large number of predefined content modules 10 as spoken content 9 with synchronous lip movements by an avatar 3, as will be explained in more detail below.
[0102] After completion of the tenth method step A10, the content 9 to be pronounced and the selected lip movement sequence 14 are thus preferably in a state that is not synchronized with each other.
[0103] In the subsequent / eleventh method step A11, the acoustic output of the content 9 and the visual output, here and preferably the lip movements of the avatar 3, are synchronized.
[0104] In the prior art, the lip movements of the avatar 3 are adapted to the content 9 to be spoken, which requires long processing times. With current computing power, it can take up to several minutes to simulate lip movements of the avatar 3 to a given text or content 9 to be spoken. Such interaction with the avatar 3 is correspondingly time-consuming and involves long waiting times for the user 6 between entering the input content 5 and the acoustic output of the content 9 to be spoken by the avatar 3. Such systems are perceived as taking too long and not very convenient.
[0105] Here, and preferably, it is provided that the pronunciation of the content 9 is adapted to the predetermined lip movement sequence 14 of the avatar 3. Within the scope of the present invention, a reverse approach is thus pursued to synchronize the lip movements and the pronunciation of the content 9.
[0106] To synchronize lip movements and the acoustic output of the content 9 by the avatar 3, the accentuation of individual sounds, syllables, letters, and / or words of the content 9 can be adapted to the predefined lip movement sequence 14. For example, the ending or a syllable of a word can be emphasized more or less strongly to match a section of the predefined lip movement sequence 14 (viseme).
[0107] Alternatively or additionally, the output speed of at least one sound, one syllable, one letter, and / or one word can be adapted, at least in certain passages, to the predefined lip movement sequence 14. For example, it is possible to adapt the output speed of, for example, a syllable, a sound, and / or individual words to the predefined lip movement sequence 14 so that a synchronous output is achieved between the pronunciation of the content 9 and the predefined lip movement sequence 14.
[0108] If the output speed of tones, spoken words and / or sounds is varied, the pitch or frequency of the output tones, spoken words and / or sounds generally changes depending on the speed difference. Here and preferably, it is provided that the pitch or frequency of the tone, spoken word and / or sound is maintained, even if the output speed changes. If, for example, the output speed of a word is increased, the frequency would also increase accordingly. By adjusting the pitch or frequency, reproduction can then preferably take place at the specified pitch, so that the specified pitch is maintained during accelerated or slowed down output. In this way, a natural emphasis can be achieved when pronouncing the content 9, even at different output speeds.Alternatively or additionally, it is possible to insert a pause between sounds, words, and / or syllables and / or to adapt it to the predefined lip movement sequence 14. For example, a pause can be additionally inserted or extended between two words to ensure synchronous lip movement of the avatar 3 with the two spoken words.
[0109] Alternatively, it is also possible to shorten the pause between consecutively pronounced syllables, sounds, and / or words. For example, consecutively pronounced words can blend into one another during pronunciation to ensure synchronized lip movement of the avatar 3 with the two words being spoken.
[0110] In particular, assignment criteria between sounds (phonemes), syllables, letters, words, and / or word sequences can be specified to lip positions (visemes) and / or passages of lip movements. This makes it possible to assign sounds (phonemes), syllables, letters, words, and / or word sequences to individual lip positions (visemes) and / or passages of lip movements. By adapting the pronunciation of the content 9, a synchronization between the specified lip movement sequences 14 and the pronunciation of the content 9 can thus be achieved.
[0111] In a subsequent, further / twelfth method step A12, the acoustic and visual output of the content 9 and the lip movement sequence 14 is carried out by the avatar 3 or screen 2 and the output device 8. The pronunciation of the content 9 is output synchronously with the predetermined lip movement sequence 14, whereby the understanding of the user 6 can be increased.
[0112] Here and preferably, the avatar 3 can be a predefined, real or fictitious person, for example a real or fictitious pharmacist. The user 6 can, for example, be a patient with a prescription for an individualized medication. Within the scope of the method, the user 6 can then, for example as a patient, ask the avatar 3, as a pharmacist known to the user 6, a question regarding the prescription or the medication or its intake. For this purpose, the user 6 can use the device 1 and speak the question or enter it otherwise using the input device 4 (method step A2). The input content 5 can then optionally be saved and / or edited (method steps A2 and A4). The question of the user 6 is preferably converted into machine-readable text orCode is converted (process step A4) and analyzed using the Large Language Model in order to recognize or analyze the input content 5 (process step 5).
[0113] Subsequently, to answer the question, at least one predefined content module 10 can be selected or processed to generate a content 9 to be pronounced (method steps A6 to A8). Based on the selected content module 10, at least one lip movement sequence 14 can be selected (method steps A9 and A10). The content 9 to be pronounced and the predefined lip movement sequence 14 are then synchronized with each other by adapting the pronunciation of the content 9 to the predefined lip movement sequence 14 (method step A11).
[0114] The aforementioned method steps for processing the input by the user 6 and generating the content 9 to be pronounced as well as for synchronizing the pronunciation of the content 9 and the predetermined lip movement sequence are preferably less than 1.0 s, more preferably less than 0.5 s, more preferably less than 0.4 s, more preferably less than 0.3 s.
[0115] It is then possible for the avatar 3 to output a response to a question from the user 6 within less than 1.0 s, with the lip movements of the avatar 3 being adapted to the pronunciation of the content 9 (method step A12). In this way, a conversation can take place without a pause between the question and the response from the avatar 3, which the user 6 finds annoying.
[0116] During input and / or until the pronunciation of the content 9 is reproduced with synchronous lip movements, the avatar 3 can perform natural listening movements, creating a familiar atmosphere for the user 6.
[0117] Avatar 3 can also be a real or fictitious employee of a company, for example, from the human resources department. User 6 can also be an employee of the same company. Using the method and / or device 1, user 6 can have relevant company-related questions answered in a particularly simple manner and quickly, without hindering or distracting other employees from their work. For example, a question about how to fill out a vacation request or how many vacation days user 6 is still entitled to can be answered particularly quickly and easily. Other use cases or constellations are also possible.
[0118] A further aspect of the present invention relates to a computer program comprising instructions that, when executed by a computer or the device 1 according to the invention, cause the computer to execute the proposed method. Reference may be made to all statements regarding the proposed method in this regard. In particular, corresponding advantages are achieved.
[0119] A further aspect of the present invention relates to a computer-readable data carrier on which the proposed computer program is stored. Reference is made to all statements regarding the proposed computer program in this regard. In particular, corresponding advantages are achieved.
[0120] List of reference symbols:
[0121] 1 Device 9 Contents
[0122] 2 Screen 10 Content block 3 Avatar 11 Category
[0123] 4 Input device 12 line
[0124] 4A Keyboard 20 13 Line
[0125] 4B Mouse 14 lip movement sequence
[0126] 4C touchpad 14A short sequence 4D camera 14B long sequence
[0127] 4E Microphone 15 Group
[0128] 5 Input content 25 16 Line
[0129] 6 users 17 lines
[0130] 7 Data processing device 8 Output device A1 -A12 Process steps
Claims
Patent claims:
1. Method for synchronizing lip movements and acoustic output of a content (9) by an avatar (3), wherein a content (9) to be pronounced is generated, characterized in that the pronunciation of the content (9) is adapted to a predetermined lip movement sequence (14) of the avatar (3).
2. Method according to claim 1, characterized in that the accentuation of individual sounds, syllables, letters and / or words of the content (9) is adapted to the predetermined lip movement sequence (14).
3. Method according to claim 1 or 2, characterized in that the output speed of at least one sound, one syllable, one letter and / or one word is adapted to the predetermined lip movement sequence (14) at least passage by passage.
4. Method according to one of the preceding claims, characterized in that at least one pause is inserted between sounds, words and / or syllables and / or adapted to the predetermined lip movement sequence (14).
5. Method according to one of the preceding claims, characterized in that defined sounds, syllables, letters and / or words are assigned to a defined lip movement and / or defined lip positions.
6. Method according to one of the preceding claims, characterized in that the output speed, emphasis and / or pauses between the defined sounds, syllables, letters and / or words are adapted to the predetermined lip movement sequence (14).
7. Method according to one of the preceding claims, characterized in that the predetermined lip movement sequence (14) is selected from a plurality of predetermined lip movement sequences (14).
8. Method according to claim 7, characterized in that the selection of the predetermined lip movement sequence (14) is carried out as a function of the length and / or number of words, the number of vowels and / or the language of the content (9).
9. Method according to one of the preceding claims, characterized in that the content (9) is selected and / or generated from a plurality of predetermined content modules (10).
10. The method according to claim 9, characterized in that the number of predetermined content modules (10) is greater than the number of predetermined lip movement sequences (14).
11. Method according to claim 9 or 10, characterized in that the number of predetermined content modules (10) is more than twice as large, more preferably more than four times as large, more preferably more than ten times as large, more preferably more than one hundred times as large, more preferably more than one thousand times as large, as the number of predetermined lip movement sequences (14).
12. Method according to one of the preceding claims, characterized in that the predetermined lip movement sequences (14) are divided into a group (15) of short sequences (14A) and a group (15) of long sequences (14B).
13. The method according to claim 12, characterized in that the short sequences (14A) have a length between 1.5 s and 5 s, preferably between 1.5 s and 4 s, more preferably between 2 s and 3 s, and / or that the long sequences (14B) have a length of 5 s to 15 s, more preferably from 5 s to 13 s, more preferably from 6 to 12 s.
14. Method according to one of the preceding claims, characterized in that the content (9) is or comprises an answer and / or reaction and / or part of an answer and / or reaction to an input content (5).
15. Method according to claim 14, characterized in that the input content (5), in particular in speech form, is entered by a user (6) and is converted into a input text and / or binary code, or that the input content (5) is entered in text form as input text and / or as binary code.
16. The method according to claim 14 or 15, characterized in that the input content (5) is analyzed using words, expressions and / or synonyms, in particular by means of a large language module.
17. Method according to one of claims 9 to 11 or according to one of claims 12 to 16 dependent on claims 9 to 11, characterized in that a predetermined content module (10) or parts thereof are selected based on the analysis of the input content (5).
18. Method according to claim 17, characterized in that the predetermined content modules (10) are divided into at least two thematic categories (11).
19. The method according to claim 18, characterized in that based on the analysis of the input content (5) a category (11) is selected for selecting a predetermined content module (10) from this category (11).
20. Method according to one of the preceding claims, characterized in that at least one method step is carried out by means of a computer.
21. A data processing device comprising means for carrying out a method according to one of the preceding claims.
22. A computer program comprising instructions which, when executed by a computer, cause the computer to carry out a method according to any one of claims 1 to 20.
23. A computer-readable data carrier on which the computer program according to claim 22 is stored.
Citation Information
Patent Citations
Methods for communication between a natural person and an artificial language system, as well as communication systems
DE102007042583B4
Method and apparatus for generating animation
US20190392625A1
Method and apparatus for providing interactive avatar services
US20230230303A1