Method for an avatar to acoustically and visually output content to be spoken
The method addresses time delays and abrupt transitions in avatar communication by using short movement sequences for immediate and synchronized acoustic and visual output, resulting in a more natural and intuitive user experience.
Patent Information
- Application Number
- PCT/EP2024/078549
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-06
- Filing Date
- 2024-10-10
- Publication Date
- 2025-06-12
AI Technical Summary
Conventional methods for acoustic and visual output of content by an avatar result in significant time delays and abrupt transitions, leading to uncomfortable and unnatural communication experiences.
A method utilizing short, predetermined movement sequences with a playback length of less than 1 second, allowing for immediate and synchronized acoustic and visual output of content by the avatar, thereby eliminating time delays and abrupt transitions.
Enables fast and smooth transitions between the avatar's listening movements and spoken content output, enhancing user experience with natural and intuitive communication.
Smart Images

Figure EP2024078549_12062025_PF_FP_ABST
Abstract
Description
[0001] Method for the acoustic and visual output of a content to be spoken by an avatar
[0002] The present invention relates to a method for the acoustic and visual output of a content to be spoken by an avatar according to the preamble of claim 1, a device for data processing, a computer program and a computer-readable data carrier.
[0003] In conventional procedures, a user first provides input. Based on the input, content to be spoken is generated as a response. The content to be spoken can then be output by the avatar. The content to be spoken is output acoustically. In addition to the acoustic output, there is usually a visual output, with the avatar performing lip movements synchronized with the acoustic output. During input and / or the generation of the content to be spoken, the avatar can be displayed statically using a still image or dynamically without speech movements using moving images. For a dynamic display, a video sequence with movements of the avatar can usually be played. Corresponding video sequences should be of sufficient length or duration to outlast the time it takes to input and generate the content to be spoken. The video sequence can take several minutes to play.
[0004] By the time the content to be spoken has been generated and can be output, the video sequence is usually already playing. Accordingly, the output of the content to be spoken must wait until the video sequence has finished. It can take several minutes for the video sequence to finish. The time between the user's input and the avatar's output is then extended in addition to the time already elapsed for the content to be generated. A correspondingly large time delay of up to several minutes between the user's input and the start of the avatar's output is detrimental to direct communication between the user and the avatar and is perceived by the user as uncomfortable.
[0005] Alternatively, it is possible to interrupt the video sequence as soon as the spoken content can be output. Such an interruption of the video sequence results in a sudden, particularly visual, transition between the video sequence and the avatar's movement becoming visible to the user when the spoken content is output. Such a sudden transition in the avatar's movement may be perceived by the user as an unnatural movement or playback of the avatar and may be perceived as disturbing.
[0006] US 2011 / 0148916 A1 discloses a communication window of a messenger through which text messages can be exchanged between two users. Each user can select an avatar, which is displayed on a background for both users when a communication window is open and / or during communication. The avatar can perform a movement and / or display an emotion depending on a sent text message. For example, the avatar can perform an animated laugh in response to a "LOL" message.
[0007] If the messenger detects that a user is absent, i.e., if a longer idle time is detected between user inputs, the avatar can depict the avatar falling asleep or falling out of the communication window 100. The moving animation of the avatar can be performed until a user initiates an input. US 2011 / 0148916 A1 thus discloses that the avatar is only displayed using a predefined motion sequence in exceptional cases, namely, when a user is inactive for a long period or in response to certain messages.
[0008] US 2017 / 0011745 A1 discloses an avatar that visually and acoustically outputs an operator's acoustic or text input, allowing communication with a customer to be conducted using an avatar. The visual output via the avatar is provided by a speech movement sequence. When the operator stops speaking, the speech movement sequence is played to the end until the next transition point. Using morphing technology, i.e., a computer-aided simulation, the avatar can then be transferred to a resting position. When the operator begins to speak, the hands are transferred from the pause position to the resting position. Using morphing technology, the head is transferred from its current position and representation to a state in which the avatar's speech can be reproduced seamlessly.
[0009] US 2009 / 0278851 A1 discloses an avatar configured to output a user's speech input in real time. To display the avatar, elementary sequences are generated that show the avatar speaking. The elementary sequences are copied, whereby the copied elementary sequences are identical to the original elementary sequences, but do not contain any speech movements. The copied elementary sequences thus show the avatar in a speaking motion without the associated mouth movements. Individual short sequences from the original elementary sequence can be exchanged for short sequences from the copied, speech-movement-free elementary sequence in order to visually represent pauses in speech.
[0010] The object of the present invention is to provide a method for the acoustic and visual output of a content to be spoken by an avatar that is improved compared to the prior art, wherein a fast and smooth transition between a movement of the avatar without speech movements to an acoustic and visual output of a content to be spoken is enabled and / or supported, wherein a particularly simple, user-friendly and / or intuitive communication is created and / or supported.
[0011] The object underlying the present invention is achieved by the method according to claim 1, the data processing device according to claim 19, the computer program according to claim 20 or the computer-readable data carrier according to claim 21.
[0012] The present invention relates to a method, in particular a computer-implemented method, for the acoustic and visual output of a content to be spoken by an avatar.
[0013] In the context of the present invention, the term “avatar” refers to digital beings with an anthropomorphic appearance that are controlled by humans or software and have the ability to interact.
[0014] Preferably, the proposed method, in particular individual or all method steps of the proposed method, is / are carried out (semi-)automatically or automatically by means of a data processing device, in particular by appropriate means for data processing and controlling the device, such as a data processing unit or the like. In the proposed method for the acoustic and visual output of content to be spoken by an avatar, an input is first provided by a user. The content to be spoken is generated in response to the input and can then be output by the avatar.
[0015] The proposed method is characterized in that a visual representation of the avatar is carried out during the input by the user and / or the generation of the content to be spoken and / or until the output of the content to be spoken by means of at least one predetermined movement sequence, wherein the predetermined movement sequence has a predetermined playback length of less than 1 s, and / or wherein the predetermined movement sequence has a predetermined number of less than 60 frames.
[0016] In the context of the present invention, the term “movement sequence” is to be understood as a particularly pre-produced or prepared video file or a file with moving images that shows a movement of the avatar or the avatar in motion.
[0017] Due to the short playback length of the movement sequence, the output of the spoken content can occur within less than 1 second after the content to be spoken has been generated and, in particular, can be output with synchronized lip movements of the avatar. Due to the short predetermined playback length, it is not necessary to interrupt the playback of the movement sequence, which avoids a sudden transition to the output of the content to be spoken by the avatar. The movement sequence can be played to its end without creating a waiting time that the user perceives as long. The representation of the avatar can then transition smoothly from a moving representation during input to the speaking movement when the content to be spoken is output.
[0018] In the context of the present invention, the term "smooth" means that the avatar is displayed smoothly, and that the user perceives a smooth, natural movement of the avatar during the transition from the movement sequence to the output of the spoken content, rather than a sudden transition. It is then also unnecessary to simulate or calculate the avatar's movements during the user's input until the output of the spoken content and / or a transition between the movements during input and output. In this way, the method can be implemented particularly efficiently and with minimal computing power.
[0019] Particularly preferably, a visual representation of a naturally listening movement of the avatar occurs during input by the user and / or the generation of the content to be spoken and / or until the output of the content to be spoken by means of at least one predetermined movement sequence. The term "naturally listening movement" is understood here to mean the representation of a natural action or natural movements of a real person. In this way, a realistic representation of a real conversation partner can be created in a simple and efficient manner, even if no output is produced by the avatar. In particular, it is not necessary to simulate the avatar in a complex manner or to resort to morphing technology or similar.The content to be spoken can then be output as quickly as possible after its generation, since only the short period of time until the end of the specified movement sequence has to be waited for until a smooth transition to the acoustic and visual output of the avatar can take place.
[0020] The waiting time for the user can be further shortened if the movement sequence has a predetermined playback length of less than 0.75 s, preferably less than 0.6 s, more preferably 0.5 s or less, and / or more than 0.25 s, preferably that the movement sequence has a predetermined playback length of 0.4 s to 0.5 s.
[0021] Alternatively or additionally, it is also possible for the movement sequence to have a predetermined number of less than 40 frames, preferably less than 30 frames, more preferably less than 25 frames, and / or more than 10 frames, in particular for the movement sequence to have a predetermined number of 13 to 23 frames.
[0022] The input by the user and the generation of the content to be spoken can take longer than 1 s. It is then possible to play the specified movement sequence repeatedly one after the other. In order to achieve a smooth transition between the end and the beginning of the movement sequence, the beginning, in particular the first frame, and the end, in particular the last frame, of the movement sequence can be identical. Alternatively, the beginning, in particular the first frame, and the end, in particular the last frame, of the movement sequence can be so similar to one another that a smooth transition between the end and the beginning of the movement sequence is ensured upon repeated playback of the movement sequence. In this way, the user can be given a feeling of natural communication.
[0023] The movement sequence preferably shows a blink, a wrinkling of the nose, a smile, an eye movement, a hand movement, a head movement and / or an adjustment of glasses and / or a piece of clothing or jewelry of the avatar, whereby the representation of other, in particular natural, movements is also possible.
[0024] According to a further aspect of the present invention, preferably several predefined movement sequences are provided, stored, or stored. The movement sequence can then be selected from a plurality of movement sequences. It is then possible to display various movements of the avatar one after the other, for example, a blink followed by a head movement. The movement sequence can be selected randomly, in a predefined order, and / or according to other selection criteria.
[0025] For example, it is possible to perform a real-time analysis of the user's input, and select the movement sequence based on the real-time analysis. This allows the avatar to provide a timely nonverbal response to the user's input, such as a smile in response to a positive statement.
[0026] If two or more different predefined movement sequences are played back one after the other, it is advantageous if a smooth transition is provided between the individual movement sequences. This can be achieved if the beginning, in particular the first frame, and the end, in particular the last frame, of each movement sequence are the same or so similar that a smooth transition is achieved between the movement sequences when two consecutive movement sequences are played back. During the visual representation of the avatar using the movement sequence, a termination criterion can be checked. In particular, if the termination criterion is met, the movement sequence can be played to the end and the content to be spoken can then be output by the avatar.Alternatively or additionally, if the termination criterion is not met, the movement sequence can be played to its end and then another movement sequence can be played, or the movement sequence can be played repeatedly. The termination criterion is particularly met if the generated content can be played back acoustically with synchronous lip movements by the avatar.
[0027] The user's input can be analyzed, in particular, using a large language model. Based on the input analysis, content to be spoken can be generated and / or selected.
[0028] For the avatar to output the content to be spoken with synchronous lip movements, a predefined lip movement sequence can be used and played or output simultaneously with the acoustic output of the content to be spoken. Within the scope of the present invention, the term "lip movement sequence" refers to a lip movement that lasts for a predefined period of time.
[0029] In particular, the lip movement sequence can be selected from a plurality of predefined lip movement sequences. It is possible for the predefined lip movement sequence to be selected depending on the content to be spoken in order to select a lip movement sequence that best matches the content.
[0030] In order to ensure a smooth transition when playing a predetermined movement sequence and a subsequent predetermined lip movement sequence, it is preferably provided that the beginning, in particular the first frame, of the predetermined lip movement sequence and the end, in particular the last frame, of the movement sequence are identical or similar in design such that a smooth transition is achieved when playing the movement sequence followed by playing the lip movement sequence.
[0031] After the output, the user can enter the input again. The avatar is then represented, in particular, by naturally listening movements. Preferably, the beginning, in particular the first frame, of the predefined movement sequence and the end, in particular the last frame, of the predefined lip movement sequence can be identical or so similar that a smooth transition is achieved when the lip movement sequence is played back followed by the movement sequence.
[0032] It is preferably provided that at least one method step is carried out by means of a computer.
[0033] According to a further aspect of the present invention, a data processing device comprising means for implementing the proposed method is proposed. Reference may be made to all statements regarding the proposed method in this regard. In particular, corresponding advantages are achieved.
[0034] According to a further aspect of the present invention, a computer program comprising instructions that, when executed by a computer, cause the computer to execute a proposed method is proposed. Reference may be made to all explanations of the proposed method in this regard. In particular, corresponding advantages are achieved.
[0035] According to a further aspect of the present invention, a particularly non-volatile, computer-readable data carrier on which the proposed computer program is stored is proposed. Reference may be made to all statements regarding the proposed computer program in this regard. In particular, corresponding advantages are achieved.
[0036] The aforementioned aspects, features and method steps as well as the aspects, features and method steps of the present invention resulting from the claims and the following description can in principle be implemented independently of one another, but also in any desired combination and / or sequence.
[0037] The aforementioned aspects, features and method steps as well as the aspects, features and method steps of the present invention resulting from the claims and the following description can in principle be implemented independently of one another, but also in any desired combination and / or sequence.
[0038] Further aspects, advantages, features, properties, and advantageous developments of the present invention will become apparent from the claims and the following description of preferred embodiments with reference to the figures. They show, in a schematic representation, not to scale:
[0039] Fig. 1 is a schematic view of a proposed data processing device designed to output a content to be spoken acoustically and visually through an avatar,
[0040] Fig. 2 shows a schematic flow diagram of a proposed method for the acoustic and visual output of a content to be spoken by an avatar or individual method steps of the proposed method,
[0041] Fig. 3 is a schematic representation of two predetermined movement sequences that can be played one after the other, the beginning and the end of the movement sequences being identical, and
[0042] Fig. 4 is a schematic representation of a lip movement sequence and a subsequently playable movement sequence with identical beginning and end.
[0043] In the figures, some of which are not to scale and are merely schematic, the same reference symbols are used for identical, identical or similar parts and components, whereby corresponding or comparable properties or advantages are achieved, even if repetition is omitted.
[0044] Fig. 1 shows a schematic view of a data processing device 1, in particular a computer. The data processing device 1 may also be a laptop, a tablet, a mobile phone, or the like.
[0045] The device 1 preferably has a visual output device, in particular a screen 2, for displaying or reproducing an avatar 3. The screen 2 is designed in particular for reproducing moving images, in particular of the avatar 3. Using the screen 2, movements of the avatar 3, and in particular movements of the avatar 3 while speaking, can be displayed, preferably smoothly.
[0046] Furthermore, the device 1 has an input device 4 for inputting an input content 5. The input device 4 can be an integral part of the device 1 or be designed as a separate input device.
[0047] The input device 4 can, for example, have a keyboard 4A, a mouse 4B, a touchpad / trackpad 4C, a camera 4D, a microphone 4E and / or a touchscreen for entering the input content 5.
[0048] For example, the microphone 4E can be integrated into the camera 4D. Alternatively or additionally, it is possible for the screen 2 to be designed as a touchscreen or for a touchscreen to be integrated into the screen 2.
[0049] A user 6 can enter input content 5 (Fig. 2) via the input device 4. The input content 5 can in particular be an acoustic input, in particular by means of the microphone 4E. For example, the user 6 can ask a question acoustically or acoustically request certain information. Alternatively or additionally, the input content 5 can also be entered using the keyboard 4A, the mouse 4B, the touchpad / trackpad 4C, the camera 4D and / or the touchscreen. A combined acoustic and motor input using the microphone 4E and keyboard 4A, mouse 4B, touchpad 4C and / or touchscreen is also conceivable. Here and preferably, the input content 5 is a text, in particular a spoken text.
[0050] The device 1, in particular a data processing device 7 of the device 1, can analyze the input content 5 and, based on the analysis, generate content 8 to be spoken, in particular text. In the context of the present invention, a “data processing device” is understood to mean a device for the automatic processing of data, in particular program code. The data processing device 7 can in particular have a processor, memory, an internal and / or external database and / or a communication device for connection to the Internet and / or another computer. In the context of the present invention, the term “content” is understood to mean information that can be presented acoustically and / or visually. For example, the content 8 can be text, in particular spoken text, tones, sounds or the like.
[0051] The device 1 preferably further comprises an acoustic output device 9 for acoustically outputting the content 8. The acoustic output device 9 may, for example, be a loudspeaker that is an integral component of the device 1 or is formed as a separate part / device of the device 1 and / or is combined with the input device 4.
[0052] The previously generated content 8 can be acoustically output by the avatar 3 via the output device 9. At the same time, the visual output can be provided by the avatar 3 via the screen 2. In this way, communication between the user 6 and the avatar 3 can be enabled.
[0053] To support the acoustic output of the content 8, the device 1 is preferably designed to synchronize the lip movements of the avatar 3 and the acoustic output of the content 8. Within the context of the present invention, the term "lip movement" refers to the lip movement and preferably the facial expression and / or facial expression of the avatar 3 when pronouncing the content 8.
[0054] The user 6 thus hears the acoustic output by the avatar 3 and sees on the screen 2 a lip movement of the avatar 3 that is synchronous with the acoustic output. In this way, the understanding of the user 6 can be increased and the communication between the user 6 and the avatar 3 can be improved.
[0055] While the user 6 enters the input content 5 via the input device 4, the avatar 3 is preferably displayed on the screen 2. The avatar 3 is preferably shown in a moving state, wherein a naturally listening movement of the avatar 3 is represented. The term "naturally listening movement" in the context of the present invention is understood to mean a movement of a natural person who is listening to a user 6. In this way, the avatar 3 represents and / or imitates a natural action or a natural movement of a real person. In particular, the avatar 3 does not perform any speaking or speech-like lip movements while the user 6 is entering the input content 5.
[0056] While the spoken content 8 is being generated, the avatar 3 is preferably displayed with natural listening movements. The transition between the natural listening movement and the output of the spoken content 8 is then preferably designed to be smooth. The display of the avatar 3 is thus smooth, and during the transition between the natural listening movement and the output of the spoken content 8, the user 6 sees or recognizes a smooth, natural movement of the avatar 3, rather than a sudden transition.
[0057] In the context of the present invention, the term “output of the content to be spoken” is to be understood as the acoustic output of the content to be spoken and the visual display of the synchronous lip movements of the avatar, unless otherwise described.
[0058] In particular, it is provided that the visual representation of the avatar 3 during the input by the user 6 and / or the generation of the content 8 to be spoken and / or until the output of the content 8 to be spoken takes place by means of at least one predetermined movement sequence 10, wherein the predetermined movement sequence 10 has a predetermined playback length of less than 1 s and / or wherein the predetermined movement sequence has a predetermined number of less than 60 frames 11.
[0059] Fig. 2 shows a schematic flow diagram of a proposed method for the acoustic and visual output of a content 8 to be spoken by the avatar 3 or individual method steps of the proposed method.
[0060] The method is preferably multi-stage or multi-step. In particular, the method comprises several method steps, whereby the individual method steps can generally be performed independently of one another and in any order, unless otherwise explained below. The proposed method is preferably carried out using the data processing device 1, in particular using the input device 4, the screen 2, and / or the output device 9.
[0061] The data processing device 1 is preferably designed to carry out the method described herein or individual or all method steps.
[0062] Preferably, the instructions or the algorithm for executing the proposed method or individual method steps of the proposed method are stored electronically in a (data) memory of the device 1, in particular the data processing device 7.
[0063] However, it is also possible that one or more process steps are carried out by means of an (external) device or an (external) device and / or that individual or more commands for executing the process or individual process steps are stored there.
[0064] The proposed method preferably includes one or more method steps and / or a program to display the avatar 3 with a naturally listening movement during the input by the user 6 and / or the generation of the content 8 to be spoken until the output of the content 8.
[0065] The proposed method is characterized in that the visual representation of the avatar 3 during input by the user 6 and / or the generation of the content 8 to be spoken and / or until the output of the content 8 to be spoken is carried out by means of at least one predetermined movement sequence 10, wherein the predetermined movement sequence 10 has a predetermined playback length of less than 1 s and / or wherein the predetermined movement sequence 10 has a predetermined number of fewer than 60 frames 11. It is then possible to begin the output of the content 8 to be spoken within less than 1 s after the content 8 to be spoken has been generated and synchronized with lip movements. At the same time, a smooth transition can take place between the naturally listening movement of the avatar 3 and the output of the content 8 to be spoken by the avatar 3.In particular, a sudden interruption of the natural listening movement of the avatar 3 can be avoided in order to promptly output the content 8 to be spoken. In this way, a conversation can take place without a pause between the input and the output by the avatar 3, which the user 6 finds annoying.
[0066] The invention is based on the finding that a simulation or calculation of a movement of the avatar 3 between the naturally listening movement and the movements of the avatar 3 during the output of the content 8 to be spoken is not necessary if the naturally listening movement regularly reaches an end in less than 1 s and thus a defined state, to which a smooth transition to the output movement of the avatar 3 is possible. An abrupt termination of the naturally listening movement can also be dispensed with, thereby avoiding a sudden transition to the movement of the avatar 3 during the output of the content 8 to be spoken.
[0067] It is then particularly possible to realize a smooth transition at defined points—namely, at the end of the predefined movement sequence 10—to output the content 8 to be spoken. In this way, a smooth transition between the naturally listening movement and the movement of the avatar 3 during the output of the content 8 to be spoken can be achieved with low computing power and / or little effort.
[0068] The method is preferably initiated by the user 6, for example, by opening a corresponding program and / or a corresponding website and / or entering a corresponding command. For example, the user 6 can open a corresponding program or enter a corresponding command using the input device 4, preferably in a first method step / process A1.
[0069] It is also possible that the method is started by the user 6 pronouncing a predefined command, for example by pronouncing “Hello Avatar” or another start command.
[0070] Optionally, in a further / second method step A2, the user 6 can enter an input content 5, in particular as spoken text, via the input device 4. The input content 5 can be a question, a request for information, and / or a statement. Here, and preferably, the input content 5 is a question. The input content 5 is preferably recorded and / or stored in order to perform the processing described below.
[0071] In a further / third method step A3, the input or input content 5 of the user 6 is preferably analyzed.
[0072] Thus, after the input or input content 5 has been captured and / or after the input or input content 5 has been saved (second method step A2), processing can be performed to improve the quality of the captured input content 5 and increase the recognition accuracy of the captured input content 5. Alternatively, processing can be started when parts of the input or input content 5 have been captured or saved. For example, background noise and interference can be eliminated. Alternatively or additionally, certain frequencies can be filtered to eliminate irrelevant frequencies. It is also possible, additionally or alternatively, to standardize the volume level of the sound. It is also possible to perform only some or all of the aforementioned processing steps, for example sequentially, in parallel, and / or iteratively.
[0073] Furthermore, in the third method step A3, phonetic matches and patterns can be searched for. The matches and patterns can then be compared with patterns stored in a database to identify at least individual components of the input content 5. The database can be an internal and / or external database, in particular a cloud-based database. The database can be embodied as a memory, in particular a hard disk. The database can be a hard disk component of the device 1 or of another computer.
[0074] In the third method step A3, the input content 5, in particular spoken text, can be converted into text and / or code, in particular machine-readable text. The conversion can be performed using a voice-to-text application. The voice-to-text application can, in particular, be a speech recognition system that transcribes spoken text.
[0075] It is also possible for the third method step A3 to be designed to support speech recognition through machine learning and / or artificial intelligence and / or to improve the accuracy and adaptability of processing. Preferably, the third method step A3 analyzes the logical content or meaning of the particularly transcribed input content 5 or the input text. A Large Language Module (LLM) is preferably used for this purpose. The term “Large Language Module” in this case refers to a generative language model for texts or text-based content. Large Language Modules use artificial intelligence, artificial neural networks and / or deep learning to process, understand and / or optionally generate natural language. Corresponding language models are used to understand complex texts, questions and / or instructions.
[0076] Here, and preferably, the input content 5 or the input text is analyzed using the Large Language Model based on expressions, synonyms, words, sentence components, and / or sentences. With the help of the Large Language Model, the meaning of the input content 5 of the user 6 can thus be captured or understood.
[0077] Subsequently, a reaction or response to the input content 5 of the user 6 is preferably generated in a further / fourth method step A4.
[0078] The device 1 or the program can have several stored predefined content modules 12. The predefined content modules 12 can be answers and / or information on various topics.
[0079] The content modules 12 can be information or reactions that have been verified, for example, by humans and are therefore objectively correct. In this way, only information or reactions based on objectively correct content modules 12 can be output. This prevents fictitious and / or inobjectively correct information or reactions from being output by the avatar 3 or the device 1. Preferably, the content module 12 is embodied as text or contains text.
[0080] The content modules 12 can be stored or saved in a database. The database can be designed in particular as a memory chip and / or hard disk. The database can be a component of the device 1 or as an internal database. It is also possible for the database to be an external database. The device 1 can then access the database, for example, via the Internet. In order to be able to answer complex questions or to respond adequately to complex input content 5, the fourth method step A4 can be run through iteratively in order to select several predetermined content modules 12. In this way, a large number of information items can be used to generate the content 8 to be spoken. In Fig. 2, the selection of two content modules 12 is shown by the line 13A.
[0081] The selected predefined content modules 12 or the selected predefined content module 12 can be processed into a content 8 to be spoken. Alternatively or additionally, the content 8 to be spoken can be formulated from the selected predefined content module 12 or the selected predefined content modules 12. The content 8 to be spoken can be a response to the input content 5 or a reaction to the input content 5. Alternatively or additionally, the content 8 to be spoken can comprise part of a response to the input content 5 or part of a reaction to the input content 5.
[0082] If several predefined content modules 12 are selected, these can be further processed individually or combined into a logical information unit in an order adapted to the input content 5 or arranged one after the other and / or nested in order to formulate the content 8 to be spoken.
[0083] Preferably, the content 8 to be pronounced corresponds to a selected content module 12. If several predetermined content modules 12 are selected, the content 9 to be pronounced preferably corresponds to the sequential and / or nested content modules 12, as shown in Fig. 2.
[0084] Subsequently, in a further / fifth process step A5, a lip movement synchronous to the content to be spoken 8 can be simulated and / or calculated. In this way, a synchronous lip movement can be generated for any content to be spoken, regardless of the language used.
[0085] Alternatively or additionally, it is also possible to use a predefined lip movement sequence 14 to generate lip movement synchronized with the content 8 to be spoken. Provision can be made for a predefined lip movement sequence 14 to be played back along with the content 8 to be spoken.
[0086] In particular, the predetermined lip movement sequence 14 can be selected from a plurality of predetermined lip movement sequences 14, as shown in Fig. 2 by line 13B. In this way, a suitable lip movement sequence 14 can be selected. In particular, it is possible to select the predetermined lip movement sequence 14 depending on the content 8 to be pronounced. In this way, a lip movement sequence 14 that best matches the content 8 to be pronounced can be selected.
[0087] Subsequently, in the fifth method step A5, the lip movement sequence 14 and the content 8 to be pronounced can be synchronized with each other in order to allow the acoustic output of the content 8 to be pronounced and the visual output of the lip movement sequence 14 to be synchronized with each other.
[0088] Subsequently, in a further / sixth method step A6, the acoustic and visual output of the content 8 to be pronounced can take place, i.e. the synchronous acoustic output of the content 8 via the acoustic output device 9 and the visual output of the lip movement sequence 14 via the screen 2.
[0089] Following the first method step A1, in a further seventh method step A7, which runs in particular parallel to method step A2, the avatar 3 is displayed or represented, preferably on the screen 2. The avatar 3 is preferably shown in a moving state, with a natural listening movement of the avatar 3 being represented. In this way, the avatar 3 imitates a natural action or movement of a real person. In particular, the avatar 3 does not perform any speaking or speech-like lip movements during the third method step A3.
[0090] In particular, it is provided that the visual representation of the avatar 3 during input by the user 6 takes place by means of at least one predetermined movement sequence 10, wherein the predetermined movement sequence 10 has a predetermined playback length of less than 1 s and / or wherein the predetermined movement sequence 10 has a predetermined number of fewer than 60 frames 11. It is then possible to have the movement sequence 10 and thus the naturally listening movement of the avatar 3 transition into the acoustic and visual output by the avatar 3 at regular intervals less than 1 s apart at defined points. In this way, a smooth transition can be created.
[0091] In particular, it is provided that the visual representation of the avatar 3 also occurs during the generation of the content 8 to be spoken and / or until the output of the content 8 to be spoken by means of the at least one predetermined movement sequence 10. The representation of the avatar 3 then preferably occurs during method steps A2 to A5 by means of the predetermined movement sequence 10.
[0092] In particular, the movement sequence 10 has a predetermined playback length of less than 0.75 s, preferably less than 0.6 s, more preferably 0.5 s or less, and / or more than 0.25 s. The movement sequence 10 can, in particular, have a predetermined playback length of 0.4 s to 0.5 s. In this way, a smooth transition to the output of the content 8 to be spoken can be generated within a particularly short period of time.
[0093] Alternatively or additionally, it is provided that the motion sequence 10 has a predetermined number of fewer than 40 frames 11, preferably fewer than 30 frames 11, more preferably fewer than 25 frames 11, and / or more than 10 frames 11. In particular, the motion sequence 10 can have a predetermined number of 13 to 23 frames 11.
[0094] In particular, during the visual representation of the avatar 3 by means of the movement sequence 10 in the seventh method step A7, a termination criterion can be checked in a further, in particular parallel, eighth method step A8. If the termination criterion is met, the movement sequence 10 currently being played can be played to its end, and the content 8 to be spoken can then be played by the avatar 3. If the termination criterion is not met, the movement sequence 10 can be played to its end, and then a new movement sequence 10 can be played.
[0095] The termination criterion is preferably met when the content 8 to be spoken and synchronous lip movements are present, or a synchronous output of the content 8 to be spoken can be achieved acoustically and the synchronous lip movements can be achieved visually. In particular, the termination criterion is only met when the fifth method step A5 is completed.
[0096] As long as the content 8 to be pronounced and / or the lip movements are not yet available and / or a synchronous output of the content 8 to be pronounced acoustically and the synchronous lip movements in visual form cannot yet be achieved, the termination criterion is preferably not met.
[0097] Accordingly, the method steps A7 and A8 can be repeated in order to output several movement sequences 10 in succession until the sixth method step A6 - the output of the content 8 to be pronounced - can be carried out.
[0098] In order to achieve a smooth and / or consistent transition during repeated playback of the movement sequence 10, it is preferably provided that the beginning and the end of the movement sequence are identical or so similar to one another that a smooth transition between the end and the beginning of the movement sequence 10 is ensured during repeated playback of the movement sequence 10 (cf. Fig. 3 and Fig. 4).
[0099] In particular, it can be provided that the first frame 11A and the last frame 11B of the motion sequence 10 are identical, as shown in Fig. 3 and Fig. 4. Alternatively, it is also possible for the first frame 11A and the last frame 11B of the motion sequence 10 to be so similar to one another that, upon repeated playback of the motion sequence 10, a smooth transition between the end and the beginning of the motion sequence 10 is ensured.
[0100] In order to avoid having to repeat one movement sequence 10 several times in immediate succession in the seventh method step A7, the predefined movement sequence 10 can be selected from a plurality of predefined movement sequences 10. In Fig. 2, the selection of a movement sequence 10 is shown by line 13C. It is then possible to display various movements of the avatar 3 one after the other, thus creating a natural appearance of the avatar 3 for the user 6. The selection of the movement sequence 10 from the plurality of movement sequences 10 can be random or in a predefined order and / or frequency, or according to other selection criteria.
[0101] The movement sequence 10 may show a blink, a wrinkling of the nose, a smile, an eye movement, a head movement, a hand movement, and / or an adjustment of glasses and / or a piece of clothing or jewelry of the avatar 3 in order to obtain a natural movement by the avatar 3. Other movements, such as a frown, a shrug, a movement of the upper body, or the like, are also possible.
[0102] Thus, during process steps A2 to A5, and thus during process step A7, it is possible to represent the movement of avatar 3 through blinking, smiling, eye movements, and / or the like. In this way, avatar 3 can be represented with varied movements.
[0103] In order to achieve a smooth transition between successive, different movement sequences 10, the beginning and the end of a movement sequence 10 are preferably identical or similar in design such that when two successive movement sequences 10 are played back, a smooth transition between the movement sequences 10 is achieved.
[0104] In particular, it is provided that the first frame 11A and the last frame 11B of a motion sequence 10 are identical (Fig. 3). Alternatively, it is also possible for the first and last frames 11A, 11B of a motion sequence 10 to be so similar to each other that, when two consecutive motion sequences 10 are played back, a smooth transition between the motion sequences 10 is achieved.
[0105] It is then easily possible to create a smooth transition between the movements of the avatar 3 or the movement sequences 10.
[0106] In particular, the transition between the naturally listening movement of the avatar 3 in the seventh method step A7 and the movements of the avatar 3 during the acoustic and visual output of the content 8 to be spoken in the sixth method step A6 can be designed or displayed smoothly. If the lip movements of the avatar 3 are simulated or calculated for the content 8 to be spoken, the beginning of the simulation, in particular a first frame 11, can be identical or similar to the end, in particular the last frame 11B, of the movement sequence 10, such that a smooth transition is achieved during the playback of the movement sequence 10 and the subsequent visual output of the lip movement.
[0107] Preferably, a predefined lip movement sequence 14 is used or played back to represent the lip movements. In particular, it is provided that the beginning, in particular the first frame 11C, of the predefined lip movement sequence 14 is identical to the end, in particular the last frame 11B, of the movement sequence 10 or is designed in such a way that a smooth transition is achieved when the movement sequence 10 is played back followed by the lip movement sequence 14, as shown in Fig. 4.
[0108] To enable further smooth communication between the user 6 and the avatar 3, the sixth method step A6 can be continued directly with the parallel method steps A2 and A7. A smooth transition between the movements of the avatar 3 during the output of the content 8 to be spoken in the sixth method step A6 and the subsequent natural listening movement of the avatar 3 in the seventh method step A7 can be achieved if the beginning of the movement sequence 10 and the end of the lip movement sequence 14 are identical or so similar to one another that a smooth transition is achieved when the lip movement sequence 14 is played back followed by the movement sequence 10.
[0109] In particular, the first frame 11A of the movement sequence 10 can be identical or similar to the last frame 11D of the predetermined lip movement sequence 14 (Fig. 4) so that a smooth transition is achieved when the lip movement sequence 14 is played back and the movement sequence 10 is subsequently played back.
[0110] The selection of the movement sequence 10 can be made depending on the input by the user 6 or depending on the input content 5. In particular, it can be provided that a real-time analysis of the input by the user 6 is performed. It is then possible for the selection of a movement sequence 10 and / or the movement sequences 10 to be made depending on the real-time analysis performed. For example, an astonished facial expression of the avatar 3 can be displayed in response to a component of the input content 5. In this way, the avatar 3 can provide visual feedback to the user 6 even while the input content 5 is being entered.
[0111] Preferably, assignment criteria can be or will be specified that enable an assignment of input content 5, components of the input content 5, analyses of the input content 5, and / or components of the analyses of the input content 5 to one or more movement sequences 10. Using the assignment criteria, one or more movement sequences 10 can then be selected.
[0112] The input content 5 is preferably analyzed in the third method step A3, in particular using a large language model, as already explained. Based on this analysis of the input or input content 5, a movement sequence 10 can be selected alternatively or additionally.
[0113] The avatar 3 can be an existing, real person or a fictitious person. For example, the avatar 3 can be a real or fictitious employee of a company, for example from the human resources department. The user 6 can also be an employee of the same company. By means of the method and / or the device 1, the user 6 can have relevant company-related questions answered in a particularly simple manner and quickly, without preventing or distracting other employees from their work. For example, a question about how to fill out a vacation request or how many vacation days the user 6 is still entitled to can be answered correctly in a particularly quick and easy manner. Other use cases or constellations are also possible.
[0114] A further aspect of the present invention relates to a computer program comprising instructions which, when executed by a computer or the device 1 according to the invention, cause the computer to execute the proposed method. Reference may be made to all explanations of the proposed method in this regard. In particular, corresponding advantages are achieved. A further aspect of the present invention relates to a computer-readable data carrier on which the proposed computer program is stored. Reference may be made to all explanations of the proposed computer program in this regard. In particular, corresponding advantages are achieved.
[0115] List of reference symbols:
[0116] Device 10 movement sequence
[0117] Screen 11 Frame
[0118] Avatar 11A Frame
[0119] Input device 11 B Frame A Keyboard 11C Frame B Mouse 11 D Frame C Touchpad 12 Content block E Camera 13A Line E Microphone 13B Line
[0120] Input content 13C line
[0121] User 14 lip movement sequence
[0122] Data processing facility
[0123] Contents A1 - A8 Procedural steps
[0124] Output device
Claims
Patent claims:
1. Method for the acoustic and visual output of content (8) to be spoken by an avatar (3), wherein an input is made by a user (6), wherein the content to be spoken (8) is generated in response to the input and output by the avatar (3), characterized in that a visual representation of a naturally listening movement of the avatar (3) during the input by the user (6) and / or the generation of the content (8) to be spoken and / or until the output of the content (8) to be spoken is carried out by means of at least one predetermined movement sequence (10), wherein the predetermined movement sequence (10) has a predetermined playback length of less than 1 s, and / or wherein the predetermined movement sequence (10) has a predetermined number of less than 60 frames (11).
2. Method according to claim 1, characterized in that a visual representation of a naturally listening movement of the avatar (3) during the input by the user (6) and / or the generation of the content to be spoken (8) and / or until the output of the content to be spoken (8) is carried out by means of at least one predetermined movement sequence (10).
3. Method according to claim 1 or 2, characterized in that the movement sequence (10) has a predetermined playback length of less than 0.75 s, preferably less than 0.6 s, more preferably 0.5 s or less, and / or more than 0.25 s, preferably that the movement sequence (10) has a predetermined playback length of 0.4 s to 0.5 s.
4. Method according to one of the preceding claims, characterized in that the movement sequence (10) has a predetermined number of less than 40 frames (11 ), preferably less than 30 frames (11 ), more preferably less than 25 frames (11 ), and / or more than 10 frames (11 ), preferably that the movement sequence (10) has a predetermined number of 13 to 23 frames (11).
5. Method according to one of the preceding claims, characterized in that the beginning, in particular the first frame (11 A), and the end, in particular the last frame (11 B), of the movement sequence (10) are identical or so similar to one another that a smooth transition between the end and the beginning of the movement sequence (10) is ensured upon repeated reproduction of the movement sequence.
6. Method according to one of the preceding claims, characterized in that the movement sequence (10) shows a blink, a wrinkling of the nose, a smile, an eye movement, a hand movement, a head movement and / or an adjustment of glasses and / or a piece of clothing or jewelry of the avatar (3).
7. Method according to one of the preceding claims, characterized in that the predetermined movement sequence (10) is selected from a plurality of predetermined movement sequences (10).
8. Method according to one of the preceding claims, characterized in that the selection of the movement sequence (10) is carried out randomly or in a predetermined order.
9. Method according to one of the preceding claims, characterized in that two or more different predetermined movement sequences (10) are played back directly one after the other, wherein the beginning, in particular the first frame (11 A), and the end, in particular the last frame (11 B), of each movement sequence (10) are the same or similar in such a way that when two successive movement sequences (10) are played back, a smooth transition between the movement sequences (10) is achieved.
10. Method according to one of the preceding claims, characterized in that during the visual representation of the avatar (3) by means of the movement sequence (10) a termination criterion is checked, wherein if the termination criterion is met, the movement sequence (10) is output to the end and then the content to be spoken is output by the avatar (3), and / or wherein if the termination criterion is not met, the movement sequence (10) is output to the end and then another movement sequence (10) is output or the movement sequence (10) is output repeatedly.
11. Method according to one of the preceding claims, characterized in that a real-time analysis of the user's input (6) is carried out and that the selection of the movement sequence (10) is carried out as a function of the real-time analysis.
12. Method according to one of the preceding claims, characterized in that the input is analyzed, in particular by means of a large language model, and that based on the analysis of the input, a content (8) to be pronounced is generated and / or selected.
13. Method according to one of the preceding claims, characterized in that a predetermined lip movement sequence (14) is played back for the content (8) to be pronounced.
14. The method according to claim 13, characterized in that the predetermined lip movement sequence (14) is selected from a plurality of predetermined lip movement sequences (14).
15. Method according to claim 14, characterized in that the predetermined lip movement sequence (14) is selected depending on the content (8) to be pronounced.
16. Method according to one of the preceding claims, characterized in that the beginning, in particular the first frame (11 C), of the predetermined lip movement sequence (14) and the end, in particular the last frame (11 B) of the movement sequence (10) are identical or similar in design such that a smooth transition is achieved when the movement sequence (10) is played back with subsequent playback of the lip movement sequence (14).
17. Method according to one of the preceding claims, characterized in that the beginning, in particular the first frame (11 A), of the movement sequence (10) and the end, in particular the last frame (11 D), of the predetermined lip movement sequence (14) are identical or so similar to one another that a smooth transition is achieved when the lip movement sequence (14) is played back with subsequent playback of the movement sequence (10).
18. Method according to one of the preceding claims, characterized in that at least one method step is carried out by means of a computer.
19. Device for data processing comprising means for carrying out a method according to one of the preceding claims.
20. A computer program comprising instructions which, when executed by a computer, cause the computer to carry out a method according to any one of claims 1 to 18.
21. A computer-readable data carrier on which the computer program according to claim 20 is stored.
Citation Information
Patent Citations
Method and system for animating an avatar in real time using the voice of a speaker
US20090278851A1
Modifying avatar behavior based on user action or mood
US20110148916A1
Real-time Animation for an Expressive Avatar
US20120130717A1
Virtual photorealistic digital actor system for remote service of customers
US20170011745A1