Apparatus and method for generating avatar lip synching animation based on multimodal biometric signal
The multimodal bio-signal-based system generates avatar lip-sync animation by processing brain and muscle activity during speech imagination, addressing limitations of existing methods by enabling silent communication and nuanced expression through avatars.
Patent Information
- Application Number
- JP2025103907
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-25
- Filing Date
- 2025-06-19
- Publication Date
- 2026-01-14
AI Technical Summary
Existing lip-sync animation techniques require recorded voice data or direct speech, limiting their use in quiet situations and for individuals with speech difficulties, and brain-computer interfaces face challenges in real-time decoding and recognition, especially for subtle emotions and nuanced facial expressions.
A multimodal bio-signal-based system that collects electroencephalograms and electromyograms during speech imagination, processes the data to generate an avatar, and implements lip-sync animation using a lip-sync reconstruction model to predict mouth and facial movements.
Enables lip-sync animation without direct speech, accurately conveying users' intentions and emotions through avatars, suitable for various applications including virtual reality and communication systems.
Smart Images

Figure 2026004252000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an apparatus and method for generating avatar lip-sync animation based on multimodal bio-signals, and more particularly to an apparatus and method for generating avatar lip-sync animation based on multimodal bio-signals, which can generate an avatar corresponding to a user's facial image and implement avatar lip-sync animation based on multimodal bio-signals when the user imagines speaking using a pre-defined lip-sync reconstruction model. [Background technology]
[0002] Brain-Computer Interface (BCI) is a technology that directly connects the brain's neurological signals to a computer system, enabling communication and control. To achieve this, various biosignal measurement technologies are used to identify the user's brain activity, such as thought, concentration, and imagination, and convert this into digital commands. Brain-computer interfaces have become an invaluable tool, especially for people with limited motor skills, allowing them to perform activities such as using a computer, moving a mechanical arm, or even controlling a wheelchair. The technology is offering new modes of interaction in fields as diverse as virtual reality and video games, neuroscience research, and even art and music creation. Recently, with the development of algorithms for analyzing brain signals, brain-computer interface technology has become more sophisticated, and this has the potential to revolutionize future human-machine interactions. Meanwhile, recent research has focused on a method for realizing a speaking human face through lip synchronization between a face synthesized by computer graphics and a human voice. As a prior art, Korean Patent Publication No. 10-2006-0031449 (published on April 12, 2006) proposes an "audio-based automatic lip-sync animation apparatus, method, and recording medium." Existing lip-sync animation techniques, including the above-mentioned conventional techniques, have mainly focused on a method of reconstructing mouth shapes based on input of voice data. However, existing methods for generating talking faces through lip-sync animation require the use of recorded voice data of the user speaking directly. Therefore, existing systems have the problem of being unable to utilize voice data in quiet situations or for patients who have difficulty speaking additionally, and have limitations in being unable to express subtle emotions such as the user's facial expressions and nuances. Meanwhile, another prior art, Korean Patent Publication No. 10-2020-0052807 (published on May 15, 2020), proposed a "brain-computer interface system and a method for recognizing user dialogue intentions using the same." However, communication systems based on brain-computer interfaces, including the above-mentioned other conventional technologies, have mainly been based on passively reading and transmitting the user's intentions, such as simple class classification and sentence generation, through brain waves during speech. Communication systems that utilize brain signals have been developed extensively in the field of brain-computer interfaces recently, and various methodologies have been developed by combining them with the field of artificial intelligence.
[0003] Among these, the speech imagining-based user communication technology has the advantage of being able to communicate the user's intentions without the user's direct speech. However, in brain-computer interfaces, there are limitations to the technology used to communicate users' intentions, such as a decline in real-time decoding performance or a low recognition rate, making it difficult to synthesize speech at an understandable level. Various methods have been proposed to improve performance, but invasive technology that measures EEG and transmits the results to doctors has limitations such as high cost and difficulty in real life, making it less useful as a method for recommending surgery to the general public. Technology has also been developed that synthesizes voice using brain waves during speech, but this has limitations that limit its use in patients who have difficulty speaking or in quiet environments where it is difficult to speak. Therefore, there is a need for the development of new technology that can output avatar lip-sync animation by receiving as input biological signals when imagining speech, rather than relying on brain waves during speech and recorded voice data as a learning base. Summary of the Invention [Problem to be solved by the invention]
[0004] The present invention has been devised to overcome the limitations of the prior art as described above and in response to demands for new technological development. An object of the present invention is to provide an apparatus and method for generating avatar lip-sync animation based on multimodal bio-signals, which can receive as input bio-signals including electroencephalograms when a user imagines speaking, and output avatar lip-sync animation. [Means for solving the problem]
[0005] In order to achieve the above object, an apparatus for generating avatar lip-sync animation based on a multimodal bio-signal according to the present invention includes a multimodal data collection unit that collects multimodal data including bio-signal data including electroencephalograms when a user imagines speaking and video data; a pre-processing unit that pre-processes the multimodal data; a feature extraction unit that extracts feature vectors including bio-signal features and facial features of the user from the pre-processed multimodal data; an avatar generation unit that generates an avatar representing the user; a lip-sync reconstruction unit that inputs the extracted feature vectors into a pre-defined lip-sync reconstruction model to predict mouth shapes and facial movements when the user imagines speaking; the lip-sync reconstruction unit for the avatar generated by the avatar generation unit; and a lip-sync animation implementation unit that applies the mouth shapes and facial movements predicted by the lip-sync reconstruction unit to the avatar generated by the avatar generation unit to implement lip-sync animation of the avatar. Here, the avatar generation unit generates a 2D or 3D avatar from the user's video data using computer vision technology, maps the user's facial features extracted by the feature extraction unit to the generated avatar to specify facial landmarks, and the lip-sync animation implementation unit implements avatar lip-sync animation by applying the mouth shape and facial movement predicted by the lip-sync reconstruction unit to the avatar generated by the avatar generation unit based on coordinate values of the facial landmarks. In addition, the multimodal bio-signal-based avatar lip-sync animation generation device according to the present invention further includes a feature fusion unit that fuses the feature vectors extracted by the feature extraction unit and converts them into an embedded fusion vector, and the lip-sync reconstruction unit inputs the embedded fusion vector into a pre-established lip-sync reconstruction model to predict mouth patterns and facial movements when the user imagines speaking. Here, the multimodal data collection unit includes a prompt transmission display module that transmits prompts to the user for speech imagination, a biosignal collection module that measures biosignals including the user's brain waves and collects biosignal data, an image collection module that captures an image of the user's face and collects image data, and a data storage module that stores the biosignal data and image data of the user who imagines speech in response to the transmitted prompts together with a trigger value recorded over time. Here, the biosignal collection module further includes an electromyogram in the biosignals of the user to be measured, and the lip sync reconstruction unit predicts the mouth shape and facial movement by inferring the movement trajectory of articulators corresponding to the speech imagery based on the electromyogram. The feature fusion unit applies weights according to a predetermined criterion to the feature vectors extracted by the feature extraction unit, fuses the feature vectors to which the weights have been applied, and converts them into an embedded fusion vector. Here, the lip sync reconstruction model is characterized by being composed of either a first prediction model that predicts the mouth pattern and facial movement of the user when imagining how to speak from the extracted feature vector, or a second prediction model that grasps and classifies the user's intention from the extracted feature vector and predicts the mouth pattern and facial movement of the user when imagining how to speak according to the classified intention. In addition, a method for generating avatar lip-sync animation based on a multimodal bio-signal according to the present invention includes: a multimodal data collecting step of collecting multimodal data including bio-signal data including electroencephalograms when a user imagines speaking and video data; a preprocessing step of preprocessing the multimodal data; a feature extraction step of extracting feature vectors including bio-signal features and facial features of the user from the preprocessed multimodal data; an avatar generating step of generating an avatar representing the user's appearance based on facial features from the extracted feature vectors; a lip-sync reconstruction step of predicting a mouth shape and a facial movement when the user imagines speaking by inputting the extracted feature vectors into a pre-prepared lip-sync reconstruction model; and a lip-sync animation realizing step of realizing a lip-sync animation of the avatar by applying the mouth shape and facial movement predicted in the lip-sync reconstruction step to the avatar generated in the avatar generating step.
[0006] In addition, the multimodal bio-signal-based avatar lip-sync animation generation method according to the present invention further includes a feature fusion step of fusing feature vectors extracted in the feature extraction step to convert them into an embedded fusion vector, and the lip-sync reconstruction step is characterized in that the embedded fusion vector is input into a pre-established lip-sync reconstruction model to predict mouth and facial movements when the user imagines speaking. [Effects of the Invention]
[0007] With the above-described configuration, the multimodal bio-signal-based avatar lip-sync animation generating device and method according to the present invention has the advantage of being able to grasp a user's intention from bio-signals when the user imagines speaking and provide it as an avatar lip-sync animation. In addition, the multimodal bio-signal-based avatar lip-sync animation generating device and method according to the present invention can be developed into a system with unlimited applications by using bio-signals including non-invasive speech-imagination EEG. By extracting speech and facial reconstruction information contained in the bio-signals, it enables lip-sync animation that can visually convey a user's intentions without the user having to speak directly, and can express and convey a user's emotions and intentions through facial expressions. It has the advantage of being able to utilize avatars to enable factual and dynamic communication in the next-generation digital world, thereby achieving future-oriented technology. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 1 is a diagram showing the configuration of a multimodal bio-signal-based avatar lip-sync animation generating device according to one embodiment of the present invention. [Figure 2] FIG. 2 is a block diagram of a multimodal data collection unit according to an embodiment of the present invention. [Figure 3] FIG. 3 is a diagram illustrating an avatar generation unit according to an embodiment of the present invention. [Figure 4] FIG. 4 is a data collection and processing flow chart for lip sync reconstruction according to one embodiment of the present invention. [Figure 5] FIG. 5 is a data processing flowchart for realizing avatar lip-sync animation according to one embodiment of the present invention. [Figure 6] FIG. 6 is a flowchart of a method for generating lip-sync animation for an avatar based on a multimodal bio-signal according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0009] Hereinafter, a multimodal bio-signal-based avatar lip-sync animation generating apparatus and method according to the present invention will be described in more detail with reference to the embodiments shown in the drawings. FIG. 1 is a configuration diagram of a multimodal biosignal-based avatar lip-sync animation generation device according to one embodiment of the present invention; FIG. 2 is a configuration diagram of a multimodal data collection unit according to one embodiment of the present invention; FIG. 3 is an example diagram of an avatar generation unit according to one embodiment of the present invention; FIG. 4 is a data collection and processing flowchart for lip-sync reconstruction according to one embodiment of the present invention; and FIG. 5 is a data processing flowchart for realizing avatar lip-sync animation according to one embodiment of the present invention. Referring to FIG. 1, a multimodal bio-signal-based avatar lip-sync animation generating device according to an embodiment of the present invention includes a multimodal data collecting unit 10, a pre-processing unit 20, a feature extracting unit 30, an avatar generating unit 40, a feature fusing unit 50, a lip-sync reconstructing unit 60, and a lip-sync animation implementing unit 70. The multimodal data collection unit 10 is configured to collect multimodal data including biosignal data such as electroencephalograms and electromyograms when the user imagines speaking, and video data such as a video of the user's face. In one embodiment of the present invention, the multimodal data collection unit 10 may include a prompt transmission display module 11, a biosignal collection module 12, an image collection module 13, and a data storage module 14, as shown in FIG. The suggested sentence transmission and display module 31 is configured to transmit suggested sentences for utterance imagination to the user via a screen. In an embodiment of the present invention, the suggested sentence transmission and display module 31 may be configured to transmit a guideline image including a reference speech pattern for the suggested sentence to the user. The biosignal collecting module 12 is configured to measure biosignals such as electroencephalograms and electromyograms of a user and collect biosignal data. Electroencephalography (EEG), a biological signal, refers to the brain's electrical activity measured through electrodes attached to the scalp. These signals are an important tool for understanding various brain states and activities, particularly for their ability to precisely track functional changes over time. EEG is widely used in neuroscience research, clinical diagnosis, and neuropsychology, as well as the development of brain-computer interfaces. In particular, brain-computer interfaces recognize users' intentions and thoughts and convert them into machine commands, enabling users to control external devices using only their imagination, providing a new method of communication and interaction. Such EEG signal utilization technology can analyze brain activity patterns related to speech imagination and be applied to complex tasks such as real-time lip synchronization for digital avatars and animations. The biosignal collection module 12 for measuring the user's brain waves may be a wearable non-invasive device for measuring brain waves based on speech imagination, and in one embodiment, may be configured to measure real-time brain wave data based on the user's biosignals by wearing a cap-shaped device with a total of 128 electrodes attached to the outside of the scalp. A gel-like conductive material is applied to the scalp to match the electrodes so that brain waves can be measured effectively.
[0010] At this time, the biosignal collection module 12 for measuring the user's brain waves also measures brain waves based on speech attempts for analysis and comparison with brain waves based on speech ideation, and measures brain waves when only the mouth moves without making a sound. The measured brain waves are recorded along with a trigger value and the required time. The brain wave data is stored in a database path designated by the data storage module 14 (described later), and is preferably backed up to an external storage device for data storage. Meanwhile, electromyography (EMG) as a biosignal refers to electrical signals related to muscle activity and measures electrical changes caused by muscle contraction and relaxation. This signal plays an important role in evaluating the functional status of muscles and the health of the nervous system, and is widely used in medical diagnosis, rehabilitation, sports science, and bioepidemiology research. In particular, EMG measurement technology is extremely useful in research on muscle control by precisely monitoring muscle activity. It can be used to integrate natural human movements into digital avatars and robotics, and can reconstruct facial movements and facial expressions based on the user's specific muscle movements. It can also be used in a variety of fields, such as intuitive communication systems that synthesize speech based on the articulatory muscle movements of a speaking situation. The electromyogram measured as described above allows the lip sync reconstruction unit 60 (to be described later) to infer the movement trajectory of articulators corresponding to the speech image based on the electromyogram, thereby predicting the mouth shape and facial movement. Articulatory Kinematic Trajectories (AKTs) are precise records of articulatory movements during speech production. This plays an important role in understanding how articulatory organs, such as the lips, tongue, and jaw, move to produce sound. Analyzing articulatory kinematics is widely used in the fields of linguistics, phonetics, medicine, and computer science. For example, AKTs can be used to analyze the articulatory patterns of people with speech disorders and develop treatment methods. In addition, when combined with artificial intelligence technology, they are important for real-time lip-sync animation, sophisticated avatar expressions, and improving the accuracy of speech recognition systems. Articulatory kinematic trajectories provide highly detailed articulatory data, deepening our understanding of the human speech process and playing an essential role in developing more natural and realistic communication technologies based on this foundation. The image collection module 13 is configured to capture an image of the user's face and collect image data. The image collection module 13 records facial images of the user when he or she attempts to speak or imagines speaking using a camera attached to the display screen. Moving and still images are recorded in real time, and the same trigger value is recorded according to time to match the EEG data recording. The facial image data is stored in a database path designated by the data storage module 14, and is preferably backed up to an external storage device for data storage. The data storage module 14 stores the biosignal data and the video data of the user who is imagining an utterance in response to the transmitted prompt, together with trigger values recorded over time.
[0011] The preprocessing unit 20 is configured to preprocess the multimodal data. For example, the pre-processing of the EEG data recorded over a continuous period of time among the multimodal data can be performed by dividing it into trigger values recorded together, separating only the EEG data during actual speech attempts and imagination for approximately 1.5 seconds, and obtaining and saving frequency value information for the corresponding portion in a two-dimensional format of time and channel. The feature extraction unit 30 is configured to extract a feature vector including biosignal features and facial features of the user from the pre-processed multi-modal data. Here, the biosignal features may include electroencephalogram features and electromyogram features, and the facial features may include facial shape features, mouth shape features, and facial movement features. The feature extraction unit 30 generally uses common spatial pattern (CSP) or linear discriminant analysis (LDA) methods to extract feature vectors. These methods are used to accurately distinguish characteristics such as movement from biometric signals and to accurately obtain information on facial features and landmark features from facial image data. The extracted feature vectors are used in learning an artificial intelligence neural network model, which is beneficial for achieving high prediction performance. For example, the feature extraction unit 30 may be configured to extract phoneme-specific and mouth shape-specific movement feature vectors from the speech imagining EEG in the case of EEG features, extract a face feature vector in the case of face features, and then mutually fuse the phoneme-specific and mouth shape-specific movement feature vectors and the face feature vector in the feature fusion unit 50 described later to convert them into an embedded fusion vector. As shown in FIG. 3, the avatar generation unit 40 is configured to generate an avatar that represents the user. Here, avatar refers to a virtual character or image that serves as a representative or spokesperson for a user in the digital world. It primarily refers to a digital representative that reflects a user's physical, emotional, or personality traits in online games, virtual reality, and social media. Avatars come in a variety of shapes and characteristics that allow users to express and interact with the digital environment, expressing their unique identity and personality. Recently, with the development of artificial intelligence technology, avatars have become increasingly realistic, mimicking the user's real-time movements, facial expressions, and even speech. These developments enrich the experience of virtual reality and augmented reality and open up new possibilities in the fields of online communication and interactive entertainment. In one embodiment of the present invention, the avatar generation unit 40 is configured to generate a 2D or 3D avatar from the user's video data using computer vision technology, and to designate facial landmarks by mapping the user's facial features, particularly facial shape features, extracted by the feature extraction unit 30 onto the generated avatar.
[0012] Here, facial landmarks are points that indicate key facial features and are widely used in computer vision research and image processing. These landmarks refer to points corresponding to important facial features such as the eyes, nose, mouth, and jawline, and play a key role in various applications such as face recognition, emotion analysis, face tracking, lip sync, and avatar generation. Accurately identifying and tracking facial landmarks is essential for understanding and interpreting facial morphology and movement, and can even capture subtle changes in facial expression and eye movements. Recent advances in artificial intelligence technology have enabled more precise and rapid facial landmark detection, enabling innovative applications in various fields such as virtual reality, augmented reality, interactive games, and security systems. The avatar generation unit 40 can utilize open source software to generate avatars, such as Apple's Mimoticon avatars. It can also be applied to character avatars other than the user's personal face, in which case it is necessary to establish a correspondence between the landmark coordinates of the user's face landmarks and the new character avatar. The facial landmarks designated by the avatar generation unit 40 will be used as a reference for coordinate values in the lip-sync animation implementation unit 70, which will be described in detail later. The feature fusion unit 50 is configured to fuse the feature vectors extracted by the feature extraction unit 30 and convert them into an embedded fusion vector. In another embodiment, the feature fusion unit 50 may be configured to apply weights based on a predetermined criterion to the feature vectors extracted by the feature extraction unit 30, fuse the feature vectors to which the weights have been applied, and convert them into an embedded fusion vector. Here, the criteria for applying the weights may be selected by the user, or the multimodal data with high importance may be set in advance. The lip-sync reconstruction unit 60 is configured to input the extracted feature vector into a pre-established lip-sync reconstruction model to predict the mouth pattern and facial movement when the user imagines speaking. In one embodiment of the present invention, the lip sync reconstruction unit 60 may be configured to input an embedded fusion vector in which features are fused by the feature fusion unit 50 into a pre-established lip sync reconstruction model, and predict the mouth shape and facial movement when the user thinks about speaking. Such a lip sync reconstruction model can be implemented as a deep learning model such as a deep neural network (DNN), a convolutional neural network (CNN), or a recurrent neural network (RNN), or an autoencoder. In addition, in one embodiment of the present invention, the lip sync reconstruction model may be configured as either a first prediction model that immediately predicts the user's mouth shape and facial movement when imagining how to speak from the extracted feature vector, or a second prediction model that grasps and classifies the user's intention from the extracted feature vector and predicts the user's mouth shape and facial movement when imagining how to speak according to the classified intention. Meanwhile, the lip sync reconstruction unit 60 can be configured to accept only the user's brain waves as input values, which are biosignals measured when the user imagines speaking, or to accept the user's brain waves and other biosignals as multiple input values. For example, in the case of an electromyogram as another biosignal, the lip sync reconstruction unit 60 may be configured to infer the movement trajectory of articulators corresponding to speech imagery based on the electromyogram, thereby predicting mouth shape and facial movement.
[0013] In this case, when other biosignals such as electromyograms are used as input values together with the user's electroencephalograms, the feature fusion unit 50 is configured to apply weights corresponding to the importance of the biosignals used as input values or user settings to the feature vectors extracted by the feature extraction unit 30, fuse the weighted feature vectors, and convert them into an embedded fusion vector. The lip sync reconstruction unit 60 may be configured to input the embedded fusion vector obtained by fusing the weighted feature vectors and predict posture and facial movements corresponding to the speech idea. FIG. 4 shows a data collection and processing flow chart for lip sync reconstruction according to one embodiment of the present invention. Below, with reference to FIG. 4, we describe a methodology for generating an avatar and lip-sync animation to visually convey a user's intentions from non-invasive electroencephalogram data obtained during speech imagining. For lip-sync animation, mouth shapes and facial movements are reconstructed from speech-imagination EEG data, and then applied to a newly generated avatar from the user's facial image data to create a speaking avatar. Compared to the previously mentioned existing speech-imagination-based brain-computer interface technology, this technology is more scalable because it can regenerate mouth shape combinations corresponding to all human voices by decoding only about 15 Vismes that indicate mouth shapes, and can create moving mouth shapes for even undefined sentences. In addition, by utilizing about 15 Vismes-related features, the regeneration performance is also higher than that of existing EEG-based systems. For the speech imagery EEG, the speech attempt situation and speech imagery attempt are all performed while the EEG measurement device is worn. The speech attempt situation is simultaneously recorded for use in additional EEG data learning for speech imagery, and feature information on the words or sentences that are attempted to be spoken by moving the mouth without speaking is added. Learning of the phoneme units that make up words or sentences is prioritized, and the speech attempt situation and speech imagery attempt each take approximately 1.5 seconds. In addition, the actual mouth shape class when actually speaking a word or sentence is defined, and additional learning of the mouth shape unit is performed. The speech attempt situation and speech imagery attempt are each performed simultaneously for the same amount of time. This process is repeated n times to collect speech imagery EEG data for each user. EEG measurements are performed using a total of 128 channels, including the Broca-Wernicke region where speech-imagination EEG is characteristically expressed. Typically, EEG data is collected at a sampling frequency of 1,000 Hz, but this value may vary depending on the situation. While measuring EEG data, blinking and unintentional muscle movements may be recorded as noise, which is removed through a preprocessing process in the preprocessing unit 20. Time-dependent EEG data is divided based on different trigger values, allowing for the collection of data that can be used for learning. The final EEG data is extracted using time-frequency characteristics.
[0014] At the same time, while measuring the speech-imagination EEG data, a camera placed in front of the user records the user's face, obtaining facial features and video data. The system can be implemented using the video information or a single 2D photo provided by the user. Information on the mouth shape and facial movement during speech attempts can be obtained and used as additional features in the learning process. Approximately 68 facial landmarks can be obtained from the user's face recorded using facial video data, which will then be used to apply lip-sync animation to the avatar. The preprocessed data is converted into feature vectors that can represent mouth shapes and facial movements. EEG data is converted into a biosignal embedding vector through a biosignal encoder, and video data is converted into a facial feature embedding vector through a facial feature encoder. The embedding vectors are stored in a fused vector format for subsequent learning. The embedding vector conversion process serves to better represent the mouth shape characteristics for phonemes and erasures of speech images, and enables high-quality reconstruction by clearly distinguishing facial features. Information about articulatory movement is analyzed from EEG data to achieve lip sync. Based on existing research findings that articulator movement trajectories are encoded in the brain's speech sensorimotor cortex, the user's articulator movement can be inferred independently. 12-dimensional articulator movement trajectories can be analyzed, which correspond to x and y displacement values for various parts of the tongue and the upper and lower parts of the lips. The transformed embedding vectors are trained through a lip-sync reconstruction model to derive results for mouth shape or facial movement. The lip-sync reconstruction model for training allows for tuning parameters such as the learning rate and placement size to achieve optimal learning results. The movement change is realized by rearranging the position values in the form of a facial landmark trajectory and reconstructing the corresponding landmarks to realize a complex speaking face lip sync animation. The generated avatar has coordinates corresponding to N landmarks (N>10), and the coordinates are adapted to match the movement of the reconstructed landmarks, resulting in a natural avatar lip sync animation. This process is based on a model trained with collected user-customized data, and when a user uses the system, their brain waves and biosignals are given to the trained model in real time, and a moving face synthesized based on the given biosignals is displayed on the screen in real time. At this time, the user's avatar can be selected and applied in the way desired by the user using a personalized avatar generated from the user's facial image data or a 2D photo provided by the user. The lip-sync animation implementing unit 70 is configured to implement avatar lip-sync animation by applying the mouth shape and facial movement predicted by the lip-sync reconstruction unit to the avatar generated by the avatar generation unit. More specifically, the lip sync animation implementing unit 70 can implement avatar lip sync animation by applying the mouth shape and facial movement predicted by the lip sync reconstruction unit 60 to the avatar generated by the avatar generation unit 40 based on the coordinate values of the facial landmarks.
[0015] FIG. 5 shows a data processing flowchart for realizing avatar lip-sync animation according to one embodiment of the present invention. Referring to FIG. 5, the pre-processed bio-signal data is passed through a bio-signal encoder, a bio-signal embedding vector, and a lip-sync reconstruction decoder to generate mouth patterns and facial movements when the user imagines speaking, and the pre-processed video data is passed through a facial feature encoder, a facial feature embedding vector, and an avatar generation decoder to generate an avatar. The generated mouth patterns and facial movements and the avatar can be inter-combined to realize the final avatar lip-sync animation. The lip sync animation implementation unit 70 can actually implement the information derived from the lip sync reconstruction unit 60 based on the landmark coordinate values of the generated avatar, and the avatar can lip sync to say the suggested sentence as imagined by the user, and can implement an avatar animation that actually says the sentence by outputting the actual voice corresponding to the lip sync. The lip sync animation implementation unit 70 can be configured to display the result on a screen in real time and receive feedback so that the user can see the result and prove that the user's intention was accurate. The above describes the multimodal bio-signal-based avatar lip-sync animation generating device according to the present invention, and the following describes the multimodal bio-signal-based avatar lip-sync animation generating method according to the present invention. FIG. 6 is a flowchart of a method for generating lip-sync animation for an avatar based on a multimodal bio-signal according to an embodiment of the present invention. Referring to FIG. 6, a method for generating avatar lip-sync animation based on multimodal bio-signals according to one embodiment of the present invention includes a multimodal data collection step (S10), a pre-processing step (S20), a feature extraction step (S30), an avatar generation step (S40), a feature fusion step (S50), a lip-sync reconstruction step (S60), and a lip-sync animation implementation step (S70). The specific content of each step has been described above in the multimodal bio-signal-based avatar lip-sync animation generation device according to the present invention, so below we will only briefly explain the characteristics of each step.
[0016] The multimodal data collecting step (S10) is a step of collecting multimodal data including biosignal data including electroencephalograms and video data when the user imagines speaking. The pre-processing step (S20) is a step of pre-processing the multi-modal data. The feature extraction step (S30) is a step of extracting a feature vector including biosignal features and facial features of the user from the pre-processed multi-modal data. The avatar generating step (S40) is a step of generating an avatar that represents the user's appearance based on facial features from the extracted feature vectors. The feature fusion step (S50) is a step of fusing the feature vectors extracted in the feature extraction step to convert them into an embedded fusion vector. The lip-sync reconstruction step (S60) is a step of inputting the extracted feature vector into a pre-established lip-sync reconstruction model to predict the mouth shape and facial movement when the user imagines speaking. In one embodiment of the present invention, the lip-sync reconstruction step (S60) may be configured to input the embedding fusion vector in which the features are fused in the feature fusion step (S50) into a pre-established lip-sync reconstruction model to predict the mouth shape and facial movement when the user imagines speaking. The lip-sync animation implementation step (S70) is a step of implementing avatar lip-sync animation by applying the mouth shape and facial movement predicted in the lip-sync reconstruction step (S60) to the avatar generated in the avatar generation step (S40). The multimodal biosignal-based avatar lip-sync animation generating apparatus and method according to the present invention, having the above-described configuration, offers innovative applicability in the fields of medicine, assistive devices, and everyday communication. In particular, it provides a customized solution for individuals with physical or language limitations, allowing them to express their thoughts in a direct and efficient manner. This technology can be integrated with various interfaces, such as voice generating devices, virtual and augmented reality systems, and next-generation communication systems, and can be used to interpret a user's thoughts in real time and control various digital devices. This brain-computer interface technology could also serve as a rehabilitation training tool for stroke and severe muscle injuries, marking a major turning point in health management and overcoming disabilities. Furthermore, because this technology can be applied to general mass-market applications such as education, entertainment, and improving personal productivity, it is expected to become deeply integrated into everyday life as technology advances, making it accessible to everyone. The multimodal bio-signal-based avatar lip-sync animation generating apparatus and method described above and shown in the drawings is merely one embodiment for carrying out the present invention and should not be construed as limiting the technical spirit of the present invention. The scope of protection of the present invention is defined solely by the matters set forth in the following claims, and improvements and modifications made without departing from the spirit of the present invention are also within the scope of protection of the present invention as long as they are obvious to those skilled in the art to which the present invention pertains. [Explanation of symbols]
[0017] 10 Multimodal Data Collection Unit 20 Pretreatment section 30 Feature Extraction Unit 40 Avatar Generation Unit 50 Feature Fusion Section 60 Lip Sync Reconstruction Unit 70 Lip Sync Animation Department
Claims
1. a multimodal data collection unit that collects multimodal data including biosignal data including electroencephalograms when the user imagines speaking and video data; a preprocessing unit that preprocesses the multimodal data; a feature extraction unit that extracts a feature vector including biosignal features and facial features of the user from the preprocessed multimodal data; an avatar generation unit that generates an avatar that represents the user; a lip sync reconstruction unit that inputs the extracted feature vector into a prepared lip sync reconstruction model and predicts the mouth shape and facial movement when the user imagines speaking; a lip sync animation implementation unit that implements avatar lip sync animation by applying the mouth shape and facial movement predicted by the lip sync reconstruction unit to the avatar generated by the avatar generation unit.
2. 2. The multimodal bio-signal-based avatar lip-sync animation generating device of claim 1, wherein the avatar generation unit generates a 2D or 3D avatar from video data of the user using computer vision technology, maps the user's facial features extracted by the feature extraction unit to the generated avatar to specify facial landmarks, and the lip-sync animation implementing unit implements avatar lip-sync animation by applying granularity and facial movements predicted by the lip-sync reconstruction unit to the avatar generated by the avatar generation unit based on coordinate values of the facial landmarks.
3. 2. The multimodal bio-signal-based avatar lip-sync animation generation device according to claim 1, further comprising a feature fusion unit that fuses the feature vectors extracted by the feature extraction unit and converts them into an embedded fusion vector, and the lip-sync reconstruction unit inputs the embedded fusion vector into a pre-established lip-sync reconstruction model to predict mouth shapes and facial movements when a user imagines speaking.
4. The multimodal data collection unit includes a prompt sentence transmission / display module that transmits a prompt sentence for a user to imagine an utterance; a biosignal collection module that measures biosignals including brain waves of a user to collect biosignal data; and an image collection module that captures a facial image of the user to collect image data; and a data storage module that stores the biosignal data and image data of the user imagining speech in response to the transmitted prompt, together with a trigger value that is recorded over time.
5. 5. The multimodal bio-signal-based avatar lip-sync animation generation device according to claim 4, wherein the bio-signal collection module further includes an electromyogram in the bio-signals of the user to be measured, and the lip-sync reconstruction unit predicts mouth patterns and facial movements by inferring articulator movement trajectories corresponding to speech imagery based on the electromyogram.
6. 4. The multimodal bio-signal-based avatar lip-sync animation generation device according to claim 3, wherein the feature fusion unit applies weights according to a predetermined criterion to the feature vectors extracted by the feature extraction unit, fuses the feature vectors to which the weights have been applied, and converts them into an embedded fusion vector.
7. the lip sync reconstruction model is a first prediction model that immediately predicts the mouth shape and facial movement of the user when imagining speech from the extracted feature vector; or a second prediction model that grasps and classifies the user's intention from the extracted feature vector and predicts the mouth shape and facial movement when the user imagines speaking according to the classified intention.
8. a multimodal data collection step of collecting multimodal data including biosignal data including electroencephalograms and video data when the user imagines speaking; a preprocessing step of preprocessing the multimodal data; and a feature extraction step of extracting a feature vector including biosignal features and facial features of the user from the preprocessed multimodal data. an avatar generating step of generating an avatar representing the user's appearance based on facial features among the extracted feature vectors; and an avatar generating step of generating an avatar representing the user's appearance based on facial features among the extracted feature vectors. A method for generating avatar lip-sync animation, comprising: a lip-sync animation realization step of realizing lip-sync animation of the avatar by applying the sphere and facial movement predicted in the lip-sync reconstruction step to the avatar generated in the avatar generation step.
9. 9. The multimodal bio-signal-based avatar lip-sync animation generation method according to claim 8, further comprising: a feature fusion step of fusing the feature vectors extracted in the feature extraction step to convert them into an embedded fusion vector; and the lip-sync reconstruction step of inputting the embedded fusion vector into a pre-established lip-sync reconstruction model to predict mouth shapes and facial movements when the user imagines speaking.
Citation Information
Patent Citations
Wearable terminal device and program
JP2016126500A
Root crops harvester
KR1020250179460A
DC series arc filure diagnosis apparatus using artificial machine learning
KR102906080B1
Real-time 3D facial animation from binocular video
US20220358719A1