A method and apparatus for generating a sequence of lip images, and an electronic device
By splitting the teeth and lip regions of the original face image into independent masks and using spatial attention to guide the generation of lip image sequences, the problem of insufficient personalization of virtual avatar lip animation in traditional methods is solved, thus improving the user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- IFLYTEK CO LTD
- Filing Date
- 2026-02-06
- Publication Date
- 2026-06-12
Smart Images

Figure CN122199755A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, apparatus and electronic device for generating lip image sequences. Background Technology
[0002] With the development of AIGC (Artificial Intelligence Generated Content) technology, the application of virtual avatars is penetrating all aspects of social life at an unprecedented speed and breadth. However, in voice-driven virtual avatar lip-syncing animation tasks, traditional methods only focus on the accuracy of audio-visual synchronization, making it difficult to meet the personalized needs of virtual avatar lip-syncing animation and resulting in a poor user experience. Summary of the Invention
[0003] To address the aforementioned issues, this application proposes a method, apparatus, and electronic device for generating lip image sequences, thereby enabling personalized generation of virtual avatar lip-shape animations and enhancing user experience.
[0004] The first aspect of this application provides a method for generating a lip image sequence, including: Acquire speech information, original face image, and lip feature information indicating target lip features for driving the generation of virtual avatar lip-splitting animation, wherein the target lip features include at least one of target tooth features and target lip features; Determine the lip region mask of the original face image, wherein the lip region mask includes a teeth region mask and a lip region mask; Using the lip region mask as spatial attention guide, a sequence of lip images conforming to the target lip features is generated based on the speech information and the original face image.
[0005] In one possible implementation, determining the lip region mask of the original face image includes: Key point detection is performed on the original face image to determine the initial lip region mask of the original face image; Based on the initial lip region mask, tooth detection is performed on the lip region of the original face image to determine the tooth region and lip region within the lip region, and a lip region mask is generated.
[0006] In one possible implementation, the step of performing keypoint detection on the original face image to determine the initial lip region mask of the original face image includes: A set of facial key points is obtained by performing key point detection on the original face image. Based on the perioral contour points in the set of facial key points, the lip region of the original face image is determined and an initial lip region mask is generated.
[0007] In one possible implementation, the step of performing tooth detection on the lip region of the original face image based on the initial lip region mask, determining the tooth region and lip region within the lip region, and generating a lip region mask includes: Based on the inner lip contour key points in the set of facial key points, determine the oral cavity opening region in the lip region of the original facial image; The teeth are detected in the oral cavity opening area of the original face image, and the tooth area and lip area in the lip area are determined according to the initial lip area mask to generate a lip area mask; the lip area includes the non-tooth area in the lip area.
[0008] In one possible implementation, the step of using the lip region mask as spatial attention guidance to generate a sequence of lip images conforming to the target lip features based on the speech information and the original face image includes: A joint condition is constructed based on the speech information, the lip feature information, and the time axis of the lip image sequence. The joint conditions, the image features of the original face image, and the lip region mask are used as guiding conditions for the lip image sequence generation model, so that the lip image sequence generation model generates a lip image sequence that conforms to the target lip features.
[0009] In one possible implementation, constructing joint conditions based on the speech information, the lip feature information, and the timeline of the lip image sequence includes: Extract the speech features of the speech frames in the speech information, and map the speech features of the speech frames to the time axis of the lip image sequence to generate the speech features of the speech information; The lip feature information is encoded to generate semantic features of the lip feature information; Based on the speech features of the speech information, the semantic features of the lip features, and the time axis of the lip image sequence, a joint condition is constructed.
[0010] In one possible implementation, the step of using the joint condition, the image features of the original face image, and the lip region mask as guiding conditions for the lip image sequence generation model to generate a lip image sequence that conforms to the target lip features includes: The joint conditions, the image features of the original face image, and the lip region mask are used as guiding conditions for the lip image sequence generation model. During the generation of the lip image sequence, the model uses the tooth region mask as a first spatial guide to ensure that the tooth features of the lip image conform to the target tooth features, and uses the lip region mask as a second spatial guide to ensure that the lip features of the lip image conform to the target lip features, thus obtaining the lip image sequence.
[0011] One possible implementation also includes: During the generation of the lip image sequence, a target historical lip image that matches the pronunciation of the current lip image to be generated is determined from at least one recently generated historical lip image, and the image features of the target historical lip image are used as a reference to generate the current lip image.
[0012] A second aspect of this application provides an apparatus for generating a lip image sequence, comprising: The information acquisition unit is used to acquire voice information, original face image and lip feature information indicating target lip features for driving the generation of virtual image lip-shaped animation, the target lip features including at least one of target tooth features and target lip features; The lip region mask determination unit is used to determine the lip region mask of the original face image, wherein the lip region mask includes a teeth region mask and a lip region mask; The lip image sequence generation unit is used to generate a lip image sequence that conforms to the target lip features based on the speech information and the original face image, using the lip region mask as spatial attention guidance.
[0013] A third aspect of this application provides an electronic device, including a memory and a processor; The memory is connected to the processor and is used to store programs; The processor is configured to implement the lip image sequence generation method as described in the first aspect of this application or any possible implementation thereof by running a program in the memory.
[0014] The fourth aspect of this application provides a chip including a processor and a data interface, wherein the processor reads and runs a program stored in a memory through the data interface to perform the lip image sequence generation method as described in the first aspect of this application or any possible implementation thereof.
[0015] The fifth aspect of this application provides a computer program product, including a computer program that, when executed by a processor, implements the lip image sequence generation method as described in the first aspect of this application or any possible implementation thereof.
[0016] The sixth aspect of this application provides a storage medium storing a computer program that, when executed by a processor, implements the lip image sequence generation method as described in the first aspect of this application or any possible implementation thereof.
[0017] This application proposes a method, apparatus, and electronic device for generating lip image sequences. The method involves acquiring speech information, an original face image, and lip feature information indicating target lip features for driving virtual avatar lip animation generation. The target lip features include at least one of target tooth features and target lip features. A lip region mask is determined from the original face image, including a tooth region mask and a lip region mask. The lip region mask is used as a spatial attention guide to generate a lip image sequence conforming to the target lip features based on the speech information and the original face image. This application separates the tooth and lip regions of the original face image into two independent masks: a tooth region mask and a lip region mask. This allows for independent control of the tooth and lip regions in the generated lip image sequence, effectively achieving personalized generation of virtual avatar lip animation and enhancing the user experience. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0019] Figure 1 This is a schematic diagram of an implementation system architecture for the lip image sequence generation method provided in this application.
[0020] Figure 2 This is a flowchart of a method for generating lip image sequences provided in an embodiment of this application.
[0021] Figure 3 A flowchart illustrating a method for determining the lip region mask of an original human face image, provided in an embodiment of this application.
[0022] Figure 4This application provides a flowchart of a method for generating a sequence of lip images that conforms to target lip features by using a lip region mask as spatial attention guidance based on speech information and the original face image.
[0023] Figure 5 This is a schematic diagram of a lip image sequence generation device provided in an embodiment of this application.
[0024] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0025] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0026] In voice-driven virtual avatar lip-syncing animation tasks, traditional methods only focus on the accuracy of audio-visual synchronization, which makes it difficult to meet the personalized needs of virtual avatar lip-syncing animation and results in a poor user experience.
[0027] To solve this technical problem, the inventors of this application discovered through research that in virtual character lip-syncing animation tasks, the lip area can be considered as a whole, and personalized generation of the entire lip area of the virtual character can be achieved by taking the lip area as a unit.
[0028] However, further research by the inventors revealed that while this method can achieve personalized generation of lip-sync animation for virtual characters to some extent, the personalized generation effect is still unsatisfactory. For example, it cannot achieve differentiated processing of the teeth and lips in the mouth area of the virtual character.
[0029] To further address the aforementioned issues, this application proposes a method, apparatus, and electronic device for generating lip image sequences. The method involves acquiring speech information used to drive the generation of virtual avatar lip-shape animation, an original face image, and lip feature information indicating target lip features. The target lip features include at least one of target tooth features and target lip features. A lip region mask is determined from the original face image, including a tooth region mask and a lip region mask. The lip region mask is used as a spatial attention guide to generate a lip image sequence conforming to the target lip features based on the speech information and the original face image. This application separates the tooth and lip regions of the original face image into two independent masks: a tooth region mask and a lip region mask. This allows for independent control of the tooth and lip regions in the lip image to be generated during the generation of the lip region image sequence, effectively achieving personalized generation of virtual avatar lip-shape animation and enhancing the user experience.
[0030] Exemplary Implementation Environment It should be understood that the lip image sequence generation method provided in this application can be applied to systems or programs that include virtual avatar lip-shape animation generation functionality (virtual avatar animation generation system or program), which typically have virtual avatar lip-shape animation generation functionality. Specifically, the virtual avatar animation generation system can run on systems such as... Figure 1 In the system architecture shown, such as Figure 1 The diagram shown is a system architecture diagram of the virtual character animation generation system. This system architecture may include a terminal 100 and a server 200.
[0031] In this embodiment, the terminal 100 can be a mobile phone, tablet computer, learning machine, teaching large screen, wearable device, augmented reality (AR) / virtual reality (VR) device, laptop computer, ultra-mobile personal computer (UMPC), netbook, personal digital assistant (PDA), etc. This embodiment does not impose any restrictions on it.
[0032] It is understandable that server 200 can include one or more servers. Figure 1 (This example uses a server as an illustration).
[0033] In this embodiment, server 200 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.
[0034] Terminal 100 and server 200 can be directly or indirectly connected via wired or wireless communication. Terminal 100 and server 200 can be connected to form a blockchain network, and this application does not impose any restrictions on this.
[0035] Either terminal 100 or server 200 can be used independently to execute the lip image sequence generation method provided in the embodiments of this application.
[0036] In addition, the terminal 100 and the server 200 can also work together to execute the lip image sequence generation method provided in the embodiments of this application.
[0037] It is understood that the aforementioned virtual character animation generation system can run as a program on the aforementioned device, or as a system component of the aforementioned device, or as a cloud service program. The specific operating mode depends on the actual scenario and is not limited here.
[0038] Exemplary methods This application proposes a method for generating a lip image sequence. This method can be executed by an electronic device, which can be any device with data and instruction processing capabilities, such as a computer, smart terminal, or server. See also... Figure 2 The method includes: S201. Obtain speech information, original face image and lip feature information indicating target lip features for driving the generation of virtual image lip-shaped animation, the target lip features including at least one of target tooth features and target lip features; An electronic device runs a virtual avatar animation generation program, which includes a virtual avatar lip-shape animation generation function. The lip image sequence generation method provided in this application is applied to the virtual avatar lip-shape animation generation function.
[0039] In this embodiment of the application, obtaining the voice information, the original face image, and the lip feature information indicating the target lip features used to drive the generation of virtual character lip-syncing animation includes: responding to a virtual character animation generation request and obtaining the voice information, the original face image, and the lip feature information indicating the target lip features indicated by the virtual character animation generation request.
[0040] For example, an electronic device runs a virtual avatar animation generation program. This program displays a personalized settings interface to the user. The personalized settings interface includes a first information fill position, a second information fill position, and a third information fill position. The first information fill position receives voice information input by the user to drive the virtual avatar's lip-sync animation generation. The second information fill position receives an original face image input by the user to drive the virtual avatar's lip-sync animation generation. The third information fill position receives lip feature information input by the user, which indicates a target lip feature. It can be understood that the lip feature indicated by the lip feature information is called the target lip feature.
[0041] After entering voice information in the first information fill position, the original face image in the second information fill position, and lip feature information in the third information fill position in the personalization settings interface, the user clicks the target button on the personalization settings interface to send a virtual avatar animation generation request to the virtual avatar animation generation program. The virtual avatar animation generation program responds to the virtual avatar animation generation request, executes a lip image sequence generation method provided in this application embodiment to generate a lip image sequence corresponding to the virtual avatar animation generation request, and constructs a virtual avatar animation based on the lip image sequence, and then displays the virtual avatar animation to the user through the electronic device screen.
[0042] It should be noted that the voice information used to drive the virtual avatar lip-sync animation generation in the virtual avatar animation generation request is the voice information entered by the user at the first information fill position, the original face image used to drive the virtual avatar lip-sync animation generation in the virtual avatar animation generation request is the original face image entered by the user at the second information fill position, and the lip feature information used to drive the virtual avatar lip-sync animation generation in the virtual avatar animation generation request is the lip feature information entered by the user at the third information fill position.
[0043] Voice information is used to drive changes in lip movements over time.
[0044] The original face image is a face image that matches the facial features.
[0045] For example, the original face image can be a real face image or a virtual face image that conforms to facial features; this application embodiment does not limit the scope.
[0046] Lip feature information refers to personalized lip settings. Lip feature information includes tooth feature information and lip feature information. Tooth feature information indicates the target tooth features, and lip feature information indicates the target lip features.
[0047] For example, lip feature information can be lip feature information described using natural language, such as "white teeth", "M-shaped upper lip", "full smiling lips", etc.
[0048] The above are merely exemplary contents of the lip feature information provided in the embodiments of this application. The specific contents of the lip feature information are not limited in the embodiments of this application.
[0049] Understandably, virtual character lip-syncing animation consists of a sequence of lip images.
[0050] In this embodiment of the application, the target tooth features characterize the personalized tooth features that the tooth regions in the lip image sequence that the user expects to generate should conform to.
[0051] For example, the target tooth features include at least one of the prior features of tooth structure and the target color features of tooth.
[0052] As a preferred implementation of this application, the prior features of tooth structure include tooth structure regularization terms.
[0053] Adding a tooth structure regularization term to the loss function of the lip image sequence generation model can constrain the physiological rationality of the tooth regions in the lip image sequence during generation, forcibly maintaining the integrity and regularity of tooth arrangement, and preventing blurring, breakage, or misalignment of tooth regions in the generated lip image sequence. The tooth structure regularization term is: in, For the prior constraint loss of tooth structure These are the structural regularization weight coefficients. To predict the tooth region in the generated lip image, Gradient features of the tooth region, A predefined tooth structure prior gradient template.
[0054] For example, the tooth structure prior gradient template includes a vertical stripe gradient template. The vertical stripe gradient template is a prior structure guide map used to explicitly constrain the texture orientation and spatial distribution of the tooth region during the generation of lip images by the lip image sequence generation model, so that the tooth region of the generated lip image presents a natural and neat tooth shape.
[0055] In this embodiment of the application, the target lip features represent the personalized lip features that the lip region in the lip image sequence that the user expects to generate should conform to.
[0056] For example, the target lip features include at least one of the following: lip target shape features, lip target texture features, and lip target color features.
[0057] It should be noted that the dental feature information and lip feature information in the oral feature information are two independent parts. The dental feature information is only used to describe the target dental features related to teeth, and the lip feature information is only used to describe the target lip features related to lips.
[0058] Correspondingly, tooth feature information is used to constrain the personalized generation of the tooth region in the lip image sequence, and lip feature information is used to constrain the personalized generation of the lip region in the lip image sequence. Based on tooth feature information and lip feature information, the tooth region and lip region in the lip image sequence can be controlled separately and independently. This allows users to personalize the tooth region and lip region of the virtual character animation according to their actual needs during the virtual character animation generation process, while also achieving differentiated control over the tooth region and lip region in the virtual character animation, thus improving the user experience.
[0059] S202. Determine the lip region mask of the original face image. The lip region mask includes the teeth region mask and the lip region mask. In this embodiment, the original face image is detected to determine the tooth region mask and lip region mask of the original face image; during the generation of the lip image sequence, personalized control of the tooth region in the lip image is realized based on the tooth region mask, and personalized control of the lip region in the lip image is realized based on the lip region mask, so as to generate a lip image sequence that conforms to both the target tooth features and the target lip features, thereby improving the user experience.
[0060] Figure 3 This is a flowchart illustrating a method for determining a lip region mask in an original face image, as provided in an embodiment of this application. Figure 3 As shown, the method includes: S301. Perform key point detection on the original face image to determine the initial lip region mask of the original face image; In this embodiment of the application, key point detection is performed on the original face image to determine the initial lip region mask of the original face image, including: performing key point detection on the original face image to obtain a set of face key points; determining the lip region of the original face image and generating an initial lip region mask based on the perioral contour points in the set of face key points.
[0061] High-density keypoint detection networks (such as DECA and HRNet-Face) are used to detect keypoints in the original face image, and facial keypoints are extracted from the original face image to obtain a set of facial keypoints.
[0062] Determine the perioral contour points in the facial key point set, and form a closed polygonal region based on the perioral contour points. The polygonal region represents the mouth region in the original face image. Use the edge filling algorithm to generate the initial lip region mask.
[0063] As can be understood, the polygonal region represents the entire mouth area in the original face image. That is, the mouth area (also known as the lip area) in the original face image.
[0064] As one implementation of this application, each perioral contour point in the set of facial key points is determined, and each perioral contour point is connected in a counterclockwise direction to form a closed polygonal contour; the area of the original face image within the polygonal contour represents the lip region in the original face image, and the lip region in the original face image is marked using a masking method to generate an initial lip region mask.
[0065] Understandably, the area within the polygonal outline of the original face image can be called the polygonal region.
[0066] For example, the perioral contour points in the facial key point set include key points at the corners of the mouth, key points at the cupid's bow, and key points at the philtrum. The above is merely a preferred content of the perioral contour points provided in the embodiments of this application. Those skilled in the art can set the specific content of the perioral contour points according to their own needs, and the embodiments of this application do not limit it.
[0067] S302. Based on the initial lip region mask, perform tooth detection on the lip region of the original face image to determine the tooth region and lip region in the lip region, and generate the lip region mask.
[0068] In this embodiment of the application, tooth detection is performed on the lip region of the original face image based on the initial lip region mask to determine the tooth region and lip region in the lip region, and a lip region mask is generated. This includes: determining the oral cavity opening region in the lip region of the original face image based on the inner lip contour key points in the face key point set; performing tooth detection on the oral cavity opening region of the original face image, and determining the tooth region and lip region in the lip region based on the initial lip region mask to generate a lip region mask; the lip region includes the non-tooth region in the lip region.
[0069] The process involves detecting teeth within the oral cavity opening region of the original face image's lip region to identify the tooth area. The non-tooth areas within the lip region are then identified as the lip region. Subsequently, tooth region masks and lip region masks are generated. The lip region mask of the original face image comprises two independent masks: a tooth region mask and a lip region mask.
[0070] Understandably, after determining the tooth region in the original face image, the lip region in the original face image can be determined based on the initial lip region mask, and the non-tooth region in the lip region of the original face image can be determined as the lip region in the original face image.
[0071] As a preferred implementation of this application, the remaining area (non-tooth area) in the lip area other than the tooth area can be defined as the lip area in the lip area.
[0072] In this embodiment, a tooth detection model is pre-set, and the tooth region in the oral cavity opening area of the original face image is detected by the tooth detection model to determine the tooth region in the oral cavity area of the original face image.
[0073] The above are merely preferred implementations of tooth detection provided in the embodiments of this application. Those skilled in the art can set specific implementations of tooth detection according to their own needs, and the embodiments of this application do not limit them.
[0074] In this embodiment, the tooth region mask and the lip region mask will be used as spatial attention guidance conditions in the lip image sequence generation model, so that when the lip image sequence generation model generates lip image sequences, it will apply different control strategies to the tooth region and lip region in the lip image respectively, thereby realizing personalized generation of teeth and lips in the lip image sequence.
[0075] S203. Using the lip region mask as a spatial attention guide, generate a sequence of lip images that conform to the target lip features based on the speech information and the original face image.
[0076] Figure 4 This application provides a flowchart of a method for generating a sequence of lip images that conforms to target lip features by using a lip region mask as spatial attention guidance based on speech information and the original face image. Figure 4 As shown, the method includes: S401. Construct joint conditions based on speech information, lip feature information, and the time axis of lip image sequence; In this embodiment, the joint conditions are constructed based on speech information, lip feature information, and the time axis of the lip image sequence, including: extracting speech features from speech frames in the speech information and mapping the speech features of the speech frames to the time axis of the lip image sequence to generate speech features of the speech information; encoding the lip feature information to generate semantic features of the lip feature information; and constructing joint conditions based on the speech features of the speech information, the semantic features of the lip feature information, and the time axis of the lip image sequence.
[0077] In this embodiment, three modalities (speech information, lip feature information, and the timeline of the lip image sequence) are integrated to generate joint conditions. These joint conditions serve as the conditional input to the lip image sequence generation model, enabling collaborative generation driven by speech, guided by text (the text being lip feature information), and time-aware.
[0078] The time axis of the lip image sequence includes the temporal position encoding of each frame of the predicted lip image sequence within the lip image sequence.
[0079] In this embodiment, a pre-trained speech model (such as HuBERT or Wav2Vec 2.0) is used to extract the speech features of the speech frames in the speech information, and the speech features of the speech frames are mapped to the time axis of the lip image sequence by forced alignment to obtain the speech features of the speech information.
[0080] It is understandable that the speech features of speech information are composed of the speech features of speech frames in the speech information, and the speech features of speech frames are mapped to the time axis of the lip image sequence.
[0081] For example, the speech features of a speech frame are the frame-level implicit representation of the speech frame. ,in, Encode the temporal position of the lip image corresponding to the speech features of the speech frame in the time axis of the lip image sequence.
[0082] In this embodiment of the application, lip feature information can be considered as an aesthetic intention instruction, which usually exists in the form of text.
[0083] The lip feature information is encoded using a text encoder to generate semantic features of the lip feature information. , Represents the user's desired aesthetic style (target lip features).
[0084] Based on the speech features of the speech information, the semantic features of the lip features, and the time axis of the lip image sequence, a joint condition C is constructed. : C in, Indicates the first Frame of lip images, Indicates the relationship with the first Speech features of speech frames aligned with lip images. Semantic features representing lip features.
[0085] for , indicating the first in the lip image sequence Frame of lip images, Encode the temporal location of the lip image. It can be viewed as a timeline of lip image sequences, used to enhance the time perception capability of lip image sequences.
[0086] S402. The joint conditions, the image features of the original face image, and the lip region mask are used as guiding conditions for the lip image sequence generation model, so that the lip image sequence generation model generates a lip image sequence that conforms to the target lip features.
[0087] In this embodiment, the joint conditions, the image features of the original face image, and the lip region mask are used as guiding conditions for the lip image sequence generation model. This allows the lip image sequence generation model to use the tooth region mask as the first spatial guide during the lip image sequence generation process, so that the tooth features of the lip image conform to the target tooth features, and the lip region mask as the second spatial guide, so that the lip features of the lip image conform to the target lip features, thus obtaining the lip image sequence.
[0088] In generating a sequence of lip images, the model uses a tooth region mask as a first spatial guide to ensure that the tooth features in the lip images match the target tooth features, and a lip region mask as a second spatial guide to ensure that the lip features in the lip images match the target lip features. This results in the generated sequence of lip images having both tooth and lip regions that match the target tooth and lip features. Because the lip region in a lip image consists of both tooth and lip regions, the lip region in the generated sequence of lip images thus matches the target lip features.
[0089] Understandably, in the process of generating lip images in a lip image sequence using a lip image sequence generation model, a target lip region mask matching the current lip image to be generated is determined. The target lip region mask includes a target tooth region mask and a target lip region mask. The target tooth region mask is used as a first spatial guide to ensure that the tooth features of the current lip image to be generated match the target tooth features. The target lip region mask is used as a second spatial guide to ensure that the lip features of the current lip image to be generated match the target lip features, thus obtaining the current frame lip image.
[0090] The target lip region mask matching the current lip image to be generated is obtained by updating the lip region mask of the original face image based on the speech features of the speech frame corresponding to the current lip image to be generated. The updated lip region mask includes an updated teeth region mask and an updated lip region mask. The updated teeth region mask is called the target teeth region mask, and the updated lip region mask is called the target lip region mask.
[0091] In this embodiment, the lip image sequence generation model can be a diffusion model. Based on the Latent Diffusion Model as the fundamental generation architecture for the lip image sequence generation model, the following improvement mechanism is designed: (1) Base map preservation mechanism The original face image is encoded to obtain its image features. During the entire denoising process, the pixels in the non-lip region of the original face image are fixed, and only the lip region is allowed to be updated.
[0092] (2) Dual-area attention guidance A dual-channel mask attention mechanism is introduced into the cross-attention layer of U-Net. Specifically, during the generation of the lip image sequence, target tooth features are injected into the tooth region, and target lip features are emphasized in the lip region.
[0093] (3) Joint guidance of lip feature information and voice information C The input to the U-Net's Cross-Attention layer allows each denoising step in the generation of the lip image sequence to be simultaneously modulated by the aesthetic intent indicated by the speech rhythm and lip feature information of the speech information.
[0094] It is understandable that lip feature information can be considered as an aesthetic intention instruction, and the target lip feature indicated by the lip feature information can be considered as the aesthetic intention indicated by the lip feature information.
[0095] (4) Prior constraints on tooth structure It should be noted that when the target tooth features include prior features of tooth structure, adding a tooth structure regularization term to the loss function of the lip image sequence generation model can prevent the tooth regions in the generated lip image sequence from being blurred, broken, or misaligned.
[0096] In this embodiment, based on the image features of the original face image and the lip region mask (including the tooth region mask and the lip region mask), during the diffusion generation process, the lip image sequence generation model uses joint conditions as the driving signal to gradually denoise and generate lip images in the latent space. The tooth region mask and the lip region mask serve as spatial guides, and a region-aware attention mechanism is used to control the influence intensity on the tooth and lip regions of the lip image in the lip image sequence, thereby achieving refined replay of the virtual avatar's lip region and generating a lip image sequence that is synchronized with the speech information and conforms to the target lip features.
[0097] Furthermore, the lip image sequence generation method provided in this application embodiment further includes: in the process of generating the lip image sequence, determining a target historical lip image that matches the pronunciation corresponding to the current lip image to be generated from at least one recently generated historical lip image, and using the image features of the target historical lip image as a reference to generate the current lip image.
[0098] For example, taking the process of generating a lip image sequence as an example, in the process of generating the current lip image to be generated, at least one historical lip image that was most recently generated in the process of generating the lip image sequence is determined, a target historical lip image that matches the pronunciation of the current lip image to be generated is determined from the at least one historical lip image, and the image features of the target historical lip image are used as reference information for generating the current lip image to be generated, so as to generate the current frame lip image.
[0099] Accordingly, at least one historical lip image is all historical lip images, or at least one historical lip image is three historical lip images, or at least one historical lip image is five historical lip images, etc.
[0100] The above is merely a preferred number of at least one historical lip image provided in the embodiments of this application. Those skilled in the art can set the specific number of at least one historical lip image according to their own needs, and the embodiments of this application do not limit it.
[0101] Determining a target historical lip image from at least one historical lip image that matches the pronunciation of the current lip image to be generated includes: determining whether the phoneme corresponding to the historical lip image is the same as the phoneme corresponding to the current lip image to be generated; if the phoneme corresponding to the historical lip image is the same as the phoneme corresponding to the current lip image to be generated, determining the historical lip image as a target historical lip image that matches the pronunciation of the current lip image to be generated; if the phoneme corresponding to the historical lip image is different from the phoneme corresponding to the current lip image to be generated, determining that the historical lip image is not a target historical lip image that matches the pronunciation of the current lip image to be generated.
[0102] In this embodiment of the application, when each frame of lip image is generated, the image features of the target historical lip image with pronunciation matching are extracted from at least one recently generated historical lip image as a reference, and are fused into the denoising process when the current frame of lip image is generated through the Memory Attention mechanism to improve the consistency between frames.
[0103] Furthermore, the lip image sequence generation method provided in this application embodiment further includes: smoothing the motion trajectory of lip key points predicted based on speech information during the lip image sequence generation process to generate the target lip key point motion trajectory.
[0104] In this embodiment, the smoothed motion trajectory of the key points of the lips is called the target motion trajectory of the key points of the lips.
[0105] In the process of generating speech-driven virtual avatar lip-sync animation, due to the discrete changes in speech features at phoneme boundaries, the lip key points predicted directly based on speech features are prone to abrupt changes between adjacent lip images, resulting in problems such as lip trembling, discontinuous opening and closing, or teeth flickering in the generated virtual avatar lip-sync animation.
[0106] A predictive keypoint interpolation and smoothing mechanism is introduced during the generation stage of the lip image sequence to continuously process the trajectories of lip and tooth-related keypoints that change over time. Specifically, for the spatial position changes of corresponding keypoints in adjacent lip images, curve interpolation or smoothing operations are performed on the time axis to construct continuous and smooth keypoint motion trajectories.
[0107] Interpolation smoothing can be achieved using Bézier curves, spline curves, or equivalent time series smoothing methods. Its application is limited to key points in the lip and teeth regions and does not affect the posture or expression changes in other areas of the face.
[0108] The keypoint trajectory after interpolation and smoothing serves as a temporal reference for subsequent lip region mask updates and lip image sequence generation model conditional inputs. This enables the lip image sequence generation model to obtain smoother spatial constraints when generating continuous lip images, thereby effectively suppressing lip shape jumps and improving the naturalness of lip movements and the temporal continuity of lip image sequences.
[0109] Furthermore, the method for generating a lip image sequence provided in this application embodiment further includes: using a lip region mask as a spatial attention guide, and combining optical flow consistency constraints to generate a lip image sequence that conforms to the target lip features based on speech information and the original face image.
[0110] In this embodiment of the application, the lip image sequence generation model combines optical flow consistency constraints to ensure that the lip movements in the generated lip image sequence conform to physical laws during the generation of the lip image sequence.
[0111] Optical flow consistency constraints are used to constrain the continuity of lip motion between adjacent lip images. The optical flow consistency constraints are as follows: in, Represents the generated lip image and The calculated actual pixel motion field; This represents the ideal pixel motion field predicted based on changes in speech features. By minimizing the difference between the two, lip tremors and abrupt changes can be effectively suppressed, improving the visual smoothness of virtual avatar lip-syncing animation.
[0112] This application proposes a method for generating lip image sequences. The method acquires speech information used to drive the generation of virtual avatar lip-shape animation, an original face image, and lip feature information indicating target lip features. The target lip features include at least one of target tooth features and target lip features. A lip region mask is determined from the original face image, including a tooth region mask and a lip region mask. The lip region mask is used as a spatial attention guide to generate a lip image sequence that conforms to the target lip features based on the speech information and the original face image. This application separates the tooth and lip regions of the original face image into two independent masks: a tooth region mask and a lip region mask. When generating the lip region image sequence, independent control of the tooth and lip regions in the lip image to be generated is achieved, effectively realizing personalized generation of virtual avatar lip-shape animation and enhancing the user experience.
[0113] Exemplary device Accordingly, this application also provides a lip image sequence generation device.
[0114] Please see Figure 5In one exemplary embodiment, a lip image sequence generation apparatus 500 is provided, the lip image sequence generation apparatus 500 comprising: The information acquisition unit 501 is used to acquire speech information, original face image and lip feature information indicating target lip features for driving the generation of virtual image lip-shaped animation, the target lip features including at least one of target tooth features and target lip features; The lip region mask determination unit 502 is used to determine the lip region mask of the original face image. The lip region mask includes the teeth region mask and the lip region mask. The lip image sequence generation unit 503 is used to generate a lip image sequence that conforms to the target lip features based on the speech information and the original face image, using the lip region mask as a spatial attention guide.
[0115] In one possible implementation, the lip region mask determination unit 502 includes: The initial lip region mask determination unit is used to perform key point detection on the original face image to determine the initial lip region mask of the original face image. The lip region mask generation unit is used to perform tooth detection on the lip region of the original face image based on the initial lip region mask, determine the tooth region and lip region in the lip region, and generate the lip region mask.
[0116] In one possible implementation, the initial lip region mask determination unit includes: The key point detection unit is used to perform key point detection on the original face image to obtain a set of face key points; The initial lip region mask generation unit is used to determine the lip region of the original face image and generate an initial lip region mask based on the perioral contour points in the set of facial key points.
[0117] In one possible implementation, the lip region mask generation unit includes: The oral cavity opening region determination unit is used to determine the oral cavity opening region in the lip region of the original face image based on the inner lip contour key points in the face key point set; The lip region mask generation subunit is used to perform tooth detection on the oral cavity opening region of the original face image, and determine the tooth region and lip region in the lip region based on the initial lip region mask to generate the lip region mask; the lip region includes the non-tooth region in the lip region.
[0118] In one possible implementation, the lip image sequence generation unit 503 includes: The joint condition construction unit is used to construct joint conditions based on speech information, lip feature information, and the time axis of the lip image sequence. The lip image sequence generation subunit is used to take the joint conditions, the image features of the original face image, and the lip region mask as the guiding conditions for the lip image sequence generation model, so that the lip image sequence generation model can generate a lip image sequence that conforms to the target lip features.
[0119] In one possible implementation, the joint conditional construction unit includes: The speech feature generation unit is used to extract speech features from speech frames in speech information and map the speech features of speech frames to the time axis of the lip image sequence to generate speech features of speech information. The semantic feature generation unit is used to encode lip feature information and generate semantic features of the lip feature information; The joint condition construction subunit is used to construct joint conditions based on the speech features of speech information, the semantic features of lip feature information, and the time axis of the lip image sequence.
[0120] In one possible implementation, the lip image sequence generation subunit is specifically used to: use joint conditions, image features of the original face image, and lip region mask as guiding conditions for the lip image sequence generation model, so that during the lip image sequence generation process, the lip image sequence generation model uses the tooth region mask as the first spatial guide to make the tooth features of the lip image conform to the target tooth features, and uses the lip region mask as the second spatial guide to make the lip features of the lip image conform to the target lip features, thereby obtaining a lip image sequence.
[0121] In one possible implementation, the lip image sequence generation apparatus 500 further includes a reference information determination unit; the reference information determination unit is used to determine, during the lip image sequence generation process, a target historical lip image that matches the pronunciation corresponding to the current lip image to be generated from at least one recently generated historical lip image, and to use the image features of the target historical lip image as a reference to generate the current lip image.
[0122] The lip image sequence generation apparatus 500 provided in this embodiment belongs to the same concept as the lip image sequence generation method provided in the above embodiments of this application. It can execute the lip image sequence generation method provided in any of the above embodiments of this application and has the corresponding functional modules and beneficial effects of executing the lip image sequence generation method. Technical details not described in detail in this embodiment can be found in the corresponding method embodiments of this application, and will not be repeated here.
[0123] The functions implemented by each unit in the above device can be implemented by the same or different processors, and this application embodiment does not limit this.
[0124] It should be understood that the units in the above device can be implemented by a processor calling software. For example, the device includes a processor connected to a memory containing instructions. The processor calls the instructions stored in the memory to implement any of the above methods or to implement the functions of each unit in the device. The processor can be a general-purpose processor, such as a CPU or microprocessor, and the memory can be internal or external to the device. Alternatively, the units in the device can be implemented as hardware circuits. By designing the hardware circuits, some or all of the unit functions can be implemented. The hardware circuits can be understood as one or more processors. For example, in one implementation, the hardware circuit is an ASIC, and the functions of some or all of the above units are implemented by designing the logical relationships between the components within the circuit. In another implementation, the hardware circuit can be implemented using a PLD, such as an FPGA, which can include a large number of logic gates. The connection relationships between the logic gates are configured through configuration files to implement the functions of some or all of the above units. All units in the above device can be implemented entirely by a processor calling software, entirely by hardware circuits, or partially by a processor calling software with the remaining parts implemented by hardware circuits.
[0125] In this application embodiment, a processor is a circuit with signal processing capabilities. In one implementation, the processor can be a circuit with instruction reading and execution capabilities, such as a CPU, microprocessor, GPU, or DSP. In another implementation, the processor can implement certain functions through the logical relationships of hardware circuits. These logical relationships are fixed or reconfigurable. For example, the processor may be a hardware circuit implemented as an ASIC or PLD, such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration document and configuring the hardware circuit can be understood as the processor loading instructions to implement the functions of some or all of the above units. Furthermore, it can also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as an NPU, TPU, or DPU.
[0126] As can be seen, each unit in the above device can be one or more processors (or processing circuits) configured to implement the above methods, such as: CPU, GPU, NPU, TPU, DPU, microprocessor, DSP, ASIC, FPGA, or a combination of at least two of these processor forms.
[0127] Furthermore, the units in the above devices can be integrated in whole or in part, or they can be implemented independently. In one implementation, these units are integrated together and implemented in the form of a System-on-Chip (SoC). The SoC may include at least one processor for implementing any of the above methods or implementing the functions of the units in the device. The at least one processor may be of different types, such as CPU and FPGA, CPU and artificial intelligence processor, CPU and GPU, etc.
[0128] Exemplary electronic devices Another embodiment of this application also proposes an electronic device. See [link to relevant documentation]. Figure 6 As shown, the electronic device may include: a memory 600 and a processor 610; wherein the memory 600 is connected to the processor 610 and is used to store a program; the processor 610 is used to implement the lip image sequence generation method disclosed in any of the above embodiments by running the program stored in the memory 600.
[0129] Specifically, the aforementioned electronic device may also include: a bus, a communication interface 620, an input device 630, and an output device 640.
[0130] The processor 610, memory 600, communication interface 620, input device 630, and output device 640 are interconnected via a bus. Among them: A bus can include a pathway for transmitting information between various components of a computer system.
[0131] The processor 610 can be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or an application-specific integrated circuit (ASIC), or one or more integrated circuits used to control the execution of the program of the present application. It can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0132] The processor 610 may include a main processor, as well as a baseband chip, modem, etc.
[0133] The memory 600 stores a program for executing the technical solution of this application, and may also store an operating system and other key business functions. Specifically, the program may include program code, which includes computer operation instructions. More specifically, the memory 600 may include read-only memory (ROM), other types of static storage devices capable of storing static information and instructions, random access memory (RAM), other types of dynamic storage devices capable of storing information and instructions, disk storage, flash memory, etc.
[0134] Input device 630 may include a device for receiving user input data and information, such as a keyboard, mouse, camera, scanner, light pen, voice input device, touch screen, pedometer, or gravity sensor.
[0135] Output device 640 may include devices that allow information to be output to a user, such as a display screen, printer, speaker, etc.
[0136] The communication interface 620 may include a device that uses any transceiver to communicate with other devices or communication networks, such as Ethernet, Radio Access Network (RAN), Wireless Local Area Network (WLAN), etc.
[0137] The processor 610 executes the program stored in the memory 600 and calls other devices, and can be used to implement each step of any of the lip image sequence generation methods provided in the above embodiments of this application.
[0138] This application also proposes a chip, which includes a processor and a data interface. The processor reads and runs a program stored in the memory through the data interface to execute any of the lip image sequence generation methods provided in the above embodiments. For the specific processing procedure and its beneficial effects, please refer to the above embodiments of the lip image sequence generation method.
[0139] Exemplary computer program products and storage media In addition to the methods and devices described above, embodiments of this application may also be computer program products, which include computer program instructions that, when executed by a processor, cause the processor to perform the steps in the lip image sequence generation method according to various embodiments of this application as described in any of the above embodiments of this specification.
[0140] The computer program product can be written in any combination of one or more programming languages to perform the operations of the embodiments of this application. The programming languages include object-oriented programming languages such as Java and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0141] Furthermore, embodiments of this application may also be storage media storing a computer program, which is executed by a processor in the steps of the lip image sequence generation method according to various embodiments of this application described in any of the above embodiments of this specification.
[0142] For the foregoing method embodiments, in order to simplify the description, they are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, because according to this application, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0143] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For apparatus embodiments, since they are basically similar to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0144] The steps in the methods of the various embodiments of this application can be adjusted, merged, or deleted in order according to actual needs, and the technical features described in each embodiment can be replaced or combined.
[0145] The modules and sub-modules in the apparatus and terminal in the various embodiments of this application can be merged, divided, and deleted according to actual needs.
[0146] It should be understood that the disclosed terminals, devices, and methods can be implemented in other ways, given the several embodiments provided in this application. For example, the terminal embodiments described above are merely illustrative. For instance, the division of modules or sub-modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple sub-modules or modules may be combined or integrated into another module, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or modules, and may be electrical, mechanical, or other forms.
[0147] The modules or submodules described as separate components may or may not be physically separate. The components that constitute a module or submodule may or may not be physical modules or submodules; that is, they may be located in one place or distributed across multiple network modules or submodules. Some or all of the modules or submodules can be selected to achieve the purpose of this embodiment's solution, depending on actual needs.
[0148] Furthermore, the functional modules or sub-modules in the various embodiments of this application can be integrated into one processing module, or each module or sub-module can exist physically separately, or two or more modules or sub-modules can be integrated into one module. The integrated modules or sub-modules described above can be implemented in hardware or in the form of software functional modules or sub-modules.
[0149] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0150] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software unit executed by a processor, or a combination of both. The software unit can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0151] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0152] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for generating a sequence of lip images, characterized in that, include: Acquire speech information, original face image, and lip feature information indicating target lip features for driving the generation of virtual avatar lip-splitting animation, wherein the target lip features include at least one of target tooth features and target lip features; Determine the lip region mask of the original face image, wherein the lip region mask includes a teeth region mask and a lip region mask; Using the lip region mask as spatial attention guide, a sequence of lip images conforming to the target lip features is generated based on the speech information and the original face image.
2. The method according to claim 1, characterized in that, Determining the lip region mask of the original face image includes: Key point detection is performed on the original face image to determine the initial lip region mask of the original face image; Based on the initial lip region mask, tooth detection is performed on the lip region of the original face image to determine the tooth region and lip region within the lip region, and a lip region mask is generated.
3. The method according to claim 2, characterized in that, The step of performing key point detection on the original face image to determine the initial lip region mask of the original face image includes: A set of facial key points is obtained by performing key point detection on the original face image. Based on the perioral contour points in the set of facial key points, the lip region of the original face image is determined and an initial lip region mask is generated.
4. The method according to claim 3, characterized in that, The step of performing tooth detection on the lip region of the original face image based on the initial lip region mask, determining the tooth region and lip region within the lip region, and generating a lip region mask includes: Based on the inner lip contour key points in the set of facial key points, determine the oral cavity opening region in the lip region of the original facial image; The teeth are detected in the oral cavity opening area of the original face image, and the tooth area and lip area in the lip area are determined according to the initial lip area mask to generate a lip area mask; the lip area includes the non-tooth area in the lip area.
5. The method according to claim 1, characterized in that, The step of using the lip region mask as spatial attention guidance to generate a sequence of lip images conforming to the target lip features based on the speech information and the original face image includes: A joint condition is constructed based on the speech information, the lip feature information, and the time axis of the lip image sequence. The joint conditions, the image features of the original face image, and the lip region mask are used as guiding conditions for the lip image sequence generation model, so that the lip image sequence generation model generates a lip image sequence that conforms to the target lip features.
6. The method according to claim 5, characterized in that, The step of constructing joint conditions based on the speech information, the lip feature information, and the timeline of the lip image sequence includes: Extract the speech features of the speech frames in the speech information, and map the speech features of the speech frames to the time axis of the lip image sequence to generate the speech features of the speech information; The lip feature information is encoded to generate semantic features of the lip feature information; Based on the speech features of the speech information, the semantic features of the lip features, and the time axis of the lip image sequence, a joint condition is constructed.
7. The method according to claim 5, characterized in that, The step of using the joint conditions, the image features of the original face image, and the lip region mask as guiding conditions for the lip image sequence generation model to generate a lip image sequence that conforms to the target lip features includes: The joint conditions, the image features of the original face image, and the lip region mask are used as guiding conditions for the lip image sequence generation model. During the generation of the lip image sequence, the model uses the tooth region mask as a first spatial guide to ensure that the tooth features of the lip image conform to the target tooth features, and uses the lip region mask as a second spatial guide to ensure that the lip features of the lip image conform to the target lip features, thus obtaining the lip image sequence.
8. The method according to claim 1, characterized in that, Also includes: During the generation of the lip image sequence, a target historical lip image that matches the pronunciation of the current lip image to be generated is determined from at least one recently generated historical lip image, and the image features of the target historical lip image are used as a reference to generate the current lip image.
9. A device for generating a sequence of lip images, characterized in that, include: The information acquisition unit is used to acquire voice information, original face image and lip feature information indicating target lip features for driving the generation of virtual image lip-shaped animation, the target lip features including at least one of target tooth features and target lip features; The lip region mask determination unit is used to determine the lip region mask of the original face image, wherein the lip region mask includes a teeth region mask and a lip region mask; The lip image sequence generation unit is used to generate a lip image sequence that conforms to the target lip features based on the speech information and the original face image, using the lip region mask as spatial attention guidance.
10. An electronic device, characterized in that, Including memory and processor; The memory is connected to the processor and is used to store programs; The processor is configured to implement the lip image sequence generation method as described in any one of claims 1 to 8 by running a program in the memory.