Method and apparatus for generating lip-shape images of virtual objects
By combining deformers and amplitude curves, virtual character lip-sync images are automatically generated, solving the problem of low efficiency in virtual character lip-sync animation generation. This achieves efficient and synchronized lip-sync image generation, suitable for scenarios such as games, movies, and live streaming.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING PERFECT WORLD SOFTWARE TECH DEV CO LTD
- Filing Date
- 2022-06-30
- Publication Date
- 2026-05-05
AI Technical Summary
In existing technologies, the automation level of virtual character lip-sync animation generation is low, resulting in poor animation production efficiency and difficulty in handling large-scale virtual character lip-sync animation generation scenarios. In particular, in game development, lip-sync animations for different characters cannot be reused, affecting development efficiency.
By acquiring dubbing materials and utilizing the deformers and amplitude curves in the deformer template, the lip-shape images of virtual objects are automatically generated. The deformer contains the mapping relationship between the pronunciation lip shape and the skeletal model. The lip shapes of initials and finals are constructed based on the rules of Chinese Pinyin. The amplitude curve indicates the audio amplitude, thereby achieving synchronous adjustment of the facial lip-shape images.
It improves the efficiency and synchronization of lip-sync image generation, enhances the accuracy of lip-sync images and dubbing materials, improves audiovisual effects, adapts to the style requirements of different virtual objects, and meets the needs of mass production.
Smart Images

Figure CN117372577B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image technology, and in particular to a method and apparatus for generating lip-shape images of virtual objects. Background Technology
[0002] In scenarios such as gaming, film and television, and live streaming, it is necessary to adapt lip-sync animations to match the audio of virtual characters. This ensures that the lip movements in the lip-sync animation match the pronunciation in the character's audio, enhancing the realism of the virtual character. Virtual characters include game characters, characters in film and television works, and the virtual avatars of streamers in live streaming.
[0003] Most related technologies do not support Chinese pronunciation rules, resulting in poor lip-sync animation effects for virtual characters. Therefore, the current mainstream approach still relies on technicians manually creating lip-sync animations for virtual characters. In this method, technicians use facial capture technology to collect facial data from actors, and then combine this data with the virtual character's design to create lip-sync animations. This method has low automation, poor animation production efficiency, and is difficult to handle large-scale virtual character lip-sync animation generation scenarios. Therefore, how to automatically generate lip-sync animations for virtual characters has become a pressing technical problem to be solved. Summary of the Invention
[0004] This invention provides a method and apparatus for generating lip-sync images of virtual objects, which can realize the automated generation of lip-sync images, greatly improve the generation efficiency of lip-sync images, improve the synchronization and accuracy of lip-sync images and dubbing materials, and optimize the audiovisual effects of lip-sync images.
[0005] In a first aspect, embodiments of the present invention provide a method for generating a lip-shape image of a virtual object, the method comprising:
[0006] Acquire the dubbing material to be processed, which includes audio data and / or text data corresponding to virtual objects;
[0007] Obtain a deformer that matches the virtual object from a pre-set deformer template. The deformer includes the mapping relationship between the pronunciation mouth shape and the skeletal model. The pronunciation mouth shape includes the mouth shape of the initial consonant and / or the mouth shape of the final vowel, which are constructed based on the combination of Chinese Pinyin rules.
[0008] An amplitude curve corresponding to the pronunciation mouth shape is generated based on the dubbing material. The amplitude curve is used to indicate the audio amplitude corresponding to each phoneme in the dubbing material. Each phoneme in the dubbing material corresponds one-to-one with the mouth shape of the initial consonant and / or the mouth shape of the final vowel in the pronunciation mouth shape.
[0009] The voice-over material is mapped onto the skeletal model of the virtual object through a deformer, generating a facial lip-sync image synchronized with the voice-over material, and then the facial lip-sync image is adjusted to match the lip-sync image of the virtual object through an amplitude curve.
[0010] In one possible embodiment, obtaining the voice-over material to be processed includes:
[0011] Receive audio and / or text data input by the user; identify multiple virtual objects from the audio and / or text data, and extract data segments corresponding to each virtual object from the audio and / or text data as dubbing material.
[0012] In one possible embodiment, obtaining a deformer matching the virtual object from a pre-set deformer template includes:
[0013] The deformer panel displays at least one pre-set deformer template, which includes a deformer and a corresponding mapping pool. The mapping pool is used to store the mapping relationship between at least one vocal lip shape and at least one skeletal model. In response to the selection command of the deformer, the skeletal model corresponding to the virtual object is determined, and a deformer that matches the skeletal model corresponding to the virtual object is selected from at least one deformer template.
[0014] In one possible embodiment, the method further includes setting a corresponding bone model for a deformer in a deformer template, wherein the corresponding bone model is reused for multiple virtual objects.
[0015] In one possible embodiment, the voice-over material is mapped onto the skeletal model of the virtual object via the deformer to generate a facial lip-sync image synchronized with the voice-over material, and the facial lip-sync image is adjusted to match the lip-sync image of the virtual object using the amplitude curve, including:
[0016] The deformer identifies each phoneme in the dubbing material; the identified phonemes are mapped to the skeletal model of the virtual object to obtain the corresponding skeletal model parameters; a facial lip shape image is calculated based on the skeletal model parameters; the amplitude curve is displayed in the amplitude panel; in response to the editing command of the amplitude curve, the change range of the amplitude curve is adjusted to change the change range of the lip shape size in the lip shape image.
[0017] In one possible embodiment, generating a corresponding amplitude curve based on the dubbing material includes: selecting keyframes from each phoneme in the dubbing material, wherein the keyframes include audio data frames corresponding to the initials and / or finals in the dubbing material.
[0018] Display amplitude curves in the amplitude panel, including: displaying the amplitude curves corresponding to keyframes in the amplitude panel.
[0019] In one possible embodiment, the method further includes: adjusting the mapping parameters of the deformer in response to an editing instruction on the deformer template to modify the mapping relationship between the lip movements and the skeletal model.
[0020] In one possible embodiment, the method further includes: adjusting the animation preset parameters in response to an editing instruction for the animation preset parameters to modify the visual effect of the lip-sync image; wherein the animation preset parameters include at least one of the following parameters: lip-sync animation style, frame rate, sampling parameters, extra duration, fade-in / fade-out.
[0021] In one possible embodiment, the method further includes: performing semantic recognition on the dubbing material; determining whether the dubbing material meets preset conditions based on the recognition result; if the dubbing material meets the preset conditions, adding specific visual elements associated with the virtual object to the facial lip-sync image, the specific visual elements including facial expressions and / or movements bound to the skeletal model.
[0022] In one possible embodiment, the association between a virtual object and a specific visual element includes: the association between a virtual object and a specific visual element; and / or the association between a preset statement of a virtual object and a specific visual element; and / or the association between a preset plot in dubbing material and a specific visual element.
[0023] In a second aspect, embodiments of the present invention provide a lip-shape image generation apparatus for a virtual object, the lip-shape image generation apparatus comprising:
[0024] The acquisition module is used to acquire dubbing materials to be processed, including audio data and / or text data corresponding to virtual objects; and to acquire deformers that match virtual objects from pre-set deformer templates. The deformers include the mapping relationship between pronunciation mouth shapes and skeletal models, and the pronunciation mouth shapes include initial consonant mouth shapes and / or final vowel mouth shapes constructed based on the combination of Chinese Pinyin rules.
[0025] The generation module is used to generate an amplitude curve corresponding to the pronunciation mouth shape based on the dubbing material. The amplitude curve is used to indicate the audio amplitude corresponding to each phoneme in the dubbing material. Each phoneme in the dubbing material corresponds one-to-one with the mouth shape of the initial consonant and / or the mouth shape of the vowel in the pronunciation mouth shape. The dubbing material is mapped to the skeletal model of the virtual object through a deformer to generate a facial mouth shape image synchronized with the dubbing material. The facial mouth shape image is adjusted to the mouth shape image of the virtual object through the amplitude curve.
[0026] This invention also provides a system including a processor and a memory, wherein the memory stores at least one instruction, at least one program, code set, or instruction set, and the at least one instruction, at least one program, code set, or instruction set is loaded and executed by the processor to implement the lip-shape image generation method for virtual objects described above.
[0027] This invention provides a computer-readable medium storing at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, at least one program, code set, or instruction set is loaded and executed by a processor to implement the lip-shape image generation method for the virtual object described above.
[0028] In this embodiment of the invention, the dubbing material to be processed is first acquired, which includes audio data and / or text data corresponding to the virtual object. Then, an amplitude curve corresponding to the lip-sync shape is generated based on the dubbing material. This amplitude curve indicates the audio amplitude corresponding to each phoneme in the dubbing material, and each phoneme in the dubbing material corresponds one-to-one with the initial consonant and / or final vowel lip-sync shapes in the lip-sync shape. A deformer matching the virtual object is obtained from a pre-set deformer template. Since the deformer includes a mapping relationship between the lip-sync shape and the skeletal model, and here, the lip-sync shape includes initial consonant and / or final vowel lip-sync shapes constructed based on the rules of Chinese Pinyin, the dubbing material can be mapped onto the skeletal model of the virtual object through the deformer to generate a facial lip-sync image synchronized with the dubbing material. This facial lip-sync image is then adjusted to match the lip-sync image of the virtual object using the amplitude curve. This invention, through a deformer matching virtual objects and amplitude curves, creates lip-shape images that conform to both Chinese Pinyin rules and the style of virtual objects. This achieves an automated lip-shape image generation process based on dubbing materials, avoiding the poor animation production efficiency caused by manual lip-shape image creation in related technologies. This significantly improves the generation efficiency of lip-shape images and helps meet the mass production needs of lip-shape images in practical applications. Furthermore, compared to the manual creation methods in related technologies, this invention, through the application of deformers and amplitude curves, can also improve the synchronization and accuracy between the final generated lip-shape image and the dubbing material, making the lip-shape image more natural and fluid, and greatly enhancing the audiovisual effect of the lip-shape image. Attached Figure Description
[0029] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0030] Figure 1 This is a flowchart illustrating a method for generating a lip shape image of a virtual object according to an embodiment of the present invention.
[0031] Figure 2 This is a schematic diagram of a text panel provided in an embodiment of the present invention;
[0032] Figure 3 This is a schematic diagram of a deformer panel provided in an embodiment of the present invention;
[0033] Figure 4 This is a schematic diagram of an amplitude panel provided in an embodiment of the present invention;
[0034] Figure 5 A schematic diagram of another amplitude panel provided in an embodiment of the present invention;
[0035] Figure 6 This is a schematic diagram of an advanced settings panel provided in an embodiment of the present invention;
[0036] Figure 7 This is a schematic diagram of an export interface provided in an embodiment of the present invention;
[0037] Figure 8 This is a schematic diagram of an exported file provided in an embodiment of the present invention;
[0038] Figure 9 This is a schematic diagram of an export confirmation interface provided in an embodiment of the present invention;
[0039] Figure 10 This is a schematic diagram of a debug panel provided in one embodiment of the present invention;
[0040] Figure 11 This is a schematic diagram of the structure of a virtual object lip-shape image generation device provided in an embodiment of the present invention;
[0041] Figure 12 To and Figure 11 The illustrated embodiment provides a schematic diagram of the electronic device corresponding to the virtual object lip-shape image generation device. Detailed Implementation
[0042] The invention will now be discussed with reference to several exemplary embodiments. It should be understood that these embodiments are described merely to enable those skilled in the art to better understand and thus implement the invention, and are not intended to imply any limitation on the scope of the invention.
[0043] As used herein, the term "comprising" and its variations are to be interpreted as open-ended terms meaning "including but not limited to". The term "based on" is to be interpreted as "at least partially based on". The terms "one embodiment" and "an embodiment" are to be interpreted as "at least one embodiment". The term "another embodiment" is to be interpreted as "at least one other embodiment".
[0044] Furthermore, the timing of the steps in the following method embodiments is merely an example and not a strict limitation.
[0045] Currently, in scenarios such as games, movies, and live streaming, it is necessary to adapt lip-sync animations to match the audio of virtual characters, thereby enhancing the realism of the virtual characters. Examples of virtual characters include game characters, characters in movies and TV shows, and the virtual avatars of live streamers.
[0046] The applicant discovered that most related technologies do not support Chinese pronunciation rules, resulting in poor lip-sync animation effects for virtual characters. Therefore, the current solution still primarily relies on technicians manually creating the lip-sync animations for virtual characters. In this lip-sync animation production solution, technicians use facial capture technology to collect facial data from actors, and then combine this facial data with the virtual character's design to create the lip-sync animation.
[0047] The applicant found that this method of generating lip-sync animation has a low degree of automation, poor animation production efficiency, and is difficult to handle large-scale scenarios involving the generation of lip-sync animation for virtual characters. For example, in game development projects, different game characters have different styles of facial expressions, so the lip-sync animations for different game characters cannot be reused. Relevant technical personnel need to create lip-sync animations separately for different game characters, resulting in poor animation production efficiency and significantly reducing game development efficiency.
[0048] In summary, how to automatically generate lip-sync animations for virtual characters has become a pressing technical problem that needs to be solved.
[0049] The lip-shape image generation scheme provided in this embodiment of the invention can be executed by an electronic device, such as a smartphone, tablet computer, PC, or laptop computer. In an optional embodiment, the electronic device may have an application program installed for executing the lip-shape image generation scheme. Alternatively, in another optional embodiment, the lip-shape image generation scheme can be executed jointly by a server device and a terminal device.
[0050] For example, suppose the first service program loads a virtual scene. The aforementioned electronic device can be implemented as a second service program for displaying virtual characters (i.e., virtual objects) in the virtual scene. This second service program can connect to the first service program, and based on the second service program, it can create and adjust the lip-sync animation of the virtual characters loaded by the first service program, and display the virtual objects loaded by the first service program in real time. Here, real-time display can be understood as displaying each frame of the lip-sync image of the virtual character in real time.
[0051] In practical applications, the first service program may be a virtual scene editor or a game editor, and the second service program may be a plugin attached to the first service program. Of course, in addition to plugins, the second service program may also be an application program independent of the first service program, and this invention is not limited thereto.
[0052] The solutions provided in this invention are applicable to various lip-sync image creation scenarios, such as the generation, modification, and optimization of virtual object lip-sync images. Examples include lip-sync image creation scenarios in fields such as games, films, and live streaming.
[0053] To address the aforementioned technical problems, a solution is provided in some embodiments of the present invention. The technical solutions provided by the embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0054] The execution process of the method for generating the lip shape image of the virtual object will be described below with reference to the following embodiments. Figure 1 This is a flowchart of a method for generating lip-shape images of virtual objects, provided as an embodiment of the present invention.
[0055] like Figure 1 As shown, the method for generating the lip shape image of this virtual object includes the following steps:
[0056] 101. Obtain the dubbing materials to be processed;
[0057] 102. Obtain a deformer that matches the virtual object from the pre-defined Morph template;
[0058] 103. Generate amplitude curves corresponding to lip movements based on dubbing materials;
[0059] 104. Map the dubbing material onto the skeletal model of the virtual object using a deformer to generate a facial lip-sync image synchronized with the dubbing material, and adjust the facial lip-sync image to match the lip-sync image of the virtual object using an amplitude curve.
[0060] The lip-sync image generation method in this embodiment of the invention is applied to an application program, which can be set up on a terminal device. This application program loads virtual objects from a virtual scene. These virtual objects can be implemented as game characters, characters from movies and television shows, or virtual avatars of live streamers.
[0061] In step 101, the voice-over material to be processed is obtained. In this embodiment of the invention, the voice-over material includes, but is not limited to, audio data and / or text data corresponding to virtual objects. Taking a virtual character in a game as an example, the voice-over material can be a voice-over file corresponding to the virtual character, or it can be the dialogue text corresponding to the virtual character.
[0062] In one optional embodiment, audio data input by the user is received, and corresponding text data is extracted from the input audio data via cloud computing as dubbing material. For example, the audio data input by the user is displayed in an audio panel, including but not limited to information such as the name, path, duration, and volume of the audio file, allowing the user to view and edit this information. In practical applications, the audio data can be an audio file of a single virtual object, or audio files of multiple virtual objects, such as audio files of various virtual characters in a game or film.
[0063] Specifically, in the above embodiments, a single character's voice-over file can be imported, and the corresponding voice-over text can be obtained through speech recognition processing. Alternatively, multiple character voice-over files can be imported, such as dialogue audio of multiple characters in a certain level, or guidance voice lines triggered by different characters in the same level; then, cloud computing can be used to extract the corresponding voice-over text from the aforementioned voice-over files according to the character. This material acquisition method can extract corresponding text data from audio data, thereby reducing the difficulty of subsequent material processing and further improving the generation efficiency of virtual character lip-sync animation. Optionally, it is also possible to... Figure 2 In the text panel shown, adjust the correspondence between the automatically acquired text data and the timeline to align them and further optimize the synchronization effect of the lip-sync animation.
[0064] In another alternative embodiment, text data edited and input by the user is received as voice-over material. For example, in Figure 2 The text panel shown receives text content edited and input by the user, and aligns the text content with the timeline. Optionally, it receives the start time, end time, or corresponding lip-sync animation duration of the user-input text content, and adjusts the correspondence between the text content and the timeline based on the aforementioned time information.
[0065] In this embodiment of the invention, whether it is audio data or text data, after receiving the user-input audio data and / or text data, multiple virtual objects can be identified from the audio data and / or text data, and data segments corresponding to each of the multiple virtual objects can be extracted as dubbing material. For example, multiple characters contained in the dubbing text can be identified through cloud computing, and their corresponding dubbing texts can be extracted from the dubbing file according to different characters. Further optionally, different characters correspond to different types of skeletal model parameters, or different skeletal models can be bound to different characters, so that the lip-sync animation of the characters has different action styles of different types of characters through skeletal model parameters or skeletal models.
[0066] In step 102, a deformer matching the virtual object is obtained from a pre-set deformer template.
[0067] In this embodiment of the invention, the deformer includes a mapping relationship between lip movements and a skeletal model. The matching between the deformer and the virtual object can be a visual style match, such as matching the lip animation style of the virtual object with the mapping relationship set in the deformer. This allows the deformer to obtain a lip image that is consistent with the virtual object in terms of lip animation style. In other words, by setting parameters related to the mapping relationship between lip movements and the skeletal model, the deformer can be made to have different lip animation styles required by different virtual objects, thus enabling the deformer to be reused for different virtual objects. In practical applications, assuming the virtual object is a game character, the skeletal model in the deformer is bound to the game character. The stylistic differences in the lip animation of the game character are mainly used to represent the differences in expression and movement between different characters or the same character in different states, thereby enhancing the visual realism of the character. Specifically, the lip animation style is, for example, the action style set for the character, which can be categorized by character personality (including but not limited to gentle, rough, and efficient), by character profession (including but not limited to assassin, mage, warrior, and craftsman), and by character level (including but not limited to beginner and expert). Based on the above classification, different characters can be configured with skeletal models or skeletal model parameters that match their movement styles, thereby improving the realism of the character's lip-syncing animation through skeletal models or skeletal model parameters.
[0068] Optionally, before step 102, corresponding skeletal models can be set for the deformers in the deformer template. Specifically, the corresponding skeletal model is related to the application scenario of the virtual object, and the skeletal model is associated with the virtual object itself. For example, the skeletal model in the deformer template can be associated with a game character. Specifically, different skeletal models can be configured for different game characters, or different skeletal model parameters can be configured for different game characters within the same skeletal model, so as to achieve differences in the lip-sync images of different game characters in terms of movement style through the skeletal model. The skeletal model parameters include, for example, smoothness, deformation amplitude, and deformation curve.
[0069] In practice, to accommodate large-scale development needs, multiple skeletal models can be batch-set for deformers to be used by multiple virtual objects, allowing the skeletal models corresponding to deformers to be reused for multiple virtual objects. For example, assuming the application scenario of virtual objects is a game development scenario, and the virtual objects are virtual characters in the game, then a common skeletal model can be set for multiple virtual characters in the game. That is, the deformers in the deformer template are bound to the base skeletal models bound to multiple virtual characters in the game, thus using the base skeletal models as the corresponding skeletal models for the deformers. Optionally, multiple deformers can use the same skeletal model, enabling multiple virtual characters in the game to reuse the deformer template. For example, the skeletal models of different virtual characters in a game can be bound to the deformers in the deformer template according to the character style, thereby further improving the efficiency of lip-syncing image production and reducing the efficiency of virtual object animation development. Of course, to ensure that the lip-syncing animation styles of different virtual characters are reflected in the lip-syncing images, optionally, in response to the editing instructions of the deformer template, the mapping parameters of the deformer can be adjusted to modify the mapping relationship between the pronunciation lip shape and the skeletal model. In other words, by editing the parameters of different deformers in the deformer template, such as the parameters related to the mapping relationship between the pronunciation mouth shape and the skeletal model mentioned above, the deformer can be adapted to the mouth shape animation style required by different virtual objects.
[0070] Specifically, in step 102, at least one pre-set deformer template can be displayed in the deformer panel. The deformer template includes a deformer and a corresponding mapping pool. The mapping pool is used to store the mapping relationship between at least one vocal lip shape and at least one skeletal model. The mapping relationship stored in the mapping pool here is similar to the mapping relationship between vocal lip shape and skeletal model described above, and will not be elaborated further here.
[0071] Furthermore, in step 102, in response to a deformer selection instruction, the skeletal model corresponding to the virtual object is determined, and a deformer matching the skeletal model corresponding to the virtual object is selected from at least one deformer template. In some embodiments, the deformer selection instruction may be user-triggered. For example, in Figure 3In the deformer panel shown, users can select the corresponding deformer from the drop-down menu displaying the deformer template, or import a deformer that matches the corresponding skeletal model of the virtual object. Of course, in practical applications, if there are many deformers, users can also use search or fuzzy matching to assist in selecting a deformer that matches the virtual object; this embodiment does not impose such limitations. Optionally, before step 102, the correspondence between the virtual object and the skeletal model is bound in the deformer template, or the mapping relationship between the character and the skeletal model parameters is bound. In this way, different virtual objects can be bound to corresponding skeletal models through deformers, thereby allowing the final generated lip-sync animation to reflect the personalized characteristics of the virtual object in terms of movement style through differences in skeletal parameters. For example, different skeletal models can be bound to female and male characters, thus reflecting the differences between female and male characters in the movement style of the lip-sync animation.
[0072] In other embodiments, the deformer selection command can be automatically triggered based on the voice-over material and / or the virtual object. Taking a game development scenario as an example, the process of automatically triggering the deformer selection command involves parsing the voice-over material obtained in step 101 to obtain the corresponding pronunciation style features, such as young or old, male or female, hoarse or clear, etc., thereby selecting a deformer that matches the pronunciation style features as the deformer that matches the virtual object. Alternatively, the attribute parameters of the virtual object to be generated can be parsed to determine the pronunciation style features of the virtual object, and similarly, a deformer that matches the pronunciation style features can be selected as the deformer that matches the virtual object. Of course, when analyzing pronunciation style features, the voice-over material and the virtual object to be generated can also be combined to improve the compatibility between the deformer and the virtual object and enhance the visual effect of the final generated lip-sync image.
[0073] In related technologies, taking game development projects as an example, the facial expressions of different game characters have different styles, so the lip-sync animations of different game characters cannot be reused. Related technicians need to create lip-sync animations separately for different game characters, which results in poor animation production efficiency and greatly reduces game development efficiency.
[0074] To address the aforementioned technical issues, in step 103, a corresponding amplitude curve is generated based on the dubbing material. In step 104, since the deformer contains a mapping relationship between the pronunciation lip shape and the skeletal model, the dubbing material to be processed can be mapped onto the skeletal model of the virtual object through the deformer, generating a facial lip shape image synchronized with the dubbing material, and adjusting the facial lip shape image to the lip shape image of the virtual object through the amplitude curve.
[0075] The amplitude curve indicates the audio amplitude corresponding to each phoneme in the dubbing material. Each phoneme in the dubbing material corresponds one-to-one with the mouth shapes of the initial consonants and / or final vowels in the pronunciation lip-sync. Since each phoneme in the dubbing material is constructed based on the rules of Pinyin, the mouth shapes of the initial consonants and / or final vowels in the pronunciation lip-sync also need to be constructed based on the rules of Pinyin to synchronize the pronunciation lip-sync with the phonemes in the dubbing material, thereby enhancing the synchronization between the lip-sync image and the dubbing material.
[0076] Specifically, the amplitude involved in the embodiments of this invention is the amplitude corresponding to the audio signal. For example, in Figure 4 In the amplitude panel shown, the x-axis represents the time corresponding to the audio signal, and the y-axis represents the amplitude intensity corresponding to the audio signal.
[0077] In an optional embodiment of the above steps, in step 103, keyframes are selected from each phoneme in the dubbing material. The keyframes include audio data frames corresponding to the initials and / or finals in the dubbing material.
[0078] In step 104, a deformer is used to identify each phoneme in the dubbing material, and then the identified phonemes are mapped onto the skeletal model of the virtual object to obtain the corresponding skeletal model parameters; the facial lip shape image is calculated based on the skeletal model parameters. The skeletal model parameters include, for example, vertex parameters.
[0079] Furthermore, in step 104, the amplitude curve is displayed in the amplitude panel; in response to editing instructions on the amplitude curve, the amplitude curve's variation range is adjusted to change the variation range of the lip shape size in the lip-sync image. Furthermore, after selecting a keyframe, the amplitude curve corresponding to the keyframe is displayed in the amplitude panel, for example... Figure 5 The amplitude curve is shown in the amplitude panel. The above steps allow adjustment of the visual effect of the lip-sync image corresponding to the keyframes, further improving debugging efficiency. These steps automatically obtain the amplitude curve corresponding to the dubbing material, enabling adjustment and optimization of the lip-sync size variation in the lip-sync image. This provides a basis for subsequent adjustments to the visual effect of the lip-sync image, adapting to differences in facial changes in virtual objects. This allows the lip-sync image to be reused in different virtual objects, further improving the efficiency of lip-sync image production and reducing the animation development efficiency of virtual objects.
[0080] In the above or following embodiments, optionally, specific visual elements are associated with virtual objects, including but not limited to facial expressions and / or actions bound to the skeletal model, thereby establishing an association between the virtual object and the specific visual elements. Specifically, the association between the virtual object and the specific visual elements includes, but is not limited to, one or more of the following associations: the association between the virtual object and the specific visual element, the association between the virtual object's preset statements and the specific visual element, and the association between the preset plot in the voice-over material and the specific visual element. The specific visual elements (such as facial expressions and / or actions) can be implemented by setting the skeletal model parameters in the skeletal model. Taking a game development project as an example, for different game characters in the game project, an association between these game characters and the facial expressions bound to their respective skeletal models can be established to obtain a list of facial expressions associated with these game characters.
[0081] It can be understood that the facial expressions and / or actions bound to the skeletal model mentioned above can be specifically set for different virtual objects. Of course, besides setting exclusive facial expressions and / or actions for virtual objects, if game characters meet certain conditions—for example, multiple game characters belonging to the same series or the same storyline—then the bound facial expressions and / or actions can be reused among the skeletal models of these game characters. This facilitates the transfer of facial expressions and / or actions between multiple game characters, further improving the efficiency of lip-sync image production. In practical applications, after the same facial expression and / or action is bound to the skeletal models of different virtual objects (such as game characters), it can be associated with different lines in the voice-over materials of different virtual objects. For example, assuming that the facial expression of raising eyebrows is associated with preset lines of multiple virtual objects, then this facial expression can be associated with different lines of different virtual objects. For instance, the facial expression of raising eyebrows can be associated with "Really?" in the voice-over material of virtual object A, and "Not necessarily?" in the voice-over material of virtual object B. Of course, the same facial expression and / or action can also be associated with the same statement or the same plot in multiple virtual objects. For example, a raised eyebrow facial expression can be associated with multiple virtual objects saying "Really?". For example, a raised eyebrow facial expression can also be associated with the plot of multiple virtual objects encountering monster 1 in level 1, that is, if any of the above virtual objects is detected to encounter monster 1 in level 1, the display of this facial expression is triggered.
[0082] In practical applications, specific facial expressions can be, for example, signature expressions designed for game characters, or expressions obtained by adjusting the attributes of game characters. They can also be personalized settings created by players for the game character, such as facial expressions obtained through character customization. Specifically, facial expressions obtained by adjusting the attributes of game characters include, but are not limited to, raising eyebrows, smiling, blinking, pouting, and so on. Similarly, specific actions can be, for example, signature actions designed for game characters, or actions obtained by adjusting the attributes of game characters. Of course, specific actions can also be personalized settings created by players for the game character, such as player-specific actions obtained through interaction with the player or by analyzing player preference data.
[0083] Optionally, semantic recognition is performed on the dubbing material, and the recognition results are used to determine whether the dubbing material meets preset conditions. If the dubbing material meets the preset conditions, specific visual elements associated with the virtual object are added to the facial lip-sync image. These specific visual elements include facial expressions and / or actions. Specifically, this can be based on the association between the virtual object and facial expressions and / or actions, adding facial expressions and / or actions associated with the virtual object to the facial lip-sync image synchronized with the dubbing material. In practical applications, preset conditions include, but are not limited to: the dubbing material contains preset lines, the dubbing material belongs to a preset virtual object, and the dubbing material belongs to a preset game development project or game character series. Through the above steps, personalized settings for the facial lip-sync image of the virtual object can be achieved, adding more visual elements associated with the virtual object's own settings or attribute parameters to the facial lip-sync image, thereby further improving the visual effect and production efficiency of the virtual object's facial lip-sync image.
[0084] For example, suppose the preset condition is that the voice-over material belongs to a preset virtual object and contains a preset statement. Suppose that the preset statement "Why?" of virtual object a is associated with raising an eyebrow (facial expression). Based on this assumption, firstly, it is detected whether the voice-over material belongs to the preset virtual object a and whether it contains "Why?" (i.e., the preset statement). If the voice-over material is detected to belong to the preset virtual object a and contains "Why?", then based on the association between the preset statement of virtual object a and raising an eyebrow (i.e., the facial expression), an eyebrow-raising expression associated with virtual object a is added to the facial lip-sync image synchronized with the preset statement "Why?" in the voice-over material.
[0085] Alternatively, the above steps could also be as follows: Assume that virtual object b is associated with a smile (i.e., a facial expression). Based on this, detect whether the voice-over material belongs to the preset virtual object b. If the voice-over material is detected to belong to the preset virtual object b, then based on the association between virtual object b and a smile, add a smile expression associated with virtual object b to the facial lip-sync image synchronized with the voice-over material. This smile expression can be added at any position in the facial lip-sync image of virtual object b, for example, at the end or beginning of each line of dialogue.
[0086] In the above or below embodiments, optionally, after step 104, a lip-sync image of the virtual object can be displayed so that the user can adjust parameters based on the visual effect of the lip-sync image. For example, if the lip-sync changes too quickly in the lip-sync image, non-keywords in the corresponding dubbing material can be deleted to reduce the number of lip-sync image frames generated in the final output. Alternatively, a function to pre-detect and automatically delete non-keywords in the dubbing material can be triggered, which can also reduce the number of lip-sync image frames generated in the final output. Non-keywords include, for example, modal particles. Optionally, in the context of displaying lip-sync images, the camera can also be switched and camera parameters adjusted. Specifically, a drop-down menu for camera switching can be selected in the display interface, and the camera can be switched during lip-sync animation playback through this drop-down menu. This drop-down menu can automatically obtain the cameras already set in the current scene, making it convenient for the user to complete the camera switching operation in the current display interface, thereby avoiding the operational complexity caused by the camera switching leaving the current interface. Of course, the user can also turn off this camera switching function and manually switch and adjust the camera in the virtual scene.
[0087] Alternatively, in other embodiments, animation preset parameters can be adjusted in the advanced settings panel to optimize the visual effect of the lip-sync image. Optionally, in response to editing instructions for the animation preset parameters, the animation preset parameters are adjusted to modify the visual effect of the lip-sync image. The animation preset parameters include at least one of the following parameters: lip-sync animation style, frame rate, sampling parameters, extra duration, fade-in / fade-out, pause interval, smoothness, word ending closure, simplification curve, and sound amplitude weight. For example, in... Figure 6 In the advanced settings panel shown, when setting the above animation preset parameters, you can select the drop-down menu next to each parameter row to switch between different system preset parameter values. You can also click the "Restore" button in the panel to revert to the currently set parameter values. Alternatively, you can manually enter the parameter values in the input boxes. In practical applications, other parameters, such as smoothing feature parameters, can be used to modify the visual effect of the lip-sync image. These other parameters can be accessed through the "More Settings" button in the panel.
[0088] Optionally, after step 104, the lip-sync image of the virtual object can be exported for later application to specific scenarios. Specifically, using... Figure 7 Taking the export interface shown as an example, assuming the lip-sync image is a lip-sync animation, you can select an export scheme that matches the lip-sync animation from the various animation export schemes built into the application. In this embodiment, the export mode that outputs to an AnimSequence format file supported by Unreal Engine (UE) is usually used as the default animation export scheme. Optionally, in the animation export mode list, you can also view the animation export schemes that are currently supported or disabled by the device, and in the mode description section, you can view a detailed explanation of the animation export scheme.
[0089] like Figure 8 As shown in the exported file, the AnimSequence format file includes a lip-sync image and its corresponding amplitude curve. Optionally, when exporting the lip-sync image, the specific settings of the aforementioned animation preset parameters can be configured separately. For example... Figure 9 The export confirmation interface shown also allows for individual confirmation of information such as animation preset parameters, the skeletal model bound to the virtual object, skeletal model parameters, export file name, and path.
[0090] Optionally, embodiments of the present invention also provide a Debug panel, which is mainly used to troubleshoot abnormalities during lip-sync image playback or to adjust the playback effect of the lip-sync image. This panel allows for real-time display of the number of deformers used in the current lip-sync movement and their corresponding amplitude curves during playback, thus providing a clear view of any issues with the amplitude curves. For example, in... Figure 10 In the Debug panel shown, users can view the amplitude curve mapped by the deformer to quickly complete the adjustment of the amplitude curve.
[0091] Figure 1 The illustrated method for generating lip-sync images of virtual objects creates images that conform to both Chinese Pinyin rules and the style of the virtual object through a deformer matching the virtual object and an amplitude curve. This achieves automated generation of lip-sync images based on dubbing materials, avoiding the poor animation production efficiency caused by manual lip-sync image creation in related technologies. This significantly improves the generation efficiency of lip-sync images and helps meet the mass production needs of lip-sync images in practical applications. Furthermore, compared to the manual creation methods in related technologies, this embodiment of the invention, through the application of deformers and amplitude curves, can also improve the synchronization and accuracy between the final generated lip-sync image and the dubbing material, making the lip-sync image more natural and smooth, and greatly enhancing the audiovisual effect of the lip-sync image.
[0092] The lip-shape image generation apparatus of one or more embodiments of the present invention will be described in detail below. Those skilled in the art will understand that these lip-shape image generation apparatuses can all be configured using commercially available hardware components through the steps taught in this solution.
[0093] Figure 11 This is a schematic diagram of a lip-sync image generation device for a virtual object, provided in an embodiment of the present invention. This lip-sync image generation device is applied to a server, such as... Figure 11 As shown, the lip-sync image generation device includes: an acquisition module 11 and a generation module 12. Optionally, the lip-sync image generation device is applied to an application that loads virtual objects.
[0094] The acquisition module 11 is used to acquire dubbing materials to be processed, including audio data and / or text data corresponding to virtual objects; and to acquire deformers that match virtual objects from pre-set deformer templates. The deformers include the mapping relationship between pronunciation mouth shapes and skeletal models, and the pronunciation mouth shapes include initial consonant mouth shapes and / or final vowel mouth shapes constructed based on the combination of Chinese Pinyin rules.
[0095] The generation module 12 is used to generate an amplitude curve corresponding to the pronunciation mouth shape based on the dubbing material. The amplitude curve is used to indicate the audio amplitude corresponding to each phoneme in the dubbing material. Each phoneme in the dubbing material corresponds one-to-one with the mouth shape of the initial consonant and / or the mouth shape of the final vowel in the pronunciation mouth shape. The dubbing material is mapped to the skeletal model of the virtual object through a deformer to generate a facial mouth shape image synchronized with the dubbing material. The facial mouth shape image is then adjusted to the mouth shape image of the virtual object through the amplitude curve.
[0096] Optionally, when module 11 acquires the dubbing material to be processed, it is specifically used for:
[0097] It can receive audio data input by users and extract corresponding text data from the input audio data through cloud computing as dubbing material; or it can receive text data edited and input by users as dubbing material.
[0098] Optionally, when module 11 acquires the dubbing material to be processed, it is specifically used for:
[0099] Receive audio and / or text data input by the user; identify multiple virtual objects from the audio and / or text data, and extract data segments corresponding to each virtual object from the audio and / or text data as dubbing material.
[0100] Optionally, when the acquisition module 11 acquires a deformer matching the virtual object from a pre-set deformer template, it is specifically used for:
[0101] The deformer panel displays at least one pre-set deformer template, which includes a deformer and a corresponding mapping pool. The mapping pool is used to store the mapping relationship between at least one vocal lip shape and at least one skeletal model. In response to the selection command of the deformer, the skeletal model corresponding to the virtual object is determined, and a deformer that matches the skeletal model corresponding to the virtual object is selected from at least one deformer template.
[0102] Optionally, it also includes a settings module for setting the corresponding bone model for the deformer in the deformer template, wherein the corresponding bone model is reused for multiple virtual objects.
[0103] Optionally, when the generation module 12 maps the dubbing material to the skeletal model of the virtual object through the deformer, generates a facial lip-shape image synchronized with the dubbing material, and adjusts the facial lip-shape image to the lip-shape image of the virtual object through the amplitude curve, it is specifically used for: recognizing each phoneme in the dubbing material through the deformer; mapping each recognized phoneme to the skeletal model of the virtual object to obtain the corresponding skeletal model parameters; calculating the facial lip-shape image based on the skeletal model parameters; displaying the amplitude curve in the amplitude panel; and adjusting the amplitude curve's variation range in response to the amplitude curve editing command to change the variation range of the lip-shape size in the lip-shape image.
[0104] Optionally, when the generation module 12 generates the corresponding amplitude curve based on the dubbing material, it is specifically used to: select key frames from each phoneme in the dubbing material, the key frames including the audio data frames corresponding to the initials and / or finals in the dubbing material.
[0105] When generating module 12 displays the amplitude curve in the amplitude panel, it is specifically used to: display the amplitude curve corresponding to the key frame in the amplitude panel.
[0106] Optionally, it also includes a mapping parameter adjustment module for: adjusting the mapping parameters of the deformer in response to editing instructions on the deformer template, so as to modify the mapping relationship between the lip movements and the skeletal model.
[0107] Optionally, it also includes a preset parameter adjustment module, which is further used to: adjust the animation preset parameters in response to the editing command of the animation preset parameters to modify the visual effect of the lip-sync image; wherein the animation preset parameters include at least one of the following parameters: lip-sync animation style, frame rate, sampling parameters, extra duration, fade-in and fade-out.
[0108] Optionally, it also includes a semantic recognition module for: performing semantic recognition on the dubbing material; determining whether the dubbing material meets preset conditions based on the recognition results; if the dubbing material meets the preset conditions, adding specific visual elements associated with the virtual object to the facial lip-sync image, the specific visual elements including facial expressions and / or movements bound to the skeletal model.
[0109] Optionally, the association between virtual objects and specific visual elements includes: the association between virtual objects and specific visual elements; and / or the association between preset statements of virtual objects and specific visual elements; and / or the association between preset plots in dubbing materials and specific visual elements.
[0110] Figure 11 The lip-shape image generation device for the virtual object shown can execute the methods provided in the foregoing embodiments. For parts not described in detail in this embodiment, please refer to the relevant descriptions in the foregoing embodiments, which will not be repeated here.
[0111] In one possible design, the above Figure 11 The structure of the lip-shape image generation device shown can be implemented as an electronic device.
[0112] like Figure 12 As shown, the electronic device may include a processor 21 and a memory 22. The memory 22 stores executable code, which, when executed by the processor 21, enables the processor 21 to implement the lip-sync image generation method for virtual objects as provided in the foregoing embodiments. The electronic device may also include a communication interface 23 for communicating with other devices or communication networks.
[0113] The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separate. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0114] As needed, the systems, methods, and apparatuses of the various embodiments of the present invention can be implemented as pure software (e.g., software programs written in Java), or as pure hardware (e.g., dedicated ASIC chips or FPGA chips), or as a system combining software and hardware (e.g., a firmware system storing fixed code or a system with general-purpose memory and processor).
[0115] Another aspect of the present invention is a computer-readable medium having stored computer-readable instructions thereon, which, when executed, implement the lip-shape image generation method for virtual objects according to various embodiments of the present invention.
[0116] The various embodiments of the present invention have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The scope of the claimed subject matter is defined only by the appended claims.
Claims
1. A method for generating lip-shape images of virtual objects, characterized in that, include: Acquire the dubbing material to be processed, the dubbing material including audio data and / or text data corresponding to virtual objects; Obtain a deformer that matches the virtual object from a pre-set deformer template. The deformer includes a mapping relationship between the pronunciation mouth shape and the skeletal model. The pronunciation mouth shape includes the mouth shape of the initial consonant and / or the mouth shape of the final vowel, which are constructed based on the combination of Chinese Pinyin rules. An amplitude curve corresponding to the pronunciation mouth shape is generated based on the dubbing material. The amplitude curve is used to indicate the audio amplitude corresponding to each phoneme in the dubbing material. Each phoneme in the dubbing material corresponds one-to-one with the mouth shape of the initial consonant and / or the mouth shape of the final vowel in the pronunciation mouth shape. The voice-over material is mapped onto the skeletal model of the virtual object through the deformer, generating a facial lip-sync image synchronized with the voice-over material, and the facial lip-sync image is adjusted to match the lip-sync image of the virtual object through the amplitude curve.
2. The method according to claim 1, characterized in that, The process of acquiring the voice-over material to be processed includes: Receive audio and / or text data input by the user; Multiple virtual objects are identified from audio data and / or text data, and data segments corresponding to each of the multiple virtual objects are divided from the audio data and / or text data as the dubbing material.
3. The method according to claim 1, characterized in that, The step of obtaining a deformer matching the virtual object from a pre-set deformer template includes: The deformer panel displays at least one pre-set deformer template, which includes a deformer and a corresponding mapping pool. The mapping pool is used to store the mapping relationship between at least one vocal lip shape and at least one skeletal model. In response to a deformer selection command, the skeletal model corresponding to the virtual object is determined, and a deformer matching the skeletal model corresponding to the virtual object is selected from the at least one deformer template.
4. The method according to claim 1, characterized in that, Also includes: Set a corresponding bone model for the deformer in the deformer template, wherein the corresponding bone model is reused for multiple virtual objects.
5. The method according to claim 1, characterized in that, The voice-over material is mapped onto the skeletal model of the virtual object using the deformer, generating a facial lip-sync image synchronized with the voice-over material. The facial lip-sync image is then adjusted to match the lip-sync image of the virtual object using the amplitude curve, including: The deformer is used to identify each phoneme in the dubbing material; Each identified phoneme is mapped to the skeletal model of the virtual object to obtain the corresponding skeletal model parameters; The facial lip shape image is calculated based on the skeletal model parameters; The amplitude curve is displayed in the amplitude panel; In response to the editing command of the amplitude curve, the change range of the amplitude curve is adjusted to change the change range of the mouth shape size in the mouth shape image.
6. The method according to claim 5, characterized in that, The step of generating the corresponding amplitude curve based on the dubbing material includes: Keyframes are selected from each phoneme in the dubbing material, and the keyframes include audio data frames corresponding to the initials and / or finals in the dubbing material; Displaying the amplitude curve in the amplitude panel includes: The amplitude curve corresponding to the keyframe is displayed in the amplitude panel.
7. The method according to claim 1, characterized in that, Also includes: In response to the editing command of the deformer template, the mapping parameters of the deformer are adjusted to modify the mapping relationship between the pronunciation mouth shape and the skeletal model.
8. The method according to claim 1, characterized in that, Also includes: In response to an editing command for the animation preset parameters, the animation preset parameters are adjusted to modify the visual effect of the lip-sync image; The animation preset parameters include at least one of the following parameters: lip-sync animation style, frame rate, sampling parameters, extra duration, and fade-in / fade-out.
9. The method according to claim 1, characterized in that, Also includes: Perform semantic recognition on the dubbing material; Based on the recognition results, determine whether the dubbing material meets the preset conditions; If the dubbing material meets the preset conditions, then specific visual elements associated with the virtual object are added to the facial lip-sync image. The specific visual elements include facial expressions and / or movements bound to the skeletal model.
10. A lip-shape image generation device for a virtual object, characterized in that, The device includes: The acquisition module is used to acquire dubbing materials to be processed, the dubbing materials include audio data and / or text data corresponding to virtual objects; and to acquire a deformer that matches the virtual object from a pre-set deformer template, the deformer including the mapping relationship between the pronunciation mouth shape and the skeleton model, the pronunciation mouth shape including the initial consonant mouth shape and / or final vowel mouth shape constructed based on the combination of Chinese Pinyin rules; The generation module is used to generate an amplitude curve corresponding to the pronunciation lip shape based on the dubbing material. The amplitude curve is used to indicate the audio amplitude corresponding to each phoneme in the dubbing material. Each phoneme in the dubbing material corresponds one-to-one with the initial consonant lip shape and / or final vowel lip shape in the pronunciation lip shape. The dubbing material is mapped to the skeletal model of the virtual object through the deformer to generate a facial lip shape image synchronized with the dubbing material. The facial lip shape image is adjusted to the lip shape image of the virtual object through the amplitude curve.
Citation Information
Patent Citations
Automatic generating system for role Chinese mouth shape cartoon
CN101826216A
Method and device for determining lip shape of virtual character, equipment and computer storage medium
CN112131988A