Lip synchronization method and apparatus, computing device cluster, and storage medium
By establishing a correlation between feature information matching in audio and media data, the problem of lip-syncing errors in multi-person dialogue videos was solved, achieving higher lip-syncing accuracy.
Patent Information
- Application Number
- PCT/CN2025/080832
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-09-14
- Filing Date
- 2025-03-05
- Publication Date
- 2025-12-26
AI Technical Summary
Existing technologies suffer from lip-sync errors and inconsistencies in multi-person dialogue video frames, making it difficult to accurately synchronize the lip movements of different individuals.
By acquiring audio and media data, and matching the feature information of audio segments and target images, a correlation is established. Lip-syncing is then performed on the audio segments and target images of the same object to ensure that the mouth area image of each object is a lip-sync image driven by the corresponding audio segment.
It improves the accuracy of lip-sync in multi-person dialogue scenarios, avoids lip-sync errors and confusion, and enhances the accuracy of lip-sync.
Smart Images

Figure CN2025080832_26122025_PF_FP_ABST
Abstract
Description
Lip-sync methods, devices, computing equipment clusters, and storage media
[0001] This application claims priority to Chinese Patent Application No. 202410796428.5, filed on June 19, 2024, entitled "Lip-syncing method and lip-syncing device", and to Chinese Patent Application No. 202410796428.5, filed on September 14, 2024, entitled "Lip-syncing method, device, computing device cluster and storage medium", the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of cloud technology, and in particular to a lip-syncing method, apparatus, computing device cluster, and storage medium. Background Technology
[0003] With the rapid development of digital media and virtual reality technologies, users have increasingly higher demands for the realism and interactivity of media data. Currently, audio data can be used to drive the lip movements of faces in video data, enabling the lip movements of faces in video data to synchronize with the audio of the corresponding individuals in the audio data.
[0004] Currently, the main approach is to align audio and video data in time, using the time-aligned audio frames to drive the lip movements of characters in the corresponding video frames. However, when a video frame includes multiple characters, such as a scene of multiple people conversing, the lip-syncing solutions in these technologies generally suffer from lip-syncing errors and inconsistencies. Summary of the Invention
[0005] This application provides a lip-syncing method, apparatus, computing device cluster, and storage medium, which can improve the accuracy of lip-syncing for different objects.
[0006] In a first aspect, this application provides a lip-syncing method. The method includes: acquiring audio data and media data, wherein the audio data includes N sets of audio segments corresponding to different objects, where N ≥ 2 and N is a positive integer; the media data includes N sets of target images corresponding to the different objects, each of the N sets of target images being an image containing the mouth region of at least one object; based on the audio data and the media data, associating a set of audio segments and a set of target images corresponding to the same object to obtain N association relationships between the N sets of audio segments and the N sets of target images; according to the N association relationships, performing lip-syncing on the N sets of target images associated with the N sets of audio segments, respectively, to obtain target media data, wherein the image of the mouth region within a set of target images of the same object in the target media data is a lip-syncing image driven by a set of audio segments associated with that set of target images.
[0007] As one implementation, audio data and media data can be acquired first. The audio data includes a first audio segment corresponding to a first object and a second audio segment corresponding to a second object. The media data includes a first target image corresponding to the first object and a second target image corresponding to the second object. The first target image is an image including the mouth region of the first object, and the second target image is an image including the mouth region of the second object.
[0008] For example, in N groups of audio segments, one group consists of a first audio segment, and the other group consists of a second audio segment. The number of first audio segments corresponding to the first object can be one or more, and the number of second audio segments corresponding to the second object can also be one or more.
[0009] For example, in N sets of target images, one set consists of a first target image, and the other set consists of a second target image. The number of first target images corresponding to the first object can be one or more, and the number of second target images corresponding to the second object can be one or more.
[0010] Then, based on the audio data and the media data, the first audio segment corresponding to the first object and the first target image can be associated, and the second audio segment corresponding to the second object and the second target image can be associated to obtain a first association relationship between the first audio segment and the first target image, and a second association relationship between the second audio segment and the second target image;
[0011] For example, in N sets of associations, the set of associations corresponding to the first object may include the first association, and in N sets of associations, the set of associations corresponding to the second object may include the second association.
[0012] Finally, lip-syncing can be performed on the first target image based on the first audio segment according to the first association relationship, and lip-syncing can be performed on the second target image based on the second audio segment according to the second association relationship, thus obtaining the target media data. The first target image of the first object in the target media data includes a lip-syncing image of the mouth region of the first object driven by the first audio segment, and the second target image of the second object in the target media data includes a lip-syncing image of the mouth region of the second object driven by the second audio segment.
[0013] The first target image of the first object in the target media data (here, the first target image after lip-syncing) may include a lip-sync image, which is an image formed by the mouth area of the first object being driven by the first audio segment; similarly, the second target image of the second object in the target media data (here, the second target image after lip-syncing) may include a lip-sync image, which is an image formed by the mouth area of the second object being driven by the second audio segment.
[0014] For example, media data can be video data, which can be used to drive the lip movements of objects in the video data.
[0015] For example, media data can be an image (e.g., an image containing multiple objects), such as a person, which can enable lip-syncing for multiple people in a single image.
[0016] For example, media data can be an image sequence, which may include two or more images, each of which may contain at least one object. For example, media data may be image 1 and image 2. Image 1 may include the face of a first object (e.g., object 1), and optionally may also include the face of another object (e.g., a second object). Image 2 may include the face of a second object, and optionally may also include the face of another object (e.g., the first object).
[0017] The object can be any object with a mouth, such as a human figure, an anthropomorphic animal figure (with a mouth), an anthropomorphic plant figure (with a mouth), or a cartoon character (e.g., Spider-Man).
[0018] Taking the human figure as an example, the target image can be a face image, a mouth image, or an image of a complete human body, etc., without any restrictions.
[0019] Any target image (e.g., the first target image or the second target image) can be a complete image in the media data, such as a frame from a video data. Alternatively, any target image can be a partially extracted image from the media data, such as a partial image of a face within a video frame of the video data; there are no restrictions on this.
[0020] Taking video data as an example, N sets of target images correspond to different objects. For example, set B1 corresponds to person 1, and set B2 corresponds to person 2. However, each target image in set B1 (an example of the first target image) may include an image of the mouth region of at least one object (including the first object). For example, set B1 includes target image b1, which includes not only the face image of person 1 but also the face image of person 2. For example, the face images of person 1 and person 2 are within a single video frame in the video data.
[0021] For example, a set of target images B1 includes target image b1 (an example of a first target image), and a set of target images B2 (an example of a second target image) includes target image b2 (an example of a second target image). Target image b2 is an image that includes the mouth region of person 2 (e.g., an image extracted from a video frame that only contains the face of person 2). Optionally, target image b2 may also be an image that further includes other people, such as an image of the mouth region of person 1 (e.g., a video frame in the video data that contains not only the face of person 1 but also the face of person 2 can be used as target image b1, and also as target image b2).
[0022] Specifically, target image b2 is an image that includes the mouth area of person 2. Optionally, target image b2 may also include the mouth areas of other objects such as person 1 and person 3.
[0023] In this context, each of the N associations represents a relationship between a set of audio segments and a set of target images corresponding to the same object, ensuring that the N objects corresponding to each association are distinct. For example, the first association is the relationship between the first audio segment and the first target image, corresponding to the first object; the second association is the relationship between the second audio segment and the second target image, corresponding to the second object.
[0024] For example, in Example 1, a set of audio clips A1 and a set of target images B1 are associated with a person 1. Each set of target images B1 only includes the face of person 1 and does not include other objects. In this embodiment, the method can use the set of audio clips A1 to synchronize the lip movements of person 1 in the set of target images B1. Thus, the image of the mouth area in the target object B1 in the synchronized target media data is the lip movement image driven by the set of audio clips A1.
[0025] Regarding the time alignment between each audio frame in the audio segment A1 and each target image in the target image B1, time alignment can be performed in the existing manner, without any restrictions.
[0026] For example, in Example 2, a set of audio segments A1 and a set of target images B1 are associated with person 1. Among them, target image b1 in the set of target images B1 not only has the face of person 1, but also the face of person 2. Then, the method of this application embodiment can use the set of audio segments A1 to synchronize the lip movements of person 1 in each target image in the set of target images B1. For example, the audio frames in the set of audio segments A1 can be used to drive the lip movements of person 1 in target image b1 in the set of target images B1, so that the image of the mouth area in each target image in the set of target images B1 in the target media data is a lip movement image formed based on the corresponding audio frame. For example, the image of the mouth area of person 1 in target image b1 in the set of target images B1 in the target media data is a lip movement image formed based on the audio frame in the set of audio segments A1 associated with the target image b1.
[0027] Thus, even if there are multiple mouth areas in the target image, the mouth area of the same object in the target media data is a lip-sync image driven by an audio segment associated with the target image. For example, in Example 2 above, the lip-sync image of only person 1 in the target image b1 is a lip-sync image driven by an audio segment associated with the target image b1, while the lip-sync image of person 2 in the target image b1 is not driven by an audio segment associated with the target image b1.
[0028] In this way, even if there are multiple faces in the same frame (e.g., video frame) in the media data, the audio frame of person 1 is directly associated with the face image of person 1, so that the audio frame of person 1 can only drive the lip movements of the face image of person 1, and will not drive the lip movements of other people in the frame.
[0029] In this embodiment, audio segments and target images corresponding to the same object in audio data and media data can be associated. According to the association, lip-syncing is performed on the target images in the media data associated with the audio segment to obtain target media data. In this way, the mouth regions in the target images belonging to the same object in the target media data are all lip-sync images driven by the audio segments associated with the target image. Even if there are mouth regions of multiple objects in the picture, such as in a multi-person dialogue scene, the method of this embodiment directly associates audio segments with target images based on the same object, so that the audio segment drives the lip-sync images in the target image. This avoids the situation where the audio of person A is used to lip-sync the face of person B, which can improve the accuracy of lip-syncing and avoid lip-syncing errors and lip-syncing confusion.
[0030] In one possible implementation, the step of associating a set of audio segments and a set of target images corresponding to the same object based on the audio data and the media data to obtain N association relationships between the N sets of audio segments and the N sets of target images includes: obtaining N sets of first feature information of different objects corresponding to the N sets of audio segments based on each set of audio segments in the audio data; obtaining N sets of second feature information of different objects corresponding to the N sets of target images based on each set of target images in the media data; and associating a set of audio segments and a set of target images whose first feature information and second feature information match based on the N sets of first feature information and the N sets of second feature information to obtain N association relationships between the N sets of audio segments and the N sets of target images.
[0031] Among them, there is a one-to-one correspondence between N sets of audio segments and N sets of first feature information, with one set of audio segments corresponding to a set of first feature information of the same object; there is a one-to-one correspondence between N sets of target images and N sets of second feature information, with one set of target images corresponding to a set of second feature information of the same object.
[0032] As one implementation, when associating the first audio segment corresponding to the first object and the first target image based on the audio data and the media data, and associating the second audio segment corresponding to the second object and the second target image to obtain a first association relationship between the first audio segment and the first target image, and a second association relationship between the second audio segment and the second target image, the first feature information of the first object corresponding to the first audio segment can be obtained based on the first audio segment in the audio data, and the first feature information of the second object corresponding to the second audio segment can be obtained based on the second audio segment in the audio data; then, the second feature information of the first object corresponding to the first target image can be obtained based on the first target image in the media data, and the second feature information of the second object corresponding to the second target image can be obtained based on the second target image in the media data; and the first audio segment and the first target image whose first feature information and second feature information match can be associated to obtain a first association relationship between the first audio segment and the first target image; the second audio segment and the second target image whose first feature information and second feature information match can be associated to obtain a first association relationship between the second audio segment and the second target image.
[0033] As an example, the first feature information and the second feature information match each other, which can be represented as the first feature information and the second feature information being similar, for example, the feature similarity between the two feature information is greater than a preset threshold.
[0034] The following relationship exists between the related audio segments and the target image: the first feature information of the object corresponding to the audio segment (e.g., the first object) and the second feature information of the object corresponding to the target image (e.g., the first object) are mutually matched.
[0035] In this embodiment, the first feature information represents feature information about the corresponding object determined based on the audio segment, while the second feature information represents feature information about the corresponding object determined based on the target image. That is, the data sources of the first feature information and the second feature information are different.
[0036] The matching between two sets of feature information may include, but is not limited to: the proportion of similar features between the two sets of feature information is greater than a certain proportion threshold (e.g., 60% and above, which is not limited here).
[0037] In this embodiment, the feature information of audio and the feature information of image can be used to associate audio segments with matching feature information and target images, thereby associating target images and audio segments corresponding to the same object and realizing automatic association between images and audio of the same object.
[0038] In one possible implementation, the first feature information includes at least one of the following: facial feature information, gender feature information, age feature information, and object type feature information.
[0039] In one possible implementation, the second feature information includes at least one of the following: facial feature information, gender feature information, age feature information, and object type feature information.
[0040] Among them, facial feature information can be a face image or information describing facial features (such as scars, double eyelids, single eyelids, thick eyebrows, etc., there are no restrictions here).
[0041] Among them, gender characteristic information can be text describing gender (e.g., gender is male or female) or features that characterize gender.
[0042] For example, "girl" and "female" are two gender-related features with semantic similarity.
[0043] Age characteristic information can be age type (e.g., infant, teenager, adult, elderly, etc.) or age group or other information that can characterize age.
[0044] Object type feature information can be feature information describing the type of an object. For example, the object type feature information is the object type, which can be human, animal, plant, or cartoon character. Among them, animals can be further subdivided into, for example, pig, cow, horse, etc.
[0045] In this application embodiment, whether it is the first feature information obtained based on the audio segment or the second feature information obtained based on the target image, it can be feature information of at least one dimension such as facial features, gender features, age features, object type features, etc. By associating the audio segment and the target image that match the feature information of the corresponding dimension, the accuracy of the association relationship obtained in this application embodiment is improved, thereby improving the accuracy of lip-syncing of the object and reducing the error rate of using the audio of character A to synchronize the lip-sync of the face image of character B.
[0046] In one possible implementation, obtaining the first feature information of the first object corresponding to the first audio segment based on the first audio segment in the audio data includes: in response to a received first user operation, obtaining the first tag of the first object corresponding to the first audio segment, wherein the first user operation indicates the first tag, and the first tag of the first object indicates the first feature information of the first object corresponding to the first audio segment.
[0047] For example, the first feature information of the first object can be the first label of the first object. The user can input the first label of the first object in the system interface. The first label can indicate the feature information of the first object, such as at least one of the above-mentioned facial feature information, gender feature information, age feature information, and object type feature information.
[0048] For example, users can input the gender (e.g., female), age (e.g., young), facial image, and object type (e.g., human) of the first object into the system's interface. Alternatively, the system can provide a list of options for the above-mentioned characteristics, which the user can select from to trigger a first user action.
[0049] Similarly, in one possible implementation, obtaining the first feature information of the second object corresponding to the second audio segment based on the second audio segment in the audio data includes: in response to a received user operation, obtaining the first tag of the second object corresponding to the second audio segment, wherein the user operation indicates the first tag of the second object, and the first tag of the second object indicates the first feature information of the second object corresponding to the second audio segment.
[0050] Thus, the method of this application embodiment can respond to user input operations to obtain a first tag set by the user for an audio segment of at least one object. The first tag indicates the feature information of the object corresponding to the audio segment, so as to improve the recognition accuracy of the first feature information of the object corresponding to the audio segment, and thereby improve the accuracy of lip-syncing of different objects in the same scene.
[0051] In one possible implementation, obtaining the first feature information of the second object corresponding to the second audio segment based on the second audio segment in the audio data includes: using an AI algorithm to identify the audio features of the second audio segment; and determining a first tag of the second object corresponding to the second audio segment based on the audio features, wherein the first tag of the second object indicates the first feature information of the second object corresponding to the second audio segment.
[0052] Similarly, in one possible implementation, obtaining the first feature information of the first object corresponding to the first audio segment based on the first audio segment in the audio data includes: using an AI algorithm to identify the audio features of the first audio segment; and determining a first tag of the first object corresponding to the first audio segment based on the audio features, wherein the first tag of the first object indicates the first feature information of the first object corresponding to the first audio segment.
[0053] The audio features can be at least one of the following: voiceprint features, pitch features, timbre features, etc. There are no restrictions here.
[0054] In this embodiment, the first feature information extracted from the audio segment corresponding to each object can be set based on user input operation as in the previous embodiment, or an AI algorithm can be used to automatically identify the audio features of the audio segment in order to determine the tag of the object corresponding to the audio segment.
[0055] In a specific implementation, when obtaining the first feature information of N objects from N sets of audio segments corresponding to N objects, it can be set by manual input as in the previous embodiment, or it can be done by AI algorithm to automatically identify audio features in order to obtain the first tags of N objects corresponding to N sets of audio segments.
[0056] In this embodiment of the application, an AI algorithm can be used to automatically identify the audio features of audio segments of one or more objects in the audio data, so as to determine the first label of the one or more objects. The first label can indicate the first feature information of the corresponding object, so as to improve the lip-sync efficiency.
[0057] In one possible implementation, obtaining second feature information of the first object corresponding to the first target image based on the first target image in the media data includes: in response to a received second user operation, obtaining a second tag of the first object corresponding to the first target image, wherein the second user operation indicates the second tag of the first object, and the second tag of the first object indicates the second feature information of the first object corresponding to the first target image.
[0058] Similarly, in one possible implementation, obtaining the second feature information of the second object corresponding to the second target image based on the second target image in the media data includes: in response to a received user operation, obtaining the second tag of the second object corresponding to the second target image, wherein the user operation indicates the second tag of the second object, and the second tag of the second object indicates the second feature information of the second object corresponding to the second target image.
[0059] For example, the second feature information can be a second label. The user can input a second label on the target image of at least one object in the system interface. The second label can indicate the second feature information of the object, such as at least one of the above-mentioned facial feature information, gender feature information, age feature information, and object type feature information.
[0060] For example, the user can input the gender (e.g., female), age (e.g., young), facial image, and object type (e.g., human) of the object corresponding to each of the N sets of target images in the system's interface. Alternatively, the system can provide a list of options for the above-mentioned features, which the user can select from to trigger the user operation in this embodiment.
[0061] Thus, the method of this application embodiment can respond to user input operations to obtain a second tag set by the user for the target object of at least one object. The second tag indicates the feature information of the object corresponding to the target image of the object, so as to improve the recognition accuracy of the second feature information of the object corresponding to the target image, and thereby improve the accuracy of lip-syncing of different objects in the same picture.
[0062] In one possible implementation, obtaining the second feature information of the second object corresponding to the second target image based on the second target image in the media data includes: using an AI algorithm to identify image features of the second target image; and determining a second label of the second object corresponding to the image features based on the image features, wherein the second label of the second object indicates the second feature information of the second object corresponding to the second target image.
[0063] Similarly, in one possible implementation, obtaining the second feature information of the first object corresponding to the first target image based on the first target image in the media data includes: using an AI algorithm to identify image features of the first target image; and determining a second label of the first object corresponding to the image features based on the image features, wherein the second label of the first object indicates the second feature information of the first object corresponding to the first target image.
[0064] The image features of the target image can be at least one of the following: the mouth features of the object, the face features of the object, the body features of the object, etc. There are no restrictions here.
[0065] In this embodiment, the first feature information extracted from the audio segment corresponding to each object can be set based on user input operation as in the previous embodiment, or an AI algorithm can be used to automatically identify the audio features of the audio segment in order to determine the tag of the object corresponding to the audio segment.
[0066] In a specific implementation, when obtaining the second feature information of N objects from N sets of target images corresponding to N objects, it can be set by manual input as in the previous embodiment, or it can be done by AI algorithm to automatically identify image features in order to obtain the second labels of N objects corresponding to N sets of target images.
[0067] In this embodiment of the application, an AI algorithm can be used to automatically identify the image features of the target image corresponding to one or more objects in the media data, so as to determine the second label of the one or more objects. The second label can indicate the second feature information of the corresponding object, so as to improve the lip-syncing efficiency.
[0068] In one possible implementation, before obtaining target media data by lip-syncing the first target image based on the first audio segment according to the first association relationship, and by lip-syncing the second target image based on the second audio segment according to the second association relationship, the method further includes: outputting information indicating the first association relationship and the second association relationship; adjusting the first association relationship and the second association relationship in response to a received third user operation to obtain an adjusted first association relationship and an adjusted second association relationship, wherein the third user operation instructs to adjust the first audio segment in the first association relationship to the second audio segment in the second association relationship, or instructs to adjust the first target image in the first association relationship to the second target image in the second association relationship.
[0069] Users can make secondary adjustments to at least two of the N relationships automatically determined by the system to ensure the accuracy of the N relationships. Specifically, users only need to indicate the adjusted audio segment or the adjusted target image within each of the at least two relationships through user input.
[0070] For example, the system automatically generates three associations: a set of audio segments 1 is associated with a set of target images 1 (e.g., described as association 1); a set of audio segments 2 is associated with a set of target images 2 (e.g., described as association 2); a set of audio segments 3 is associated with a set of target images 3 (e.g., described as association 3);
[0071] For example, the user input operation is: adjust a set of target images associated with a set of audio clips 1 to a set of target images 3, where the set of target images 3 was originally associated with a set of audio clips 3. Therefore, the user operation indicates the adjusted set of target images in association 1 (here, a set of target images 3) and the adjusted set of target images in association 3 (here, a set of target images 1).
[0072] In this way, if the automatically generated association is not accurate enough, the user can make a secondary adjustment to the association to ensure the accuracy of lip-sync.
[0073] In one possible implementation, the first target image is a high dynamic range (HDR) image, and the step of lip-syncing the first target image based on the first audio segment according to the first association relationship includes: obtaining a first low dynamic range (LDR) lip-sync image based on the first audio segment, wherein the first LDR lip-sync image is a lip-sync image of the first object driven by the first audio segment; converting the first LDR lip-sync image into a first HDR lip-sync image; and lip-syncing the lip-sync images of the first object in the first target image that is associated with the first audio segment based on the first HDR lip-sync image according to the first association relationship, thereby obtaining target media data.
[0074] Similarly, in one possible implementation, the second target image is an HDR image, and the step of lip-syncing the second target image based on the second audio segment according to the second association relationship includes: obtaining a second LDR lip-sync image based on the second audio segment, wherein the second LDR lip-sync image is a lip-sync image of the second object formed based on the second audio segment; converting the second LDR lip-sync image into a second HDR lip-sync image; and, according to the second association relationship, performing lip-syncing (e.g., replacing the image of the mouth area) on the lip-sync image of the second object in the second target image that is associated with the second audio segment based on the second HDR lip-sync image to obtain target media data.
[0075] Alternatively, in other embodiments, the media data may also be an image that combines HDR and LDR. For example, if the first target image is an HDR image and the second target image is an LDR image, then it is only necessary to perform HDR to LDR conversion processing on the first target image according to the previous implementation method.
[0076] In this embodiment, lip-syncing of HDR media data can be supported. Specifically, after obtaining an LDR lip-sync image using an audio clip, the LDR lip-sync image can be converted into an HDR lip-sync image. In the HDR space, the HDR lip-sync image generated by the audio driver can be used to perform lip-syncing on the target image of the corresponding object. For example, the HDR lip-sync image obtained by the audio driver and converted can be used to replace the mouth area of the corresponding object in the target HDR image. In this way, it is only necessary to convert the lip-sync image obtained by the audio driver from the LDR space to the HDR space to support lip-syncing of the image in the HDR space.
[0077] In one possible implementation, the media data is HDR media data. The step of associating the first audio segment corresponding to the first object and the first target image, and associating the second audio segment corresponding to the second object and the second target image, based on the audio data and the media data, to obtain a first association relationship between the first audio segment and the first target image, and a second association relationship between the second audio segment and the second target image, includes: converting the brightness range of the HDR media data from HDR to LDR to obtain LDR media data, wherein the LDR media data includes a first LDR target image corresponding to the first object and a second LDR target image corresponding to the second object; and associating the first audio segment corresponding to the first object and the first LDR target image, and associating the second audio segment corresponding to the second object and the second LDR target image, based on the audio data and the LDR media data, to obtain a first association relationship between the first audio segment and the first LDR target image, and a second association relationship between the second audio segment and the second LDR target image.
[0078] In this embodiment, HDR media data can be converted from HDR to LDR to obtain LDR media data. Then, the target image is also an LDR target image (referred to as a target LDR image). In this way, LDR media data can be obtained, but lip-syncing of the input HDR media data is supported.
[0079] In one possible implementation, the step of lip-syncing the first target image based on the first audio segment according to the first association relationship, and lip-syncing the second target image based on the second audio segment according to the second association relationship, to obtain target media data, includes: lip-syncing the first LDR target image based on the first audio segment according to the first association relationship, and lip-syncing the second LDR target image based on the second audio segment according to the second association relationship, to obtain target LDR media data; and converting the brightness range of the target LDR media data from LDR to HDR to obtain target HDR media data.
[0080] In conjunction with the previous implementation method, the lip-synced LDR media data can be converted from LDR to HDR to obtain the target HDR media data, thus supporting lip-syncing of HDR media data.
[0081] In one possible implementation, the media data is HDR media data, the first target image in the HDR media data is a first HDR target image, and the second target image in the HDR media data is a second HDR target image; the step of performing lip-syncing on the first target image based on the first audio segment according to the first association relationship, and performing lip-syncing on the second target image based on the second audio segment according to the second association relationship, to obtain target media data, includes: performing lip-syncing inference on the first audio segment using an AI model, and performing lip-syncing inference on the second audio segment to obtain a first HDR lip-syncing image corresponding to the first audio segment and a second HDR lip-syncing image corresponding to the second audio segment; performing lip-syncing on the lip-syncing image of the first object in the first HDR target image that is associated with the first audio segment according to the first association relationship, and performing lip-syncing on the lip-syncing image of the first object in the second HDR target image that is associated with the second audio segment according to the second association relationship, to obtain target HDR media data.
[0082] The first HDR lip-sync image in this embodiment is obtained directly using an AI model and audio inference, which differs from the first HDR lip-sync image obtained by converting between HDR and LDR spaces in the above embodiment.
[0083] In this embodiment of the application, an AI model can be used to perform lip-syncing inference on audio segments to obtain lip-syncing images in HDR space. Then, based on N association relationships, lip-syncing can be performed on HDR target images (e.g., first HDR target image, second HDR target image) according to the HDR lip-syncing images. An AI model can be used to support lip-syncing of HDR media data.
[0084] In one possible implementation, the method further includes: performing audio segmentation on the audio data for different objects to obtain a first audio segment corresponding to a first object and a second audio segment corresponding to a second object.
[0085] In other embodiments, the method of this application embodiment can also directly obtain the segmented N groups of audio segments from other terminals or devices, which is not limited here.
[0086] In one possible implementation, the method further includes: identifying different objects in the media data to obtain a first target image corresponding to the first object and a second target image corresponding to the second object.
[0087] In other embodiments, the method of this application embodiment may also directly obtain the N sets of target images corresponding to different objects from other terminals or devices without performing different object recognition processing, which is not limited here.
[0088] Secondly, this application provides a lip-syncing device, comprising: an acquisition module for acquiring audio data and media data, wherein the audio data includes a first audio segment corresponding to a first object and a second audio segment corresponding to a second object, and the media data includes a first target image corresponding to the first object and a second target image corresponding to the second object, wherein the first target image is an image including the mouth region of the first object, and the second target image is an image including the mouth region of the second object; and an association module for associating the first audio segment corresponding to the first object and the first target image based on the audio data and the media data, and associating the second audio segment corresponding to the second object and the second target image. The system obtains a first association relationship between the first audio segment and the first target image, and a second association relationship between the second audio segment and the second target image; a lip-sync module is used to perform lip-sync on the first target image based on the first audio segment according to the first association relationship, and to perform lip-sync on the second target image based on the second audio segment according to the second association relationship, to obtain target media data, wherein the first target image of the first object in the target media data includes a lip-sync image formed by the mouth region of the first object based on the first audio segment, and the second target image of the second object in the target media data includes a lip-sync image formed by the mouth region of the second object based on the second audio segment.
[0089] In one possible implementation, the association module is specifically configured to: obtain first feature information of the first object corresponding to the first audio segment in the audio data; obtain first feature information of the second object corresponding to the second audio segment in the audio data; obtain second feature information of the first object corresponding to the first target image in the media data; obtain second feature information of the second object corresponding to the second target image in the media data; associate the first audio segment and the first target image whose first feature information and second feature information match to obtain a first association relationship between the first audio segment and the first target image; associate the second audio segment and the second target image whose first feature information and second feature information match to obtain a first association relationship between the second audio segment and the second target image.
[0090] In one possible implementation, both the first feature information and the second feature information include at least one of the following: facial feature information, gender feature information, age feature information, and object type feature information.
[0091] In one possible implementation, the association module is specifically used to: in response to a received first user operation, obtain a first tag of the first object corresponding to the first audio segment, wherein the first user operation indicates the first tag, and the first tag of the first object indicates feature information of the first object corresponding to the first audio segment.
[0092] In one possible implementation, the association module is specifically used to: use an AI algorithm to identify the audio features of the second audio segment; and based on the audio features, determine a first tag of the second object corresponding to the second audio segment, wherein the first tag of the second object indicates the feature information of the second object corresponding to the second audio segment.
[0093] In one possible implementation, the association module is specifically used to: in response to a received second user operation, obtain a second tag of the first object corresponding to the first target image, wherein the second user operation indicates the second tag, and the second tag of the first object indicates feature information of the first object corresponding to the first target image.
[0094] In one possible implementation, the association module is specifically used to: use an AI algorithm to identify image features of the second target image; and based on the image features, determine a second label of the second object corresponding to the image features, wherein the second label of the second object indicates feature information of the second object corresponding to the second target image.
[0095] In one possible implementation, the apparatus further includes: an output module for outputting information indicating the first association relationship and the second association relationship; the association module is further configured to adjust the first association relationship and the second association relationship in response to a received third user operation to obtain an adjusted first association relationship and an adjusted second association relationship, wherein the third user operation instructs to adjust a first audio segment in the first association relationship to a second audio segment in the second association relationship, or instructs to adjust a first target image in the first association relationship to a second target image in the second association relationship.
[0096] In one possible implementation, the first target image is an HDR image, and the lip-sync module is specifically used to: obtain a first LDR lip-sync image based on the first audio segment, wherein the first LDR lip-sync image is a lip-sync image of the first object formed based on the first audio segment; convert the first LDR lip-sync image into a first HDR lip-sync image; and, according to the first association relationship, perform lip-sync on the lip-sync image of the first object in the first target image that is associated with the first audio segment based on the first HDR lip-sync image to obtain target media data.
[0097] In one possible implementation, the media data is HDR media data, and the association module is specifically configured to: convert the brightness range of the HDR media data from HDR to LDR to obtain LDR media data, wherein the LDR media data includes a first LDR target image corresponding to the first object and a second LDR target image corresponding to the second object; based on the audio data and the LDR media data, associate the first audio segment corresponding to the first object with the first LDR target image, and associate the second audio segment corresponding to the second object with the second LDR target image to obtain a first association relationship between the first audio segment and the first LDR target image, and a second association relationship between the second audio segment and the second LDR target image.
[0098] In one possible implementation, the lip-sync module is specifically configured to: perform lip-sync on the first LDR target image based on the first audio segment according to the first association relationship, and perform lip-sync on the second LDR target image based on the second audio segment according to the second association relationship to obtain target LDR media data; and convert the brightness range of the target LDR media data from LDR to HDR to obtain target HDR media data.
[0099] In one possible implementation, the media data is HDR media data, the first target image in the HDR media data is a first HDR target image, and the second target image in the HDR media data is a second HDR target image; the lip-sync module is specifically used to: perform lip-sync inference on the first audio segment using an AI model, and perform lip-sync inference on the second audio segment to obtain a first HDR lip-sync image corresponding to the first audio segment and a second HDR lip-sync image corresponding to the second audio segment; according to the first association relationship, based on the first HDR lip-sync image, perform lip-sync synchronization on the lip-sync image of the first object in the first HDR target image that is associated with the first audio segment, and according to the second association relationship, based on the second HDR lip-sync image, perform lip-sync synchronization on the lip-sync image of the first object in the second HDR target image that is associated with the second audio segment, to obtain target HDR media data.
[0100] In one possible implementation, the acquisition module is specifically used to perform audio segmentation on the audio data for different objects, to obtain a first audio segment corresponding to a first object and a second audio segment corresponding to a second object.
[0101] In one possible implementation, the acquisition module is specifically used to identify different objects in the media data to obtain a first target image corresponding to the first object and a second target image corresponding to the second object.
[0102] The effects of the lip-syncing devices in the above embodiments are similar to those of the lip-syncing methods in the above embodiments, and will not be repeated here.
[0103] Thirdly, embodiments of this application provide a computing device cluster, including at least one computing device, each computing device including a processor and a memory. The processor of the at least one computing device is used to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster performs the lip-sync method in the first aspect or any possible implementation of the first aspect.
[0104] The effect of the computing device cluster in this embodiment is similar to that of the lip-syncing methods in the above embodiments, and will not be described again here.
[0105] Fourthly, embodiments of this application provide a computer program product containing instructions that, when executed by a cluster of computing devices, cause the cluster of computing devices to perform the lip-sync method in the first aspect or any possible implementation thereof.
[0106] The effect of the computer program product in this embodiment is similar to that of the lip-syncing methods in the above embodiments, and will not be described again here.
[0107] Fifthly, embodiments of this application provide a computer-readable storage medium including computer program instructions, which, when executed by a cluster of computing devices, enable the cluster of computing devices to perform the lip-sync method in the first aspect or any possible implementation thereof.
[0108] The effect of the computer-readable storage medium in this embodiment is similar to that of the lip-syncing methods in the above embodiments, and will not be repeated here. Attached Figure Description
[0109] Figure 1a is a schematic diagram illustrating the training process of an AI model as an example;
[0110] Figure 1b is a schematic diagram illustrating the reasoning process of an AI model as an example;
[0111] Figure 2 is a schematic diagram illustrating the lip-syncing process of a related technology;
[0112] Figure 3 is a schematic diagram of the architecture of an exemplary lip-sync system;
[0113] Figure 4 is a schematic diagram of the lip-syncing method as an example;
[0114] Figure 5a is a schematic diagram illustrating an application scenario;
[0115] Figure 5b is a schematic diagram illustrating an application scenario;
[0116] Figure 5c is a schematic diagram illustrating an application scenario;
[0117] Figure 5d is a schematic diagram illustrating an application scenario;
[0118] Figure 5e is a schematic diagram illustrating an application scenario;
[0119] Figure 6 is a schematic diagram illustrating an example application scenario;
[0120] Figure 7 is a schematic diagram of the lip-syncing method as an example;
[0121] Figure 8 is a schematic diagram of the lip-sync method as an example;
[0122] Figure 9 is a schematic diagram of the structure of an exemplary lip-syncing device;
[0123] Figure 10 is a schematic diagram of the structure of an exemplary computing device;
[0124] Figure 11 is a schematic diagram of the structure of an exemplary computing device;
[0125] Figure 12 is a schematic diagram of the structure of a computing device cluster as an example. Detailed Implementation
[0126] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.
[0127] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone.
[0128] The terms "first" and "second," etc., used in the specification and claims of this application are used to distinguish different objects, not to describe a specific order of objects. For example, "first target object" and "second target object," etc., are used to distinguish different target objects, not to describe a specific order of target objects.
[0129] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0130] In the description of the embodiments in this application, unless otherwise stated, "multiple" means two or more. For example, multiple processing units means two or more processing units; multiple systems means two or more systems.
[0131] Before describing the technical solutions of the embodiments of this application, a brief introduction to the background technology and technical terms involved in the embodiments of this application will be given first:
[0132] Lip-sync, based on audio data, drives the lip movements in video data or images to change synchronously.
[0133] Low-Dynamic Range (LDR) with a brightness range of [0, 255].
[0134] High-Dynamic Range (HDR) has a brightness range of [0, +∞).
[0135] Public cloud is a cloud platform provided by a third-party public cloud provider to a wide range of individuals or businesses. In a public cloud, the hardware, software, and other infrastructure are owned and managed by the third-party public cloud provider.
[0136] A private cloud is a dedicated cloud platform provided to a single enterprise or organization. It can be operated internally by the respective enterprise or organization. Private clouds are primarily geared towards enterprise users and are also known as enterprise clouds.
[0137] Hybrid cloud refers to a cloud platform comprised of different cloud platforms. A hybrid cloud typically includes at least two cloud platforms, also known as a multi-cloud platform or multi-cloud. Optionally, a hybrid cloud integrates public and private clouds. For security reasons, some enterprise users prefer to store data in a private cloud while simultaneously wanting access to the computing resources of a public cloud. In this context, hybrid clouds, which combine public and private clouds, are increasingly being adopted. Hybrid clouds combine and match public and private clouds to achieve optimal performance.
[0138] The technical solutions of the embodiments of this application will be described below using a public cloud as an example. In other embodiments, the technical solutions of the embodiments of this application can also be applied to private clouds and / or hybrid clouds, and this application does not limit them.
[0139] A cloud management platform, also known as a cloud platform or simply a cloud management platform, is a software system provided by a cloud provider for managing the infrastructure that provides cloud services. Specifically, the cloud management platform provides an interface related to cloud services for tenants to remotely access them. Tenants can log in to the cloud management platform using a pre-registered account and password on the cloud service access page, and after successful login, select and purchase the corresponding cloud services. In this embodiment, a tenant includes at least one user. The tenant has a main account, and the user has sub-accounts. In this embodiment, the tenant and user can be arbitrarily replaced, and this will not be repeated below.
[0140] Cloud services include computing services, storage services, virtual machine services, and network services. Any device or function that a tenant's client can access on the cloud management platform can be considered a service provided by the cloud management platform.
[0141] Infrastructure refers to the hardware devices used to implement cloud (taking public cloud as an example) systems and provide various cloud services. It may include multiple cloud data centers (DCs) located in different geographical regions, with at least one data center in each region. Each cloud data center contains multiple physical servers, each of which can support various cloud services such as virtual machines (VMs), containers (Docker), bare metal servers, and cloud disks. The physical servers within a cloud data center can be compute nodes, storage nodes, and network nodes. Compute nodes can be compute servers, storage nodes can be storage servers, and network nodes can include, but are not limited to, network facilities such as switches. For example, the hardware resources corresponding to software resources such as compute services, storage services, or network services include compute resources, storage resources, or network resources. Compute resources can be deployed on compute nodes, storage resources can be deployed on storage nodes, and network resources can be deployed on network nodes. Furthermore, the cloud management platform communicates with the infrastructure, thus enabling the cloud management platform to provide various cloud services supported by the infrastructure to tenants.
[0142] Computing resources may include central processing unit (CPU) resources, memory resources, and / or hard disk resources. For example, a user (also known as a tenant) purchases a large number of virtual machine resources, and a large number of applications (also known as cloud applications) are deployed on these virtual machine resources. In this embodiment, both the virtual machines and the applications within them belong to cloud services (i.e., software resources), while the servers and other devices to which the virtual machines belong are the corresponding hardware resources.
[0143] With the rapid development of digital media and virtual reality technology, users have increasingly higher requirements for the realism and interactivity of media data. Artificial intelligence (AI) lip-sync (LS) technology has emerged. The core principle of AI lip-sync technology can be mainly divided into the following steps: (1) Audio analysis: Analyze and process the captured audio data to extract audio features such as phonemes, rhythm, and intonation, so as to provide a reference for lip-sync. (2) Lip generation: Based on the trained AI model, generate lip movements synchronized with the audio features obtained from audio analysis. The lip movement can be a two-dimensional image of the lip shape (referred to as a two-dimensional lip image) or a three-dimensional asset of the lip shape. The three-dimensional asset can include geometric data and material maps. The geometric data can be a triangular mesh or a position map. (3) Synchronization adjustment: Based on the two-dimensional lip image corresponding to the lip movement generated by the AI model, replace the lip region in the video data to be processed or the lip region in the image to be processed with the two-dimensional lip image. When the mouth movement generated by the AI model is a two-dimensional mouth image, the mouth area can be replaced with the two-dimensional mouth image generated by the AI model in step (3); when the mouth movement generated by the AI model is a three-dimensional mouth asset, the three-dimensional mouth asset can be rendered as a two-dimensional mouth image, and then the mouth area in step (3) can be replaced with the rendered two-dimensional mouth image.
[0144] The following section, with reference to the accompanying diagram, describes the specific implementation process of AI lip-sync technology.
[0145] Figures 1a and 1b illustrate the training process of the AI model of this application, and the process of using the trained AI model to achieve AI lip-syncing, respectively.
[0146] Referring to Figure 1a, the training process of an AI model may include the following steps:
[0147] S101, analyzes and processes the audio data to obtain audio features.
[0148] The audio data can be speech data, such as a speech signal.
[0149] For example, key speech features such as phonemes, rhythm, and intonation can be extracted from speech signals to provide accurate speech features for lip-syncing.
[0150] S102, Obtain a reference lip shape image.
[0151] The reference lip shape image is a two-dimensional image that provides a reference for the lip shape image to be generated.
[0152] Furthermore, the reference lip-sync image is a lip-sync image synchronized with the audio feature.
[0153] In this way, through S101 and S102, mutually synchronized reference lip-sync images and audio features can be obtained. A set of mutually synchronized reference lip-sync images and audio features can constitute a training data set. Multiple training data sets can be obtained in this way, thus obtaining a training dataset.
[0154] In this training dataset, an audio feature from a training data point can be used as a label for a reference lip-sync image to guide the training of the AI model.
[0155] S103, the audio features in the training dataset are input into the AI model, and the AI model performs forward calculations to obtain the predicted lip-sync image.
[0156] The AI model can be a Transformer-based model, such as the large language model Meta AI (LLaMA), bidirectional encoder representations from transformers (BERT), generative pre-trained Transformer 3 (GPT-3), deep learning models (e.g., convolutional neural networks (CNN) or recurrent neural networks (RNN)), or other types of models; there are no restrictions on this. AI models involve a large number of parameters, such as hundreds of billions. Correspondingly, different parameters in an AI model can be distributed across different network layers.
[0157] For example, if the AI model is a deep learning model, its powerful learning capabilities can be leveraged to improve the accuracy of lip-reading prediction and the generalization ability of the AI model.
[0158] In process S103, the AI model can predict the lip-sync image that matches the audio feature based on the audio feature, and output the predicted lip-sync image.
[0159] S104, For the reference lip-sync image synchronized with the audio feature, perform loss calculation with the predicted lip-sync image to obtain the loss.
[0160] S105, use this loss to optimize the AI model.
[0161] For example, this loss can be used to adjust the parameters of an AI model.
[0162] Based on the training dataset, the AI model is trained through multiple rounds of training from S103 to S105 to obtain the trained AI model, which can be used as a two-dimensional lip shape generation model as shown in Figure 1b.
[0163] This AI model, once trained, is able to identify and predict two-dimensional lip-sync images corresponding to different audio features.
[0164] It should be understood that Figure 1a is only an example of the training process of the AI model. In other embodiments, the above training dataset can also be used to train the AI model through other training methods so that the trained AI model can fit two-dimensional lip-sync images based on audio features.
[0165] The following description, in conjunction with Figure 1b, illustrates the process of using this two-dimensional lip-shape generation model for model inference to achieve lip-shape synchronization.
[0166] As shown in Figure 1b, the process may include the following steps:
[0167] S201, analyze and process the audio data to obtain audio features.
[0168] The audio data may be the same as or different from the audio data used for AI model training in Figure 1a; no restrictions are imposed here.
[0169] The audio data includes information describing the audio features used to generate the lip-sync image.
[0170] The audio data in this reasoning scenario is audio data used for lip-syncing, such as the dubbing file of a movie.
[0171] The implementation principle of S201 is the same as that of S101 shown in Figure 1a, and will not be repeated here.
[0172] S202 generates a two-dimensional lip shape image based on audio features using a two-dimensional lip shape generation model.
[0173] Optionally, S203 may be included after S202.
[0174] S203, post-process the two-dimensional lip shape image to obtain the processed lip shape image.
[0175] This post-processing may include, but is not limited to, adjusting the color, shape, texture, and other information of the lip-sync image to meet artistic requirements.
[0176] For example, the color of the lips in a two-dimensional mouth shape image can be adjusted, as can the structure of the teeth in the mouth shape image.
[0177] S204: Use the processed lip-sync image to synchronize the lip-sync of the video (or image) to obtain a synchronized video (or image).
[0178] For example, the audio data in S201 is aligned with the video data here, wherein an aligned audio segment and a video frame have the following relationship: a two-dimensional lip-sync image generated based on the audio features of the audio segment is used to synchronize the lip-sync of the video frame.
[0179] In S204, the lip-sync region in the video frame aligned with the audio segment can be replaced with a two-dimensional lip-sync image generated based on the audio segment, and the audio information of the video frame can be replaced with an audio segment aligned with the video frame to obtain a synchronized video.
[0180] Alternatively, if the object to be synchronized with lip movements is not a video, but an image or an image sequence, taking an image sequence as an example, then you only need to replace the lip movement region in the image frame that is aligned with the audio segment with the processed lip movement image generated based on the audio segment to obtain the synchronized image sequence.
[0181] In this way, AI technology is used to achieve lip-syncing between audio data and video data (or image sequences).
[0182] In Figures 1a and 1b, the lip movements to be synchronized are illustrated using two-dimensional lip images. When the lip movements are three-dimensional assets, the training process of the AI model and the process of using the trained AI model for lip synchronization are largely the same as those described in Figures 1a and 1b, with the only difference being:
[0183] The training process of AI models:
[0184] Difference 1: In the training dataset, the reference data is replaced by a reference lip-shape image with a reference 3D asset for the lip shape. This reference 3D asset can include reference geometric data and reference texture maps for the lip shape. The reference geometric data can be a reference triangle mesh or a reference position map for the lip shape. This reference 3D asset serves as a reference for the lip-shape image to be generated.
[0185] Difference 2: The AI model uses audio features to predict data by replacing the predicted lip-sync image with a predicted 3D asset of the lip-sync.
[0186] The predicted 3D asset for this lip shape also includes the predicted geometry data and the predicted texture map of the lip shape. For details, please refer to the introduction of the reference 3D asset above, which will not be repeated here.
[0187] Difference 3: The trained AI model is a three-dimensional lip-shape generation model, rather than the two-dimensional lip-shape generation model shown in Figure 1a.
[0188] The process of using a trained AI model for lip-syncing:
[0189] Difference 1: The 3D lip-sync generation model is based on audio features to generate 3D assets of the lip shape (specifically including the geometric data and texture maps of the lip shape).
[0190] Difference 2: In the post-processing of S203 shown in Figure 1b, at least one of the data in the lip shape geometry and the lip shape material texture is adjusted to achieve adjustments in color, shape and texture.
[0191] Difference 3: In the lip-sync process of S204 shown in Figure 1b, the lip-sync process can be performed after rendering the above-mentioned 3D lip assets into a 2D lip image; or, in the lip-sync process, the lip image is not replaced, but the 3D lip assets are replaced to obtain the synchronized video (or image).
[0192] The AI lip-syncing technology described above can improve the speed of lip-syncing generation and shorten the generation cycle compared to manual lip-syncing of video and audio.
[0193] Based on this, an AI lip-sync system is provided in the related technology. Figure 2 illustrates the process of lip-syncing by the AI lip-sync system.
[0194] As shown in Figure 2, the process may include the following steps:
[0195] S301 performs time alignment on audio and LDR video data uploaded to the system by users based on user operations.
[0196] For example, in a dubbing scenario, the audio data uploaded by the user is a dubbing file used to dub the characters in the LDR video data.
[0197] Users can manipulate the audio track of the audio data and the video track of the LDR video data, causing the system to respond to the user's actions and align the audio data and LDR video data.
[0198] In Figure 2, dashed double-headed arrows represent the two types of aligned data.
[0199] In this way, audio frames in the audio data can be aligned with video frames in the LDR video data to achieve temporal alignment. Specifically, the relationship between an aligned audio frame and a video frame is as follows: an audio frame is used to synchronize the lip movements of a face in a video frame.
[0200] S302, perform face recognition on the LDR video data uploaded by the user to the system to obtain a face image sequence.
[0201] The system can recognize faces in LDR videos and segment the face regions (e.g., face regions outlined by rectangles) to obtain face images. Each face image contains only one face. Thus, the system can segment a sequence of face images from an LDR video, referred to as a face image sequence. This face image sequence includes multiple face images.
[0202] Furthermore, regardless of whether the LDR video contains the face of one person or the faces of multiple people, the system does not distinguish between people, but only segments according to the individual face to obtain a sequence of face images.
[0203] For example, if a video frame 1 in LDR video data contains the faces of two or more people, the system can segment each face in video frame 1 to obtain two or more face images (e.g., face image A and face image B) corresponding to video frame 1.
[0204] Because LDR video and audio data are time-aligned, and facial image sequences are also aligned with the audio data, when two or more facial images exist in a video frame 1, the audio frame 1 aligned with that video frame 1 will also be time-aligned with those two or more facial images. In other words, the related technology will use the audio frame 1 to synchronize lip movements for the two or more facial images aligned with it.
[0205] S303 uses an AI model to perform lip-reading inference on audio data to obtain a sequence of lip-reading images.
[0206] For example, the lip-sync image sequence corresponding to the audio data (including multiple audio frames) can be obtained by following the steps S201, S202, and S203 shown in Figure 1b. The lip-sync image sequence includes multiple lip-sync images, where each lip-sync image is obtained by the two-dimensional lip-sync generation model inferring the audio features of an audio frame.
[0207] Because the audio data is aligned with the face image sequence, the lip-sync image sequence obtained based on the audio data is also time-aligned with the face image sequence.
[0208] Continuing with audio frame 1 as an example, lip-reading inference using audio frame 1 yields a lip-reading image A'. Since audio frame 1 is time-aligned with video frame 1, it aligns with both face images A and B in video frame 1. Therefore, the lip-reading image A' is aligned with both face images A and B from video frame 1.
[0209] S304, Based on the lip-shape image sequence, replace the lip-shape region in the face image sequence to obtain a face image sequence with synchronized lip-shape.
[0210] In this system, the lip-sync image sequence and the face image sequence are time-aligned. Therefore, the system can use a single aligned lip-sync image (e.g., lip-sync image A') to replace the lip-sync regions of face images aligned with that lip-sync image in the face image sequence (here, face image A of person A from video frame 1 and face image B of person B), thus obtaining a lip-sync synchronized face image. Similarly, by replacing the lip-sync regions in each face image of the face image sequence, a lip-sync synchronized face image sequence can be obtained.
[0211] In cases where a lip-sync image is time-aligned with multiple face images from the same video frame, this scheme cannot determine which face image's lip-sync region should be replaced by that lip-sync image. For example, if the audio data is lip-sync driven by person A's face, it might replace the lip-sync region in person B's face image B with that lip-sync image A', leading to lip-sync errors.
[0212] S305, pastes the synchronized lip-sync face image sequence back into the LDR video data to obtain synchronized lip-sync video data.
[0213] As described in S301, the face image sequence is an image region segmented from LDR video data. In this step, each face image in the face image sequence after lip-syncing can be filled into the corresponding video frame in the LDR video data, thereby realizing the synchronization of audio data with the lip movements of the person in the LDR video data and obtaining video data after lip-syncing.
[0214] Finally, users can download the synchronized lip-sync video data from the system to their local device.
[0215] In this related technology, in a scenario where a video frame in LDR video data contains only an image of a face, the lip-sync system synchronizes the lip movements of the face image in the video frame according to the time sequence, based on an audio frame aligned with the video frame.
[0216] However, when multiple face images exist in the same video frame of LDR video data, such as in a dialogue scene where the same video frame contains face images of two people, A and B, the related technology, after inferring the lip-sync image based on the audio features of the audio frame aligned with the video frame, cannot determine whether to paste the lip-sync image into the face area of person A or person B in the video frame. This can lead to system errors or lip-sync confusion (for example, using the audio frame used to synchronize the lip-sync of person A to synchronize the lip-sync of person B).
[0217] Furthermore, in related technologies, the AI model used to implement lip-sync inference in S303 shown in Figure 2 uses LDR lip-sync images as reference data (e.g., reference lip-sync images) in its training data during training. Because HDR image sample data is limited, it is not conducive to fitting the lip-sync images to the AI model. Therefore, the AI model trained based on LDR lip-sync images can only infer LDR lip-sync images and cannot infer HDR lip-sync images. Thus, related technologies only support AI lip-sync synchronization in LDR space and do not support AI lip-sync synchronization in HDR space. Even if the user-uploaded video data is HDR video data, related technologies cannot synchronize the lip-sync in the HDR video data based on the audio data in HDR space.
[0218] However, in the film and television industry, there are a large number of film and television resources that require post-production dubbing and lip-syncing. Most of these film and television resources are HDR video data, which will cause the relevant technology system to be unable to support lip-syncing of HDR video or images.
[0219] Therefore, the system in the related technology does not support lip-syncing of images or videos in multi-person dialogue scenarios, and is prone to lip-syncing failure or lip-syncing chaos in such scenarios; in addition, the system does not support lip-syncing of images or videos in HDR space.
[0220] In order to solve the aforementioned technical problems existing in the systems of related technologies, this application provides a lip-syncing method, apparatus and system that can solve the aforementioned technical problems.
[0221] As shown in Figure 3, the lip-syncing method and apparatus of this application embodiment can be applied to the cloud. Specifically, the lip-syncing system 10 may include, but is not limited to, a public cloud 200, tenant clients 300, etc. The public cloud 200 may include, but is not limited to, a cloud management platform 201 and infrastructure 202. This application uses a public cloud as an example for system 10. In other embodiments, system 10 is also applicable to private clouds or hybrid clouds, with the same principle, and will not be described further here.
[0222] In some embodiments, the cloud management platform 201 may provide various interfaces, such as login interfaces and data processing interfaces, for tenant clients (e.g., client 300) to access. The client may be a terminal used by the tenant or a browser on that terminal, etc., without limitation. The tenant's terminal is also referred to as an electronic device, user device, or terminal device, and this application does not limit this. The terminal may include, but is not limited to, mobile phones, tablets, computers, personal computers (PCs), etc.
[0223] The public cloud 200 in this application embodiment can provide lip-sync cloud services. After an organization or individual purchases the cloud service, an account and password for using the cloud service can be assigned to that enterprise.
[0224] For example, the cloud management platform 200 can receive the account and password information entered by the tenant through the client 300 via the login interface to authenticate the tenant's client 300. After successful authentication, the tenant's client 300 can be allowed to log in to the cloud management platform 200 to use the lip-sync cloud service.
[0225] In addition, the cloud management platform 200 can also provide a user interface through which the tenant's client 300 can interact with the cloud management platform 201 to realize the lip-sync method of this application.
[0226] For example, after logging into the cloud management platform 200, the tenant's client 300 can input the video and audio data to be used for lip-syncing into the cloud management platform 201 through the user interface.
[0227] The video data can be either LDR or HDR video; there are no restrictions.
[0228] In this context, LDR video indicates that the brightness range of the video frame is LDR, and HDR video indicates that the brightness range of the video frame is HDR.
[0229] In other words, the system 10 of this application supports users uploading HDR videos for lip-syncing in HDR space.
[0230] In other embodiments, the video data can also be replaced with LDR images, HDR images, or LDR image sequences, HDR image sequences, etc. Thus, the system 10 of this application can utilize audio data to perform lip-syncing on the image sequence.
[0231] The audio data uploaded by the tenant's client 300 through the aforementioned user interface can be one or more audio files, and there are no restrictions here.
[0232] Tenants can input the aforementioned video and audio data through the interface, by importing files, or by configuring the software to input the binding relationship between the video and audio data and the character ID to the cloud management platform 201. This application does not impose any restrictions on this. In this way, the cloud management platform 201 can receive the video and audio data to be synchronized with lip movements.
[0233] Then, the cloud management platform 201 can use the lip-sync cloud service provided by the infrastructure 202 to synchronize the lip movements of objects in the video data based on the received audio data and optionally in combination with the above binding relationship, so as to obtain the lip-synced target video.
[0234] The object can be any object with a mouth, such as a human figure, an anthropomorphic animal figure (with a mouth), an anthropomorphic plant figure (with a mouth), or a cartoon character (e.g., Spider-Man).
[0235] Finally, the cloud management platform 201 can send the lip-synced target video to the client 300 through the aforementioned user interface, so that the user can view or download the target video from the client 300.
[0236] In this embodiment, a set of audio segments and a set of target images corresponding to the same object in audio data and media data can be associated. According to the association, lip-syncing is performed on a set of target images in media data associated with the set of audio segments to obtain target media data. In this way, the image of the mouth area in a set of target images of the same object in the target media data is a lip-sync image driven by a set of audio segments associated with the set of target images. Even if there are mouth areas of multiple objects in the image, such as in a multi-person dialogue scene, the method of this embodiment directly associates audio segments with target images of the same object so that the audio segments drive the lip-sync images in the target images. This avoids the situation where the audio of person A is used to lip-sync the face of person B, thus improving the accuracy of lip-syncing and avoiding problems such as lip-syncing errors and lip-syncing confusion.
[0237] Figure 3 illustrates the implementation process of the lip-sync method and apparatus of this application, taking the application of the lip-sync method and apparatus in the cloud as an example. In other embodiments, the lip-sync method and apparatus of this application are not limited to the cloud, but can also be applied to edge devices, terminal devices, and other devices under the cloud.
[0238] Furthermore, the lip-syncing device of this application can be implemented by software or hardware.
[0239] In the first example, when implemented in software, the lip-sync device can be, for example, an application running on a computing instance, such as a virtual machine, container, or host.
[0240] In the second example, when implemented in hardware, the lip-sync device can be implemented through at least one physical device including a processor, such as a server. The processor can be a central processing unit (CPU), or any combination of an application-specific integrated circuit (ASIC), a programmable logic device (PLD), a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), a system-on-chip (SoC), a software-defined infrastructure (SDI) chip, an artificial intelligence (AI) chip, or a data processing unit (DPU). Furthermore, the number of processors included in the lip-sync device can be arbitrary, and the types of processors can be one or more. The specific number and types of processors can be determined according to the actual application's business requirements; this application does not impose any limitations on this.
[0241] In the third example, when implemented in hardware, the lip-sync device can be a computing cluster comprising multiple computing nodes. Further, these multiple computing nodes can communicate through the at least one switching node. Each computing node can be a node with model training / inference capabilities for lip-syncing. Exemplarily, the computing node can be an accelerator card, such as a deep-learning processing unit (DPU), data processing unit (DPU), graphics processing unit (GPU), neural-network processing unit (NPU), or tensor processing unit (TPU), or other types of accelerator cards. Alternatively, the computing node can be a general-purpose processor, such as a central processing unit (CPU).
[0242] Furthermore, the lip-syncing method, apparatus, and system 10 provided in this application embodiment can be applied to multiple scenarios such as film and television, games, and virtual reality.
[0243] 1. Production scenes of film and television resources
[0244] The film and television resources can include movies, TV series, short videos, animations, etc.
[0245] Taking a movie as an example, in order to promote the movie to countries with different languages for broadcast, it may involve using dubbing files in different languages to dub multiple characters in the movie based on the dubbing files and to synchronize the lip movements of multiple characters according to the lip-syncing method provided in the embodiments of this application.
[0246] In this way, regardless of the language of the dubbing data in the movie's dubbing file, the lip movements of the dubbed characters in the movie can be synchronized with the corresponding language dubbing file after the dubbing and lip-syncing are synchronized. This achieves lip-syncing of multiple characters in film and television resources and improves the post-production efficiency of film and television resources.
[0247] 2. Game Scene
[0248] In this game scenario, the video to be synchronized with lip movements is the game's video, which can be a role-playing game or an interactive narrative game. Taking a role-playing game as an example, during gameplay, the player can input voice into an electronic device (such as a mobile phone). The mobile phone can then use the lip-syncing method provided in this application to perform AI inference on the player's character's lip movements based on the input voice, and render the inferred lip-sync images in real time and paste them onto the game's video screen. Synchronization between the character's lip movements and the player's voice in the game can enhance the player's immersion and gaming experience.
[0249] 3. Virtual Reality (VR) / Augmented Reality (AR) / Mixed Reality (MR) Scenes
[0250] In a VR / AR / MR environment, VR / AR / MR devices can use the lip-syncing method provided in the embodiments of this application to synchronize the lip movements of virtual characters in virtual images or virtual characters in images synthesized from virtual and real data, based on the user's input voice, thereby enhancing the user's interactive experience and making the virtual characters more realistic.
[0251] Of course, the above are merely exemplary application scenarios of the lip-syncing method, apparatus and system 10 of this application. The lip-syncing method, apparatus and system 10 can also be applied to other scenarios where there is a need for lip-syncing. This application does not limit the application scenario.
[0252] The implementation process of the lip-sync method of this application will be described below with reference to specific embodiments.
[0253] Please refer to Figure 4, which is a schematic diagram illustrating an exemplary lip-syncing method of this application. The lip-syncing method of this application will be described using the example of its application in the production of film and television resources.
[0254] The method shown in Figure 4 can be implemented based on the cloud architecture shown in Figure 3, but is not limited to this architecture. It can also be implemented through other architectures, which are not restricted here.
[0255] As shown in Figure 4, the process may include the following steps:
[0256] S401a, acquire video data.
[0257] The video data can be either HDR video data or LDR video data.
[0258] The video data obtained in this application may be a video file uploaded by the tenant's client 300 through the user interface as shown in Figure 3, or a video file pre-stored in the local storage device or external storage device of the lip-sync system of this application, or video data obtained after converting user-uploaded LDR video data (e.g., converting LDR video data to HDR video data), and there are no restrictions here.
[0259] This application does not restrict the method of acquiring video data. The video data can be video data obtained directly from external sources, or video data obtained from external sources after certain processing.
[0260] In some embodiments, the lip-sync method of this application may also support uploading a single image (including multiple objects) or an image sequence (including multiple images) for lip-sync. In this way, the video data in S401a can be replaced with the image or image sequence here. The image or image sequence can be data in HDR space or LDR space. The implementation principle of other processes is similar to that of acquiring video data, and will not be elaborated here.
[0261] S401b, acquire audio data.
[0262] The audio data acquired by the lip-sync method of this application can be audio data directly obtained from external sources (such as one or more audio files uploaded by the user), or audio data obtained after certain processing (such as noise reduction) of the audio data directly obtained from external sources. There are no restrictions here.
[0263] This application does not restrict the execution order of S401a and S401b; they can be executed serially or in parallel.
[0264] Similar to related technologies, the lip-syncing method of this application can align audio data obtained through S401b and video data obtained through S401a based on user input to determine which part of the audio data is used for lip-syncing with which part of the video data. Specifically, the alignment can be performed by the user manipulating the audio track and video track (for example, the system 10 shown in Figure 3 receives this operation through the user interface), so that the system 10 of this application responds to the user operation to align the audio and video data.
[0265] Figure 5a illustrates an example of a timing diagram of aligned audio and video data.
[0266] As shown in Figure 5a, the audio data includes audio file 1 and audio file 2. Audio file 1 is aligned with video segment 1 in the video file during time period t1 to t2. Audio file 2 is aligned with video segment 2 in the video file during time period t2 to t3, video segment 3 in the video file during time period t3 to t4, and video segment 4 in the video file during time period t4 to t5, respectively.
[0267] Thus, audio file 1 is used to synchronize and dub the lip movements of characters in video segment 1 during time period t1 to t2 in video file 1; audio file 2 is used to synchronize and dub the lip movements of characters in video segment 1 during time period t2 to t5 in video file 1.
[0268] In the example shown in Figure 5a, the audio and video segments that are aligned with each other have the same duration. For example, the duration of audio file 1 is the same as the duration of time segment t1 to t2 of video file.
[0269] In other embodiments, the durations of the aligned audio and video segments may differ. For example, as shown in Figure 5a, audio file 1 is aligned with a video segment of time period t1 to t2 of the video file, for example, the duration of time period t1 to t2 is 1 minute, but the duration of audio file 1 may be more or less than 1 minute. When the duration of audio file 1 is more than 1 minute, there may be at least two audio frames aligned with the same video frame; when the duration of audio file 1 is less than 1 minute, there may be the same audio frame aligned with at least two video frames; when the duration of audio file 1 is 1 minute, there may be one audio frame aligned with one video frame. Here, the duration of one audio frame is the same as the duration of one video frame.
[0270] Furthermore, the timing of aligning audio and video data can be before or after S402a and S402b, such as during process S403 as shown in Figure 4. This application does not impose any restrictions on this.
[0271] S402a can be executed after S401a.
[0272] S402a, based on video data, identifies N sets of local images corresponding to different objects.
[0273] The local image is an image that includes the mouth region. For example, the local image can be a mouth image or a face image that includes the mouth region. N is an integer greater than or equal to 2.
[0274] The following text uses the image of a human (e.g., a character in a movie scene) and the local image as a face image as an example to illustrate the lip-sync method of this application. When the object is another type of image and the local image is a mouth image or other images including a mouth region, the principle of the implementation process of the lip-sync method is similar, and will not be repeated here.
[0275] The method in this application embodiment can perform face recognition on the video data and segment the recognized face images from the video data. Furthermore, the method in this application embodiment can group the segmented face images according to the biometric features of the faces (e.g., face images with high biometric similarity are grouped together), so as to divide multiple face images of the same person into a single group. In this way, multiple groups of face images corresponding to different people can be obtained.
[0276] In one example, even if the video data includes similar facial images of twins or multiples, the method of this application embodiment can still perform facial recognition and differentiation of different objects, so as to identify, for example, the two faces of twins as facial images of different people.
[0277] Please refer to Figure 5a for further explanation, using the video file shown in Figure 5a as an example of the video data obtained in this application.
[0278] As shown in Figure 5a, the video file consists of video frames of person A during the time period t1 to t2 (excluding time t2), video frames of person B during the time period t2 to t3 (excluding time t3), and video frames of person C during the time periods t3 to t3' (excluding time t3') and t3' to t4. The video frame corresponding to time t3' contains both person A and person C. In other words, time t3' is a video frame from the multi-person conversation scenario described in this application, and this video frame contains two or more facial images appearing simultaneously.
[0279] In S402a, the method of this application embodiment can perform person recognition, segmentation and grouping on the video file shown in FIG5a to obtain a set of face images corresponding to person A, a set of face images corresponding to person B and a set of face images corresponding to person C.
[0280] Figure 5b shows a schematic diagram of the three sets of face images obtained after performing operation S402a on the video file shown in Figure 5a.
[0281] The frames shown in Figure 5b are face images segmented from a video file, not video frames from the video file.
[0282] As shown in Figure 5b, a set of face images A includes 51 frames of face images of person A, which are face images of person A corresponding to the video segment from time t1 to t2 in the video file, represented by frames P1 to P50; and the set of face images A also includes a face image of person A segmented from the video frame at time t3' as shown in Figure 5a, represented by frame P103'.
[0283] It should be understood that this application does not limit the number of frames of facial images contained in the video segment of each time period. The video segment of the time period t1 to t2 mentioned above, which includes 50 frames of facial images, is only an example, and the same applies below.
[0284] As shown in Figure 5b, a set of face images B includes 100 frames of face images of person B, namely: face images of person B corresponding to the video segment of time period t2 to t3 in the video file, represented by frames P51 to P100, and face images of person B corresponding to the video segment of time period t4 to t5 in the video file, represented by frames P131 to P180.
[0285] As shown in Figure 5b, a set of face images C contains 30 frames of face images of person C, which are face images of person C in the video segment corresponding to the time period t2 to t3 (including t3') in the video file, represented by frames P101 to P130.
[0286] The method of this application embodiment can be implemented by using face recognition algorithms or AI algorithms such as AI models when identifying N sets of local images corresponding to different objects based on video data. This application does not limit the implementation process.
[0287] In some embodiments, the method of this application can output the N groups of local images as a group, so that users can browse the grouping results of people in the video file through the system of this application.
[0288] S402b can be executed after S401b.
[0289] S402b, based on audio data, obtains N sets of audio segments corresponding to different objects.
[0290] As an example of implementation, the N sets of audio segments can be data obtained directly based on user input.
[0291] For example, if a user wants to use three audio files (e.g., dubbing files) for three different characters to dub and synchronize lip movements for three characters in a video file uploaded to this system, the user can upload three dubbing files through the aforementioned user interface, with each dubbing file corresponding to a different character. In this way, the method of this embodiment can directly obtain three sets of audio segments corresponding to different characters based on user input. These three sets of audio segments are the aforementioned three audio files, without needing to segment a single audio file for different characters, thus ensuring the accuracy of the N sets of audio segments.
[0292] In this way, when the voiceprints of at least two people are extremely similar, making it difficult to identify the audio segments of each of the two people from the same audio file, multiple audio files uploaded by the user through the user interface, which have been segmented according to different people, can be used to directly obtain multiple sets of audio segments corresponding to different people, avoiding the segmentation error caused by segmenting the audio data of different people.
[0293] As another implementation example, the N sets of audio segments can be obtained by using an AI algorithm to identify different people in the audio data obtained in S401b, in order to segment the audio data into N sets of audio segments corresponding to different people.
[0294] It should be noted that, since the audio data is used for lip-syncing and synchronization of the video data, the number N of the N audio segments corresponding to different objects obtained in this embodiment is the same as the number N of the N face images. This ensures that the number N of people corresponding to the face images identified and segmented from the video data by the method of this application is the same as the number N of people corresponding to the audio segments identified from the audio data. Each face image corresponds to only one person ID, and each audio segment also corresponds to only one person ID.
[0295] It should be understood that the above implementation examples are merely illustrative and are not intended to limit the technical solutions of this application. This application can also obtain the above N sets of audio segments corresponding to different objects based on the input and AI algorithms.
[0296] In some embodiments, the quantity N associated with the aforementioned audio and video data can also be specified by the user to the system of this application through a user interface. For example, if a user wishes to use three voice-over files for three characters (each character having a corresponding voice-over file) to dub and synchronize the lip movements of three main characters in a video file, the user can specify the value of the quantity N as 3 through the user interface. The value of N is the same as the number of voice-over files and the number of characters. In this way, the method of this application embodiment can perform at least one step in S402a and S402b based on the quantity N.
[0297] For example, if there are more than three people in the video data, but the user specifies that the number N is 3 through the user interface, then the method of this application embodiment can use AI algorithms to identify and filter three sets of face images corresponding to the three people from the video data based on factors such as the frequency of appearance of the people and the scene in which they appear. For example, the three people are people who appear frequently in the video data, or people who appear in key scenes in the video data.
[0298] As an implementation example, the implementation process of S402b will be explained below using the audio data shown in Figure 5a as an example.
[0299] The user uploads audio file 1 and audio file 2 to the system of this application embodiment through the user interface, and the user uploads information indicating that audio file 1 is audio data of a single person and information indicating that audio file 2 is audio data of two people through the user interface.
[0300] Subsequently, the method of this application embodiment can identify audio features such as voiceprint and timbre in audio file 2 based on information uploaded by the user, to determine two sets of audio segments corresponding to different people in audio file 2, thereby dividing audio file 2 into two sets of audio segments, each set consisting of audio from a different person. Furthermore, the method of this application embodiment can also, based on audio file 1 uploaded by the user and information indicating that audio file 1 represents audio data from a single person, designate audio file 1 as a set of audio segments corresponding to a third person.
[0301] Alternatively, this method can also use AI algorithms to identify different people in audio file 2, in order to determine the two sets of audio segments in audio file 2 that correspond to different people.
[0302] Specifically, referring to Figures 5a and 5c, the method of this embodiment can divide audio file 1 and audio file 2 into three groups of audio segments according to the above process. One group of audio segments A is audio file 1 aligned with the time period t1 to t2 of the video file. For example, audio file 1 includes 50 audio frames, represented as frames A1 to A50. For example, an audio frame is aligned with a video frame in the video file. For example, frame A1 is aligned with frame P1 shown in Figure 5b. Frame A1 is used to drive the lip movement of frame P1. The other frames are similar and will not be described in detail.
[0303] Please continue to refer to Figures 5a and 5c. Another set of audio segments among the three sets of audio segments is a set of audio segments B. A set of audio segments B includes 100 audio frames aligned with time periods t2 to t3 and t4 to t5 in the video file, represented as frames B1 to B100.
[0304] Please refer to Figures 5a and 5c. Among the three sets of audio segments, another set of audio segments is called audio segment C. Audio segment C includes 50 audio frames aligned with the time period t3 to t4 in the video file, represented as frames C1 to C50.
[0305] Among them, a set of audio segments B and a set of audio segments C are two sets of audio segments corresponding to two people, which are identified and segmented from audio file 2 by the method of this application embodiment. One set of audio segments is the audio of the same person.
[0306] This application does not limit the specific implementation algorithm for segmenting audio data to obtain N groups of audio segments corresponding to different objects. It can be the voiceprint recognition algorithm, timbre recognition algorithm, AI algorithm mentioned above, or other algorithms that can be used to identify audio from different objects.
[0307] In some embodiments, when the user does not upload information about the number of people corresponding to the audio file, the method of this application embodiment can also use AI algorithms or voiceprint recognition algorithms to distinguish the audio of different people. After the two audio files uploaded by the user as shown in FIG5a are identified as having different people, the two audio files are divided into three groups of audio segments. One group of audio segments is the audio of the same person, and the three groups of audio segments are the audio of different people.
[0308] It should be understood that the above examples are only used to illustrate the methods of the embodiments of this application, but the implementation process of the methods of the embodiments of this application is not limited to the above examples.
[0309] Following S402a and S402b, the method of this embodiment can segment N sets of local images corresponding to different objects from video data (e.g., three sets of face images corresponding to three people as shown in FIG. 5b), and segment N sets of audio segments corresponding to different objects from audio data (e.g., three sets of audio segments corresponding to three people as shown in FIG. 5c). However, the method of this embodiment only groups the face images according to different people and segments and groups the audio data in S402a and S402b, but cannot determine which person each set of face data corresponds to, and may also be unable to determine which person each set of audio data corresponds to. Therefore, the method of this embodiment cannot yet associate audio data with the same person ID with local images.
[0310] Therefore, S403 can be executed after S402a and S402b.
[0311] S403, based on N sets of local images corresponding to different objects and N sets of audio segments corresponding to different objects, determine the association relationship between the N sets of local images and the N sets of audio segments.
[0312] This step S403 aims to associate a set of audio segments and a set of local images (e.g., face images) corresponding to the same person in N sets of audio segments and N sets of local images (e.g., face images) to obtain the association relationship between N sets of local images and N sets of audio segments, wherein a set of local images is mapped to a set of audio segments in the association relationship.
[0313] In one possible implementation, the method of this application embodiment can establish the association between N sets of audio segments and N sets of local images based on the audio features (such as pitch, timbre, etc.) of N sets of audio segments and the features of objects (such as facial features, etc., without limitation) of N sets of local images.
[0314] As an implementation example, the method of this application embodiment can perform image feature recognition on N sets of local images to determine information of at least one dimension such as gender, age type, and object type of the N sets of local images, and mark each set of local images in the N sets of local images with the identified information of at least one dimension such as gender, age type, and object type.
[0315] Furthermore, the method in this application embodiment can identify audio features such as timbre and pitch in N groups of audio segments to determine at least one dimension of information (also called a tag) such as gender, age type, and object type for each group of audio segments in the N groups of audio segments, and mark the N groups of audio segments with the identified at least one dimension of information such as gender, age type, and object type.
[0316] Then, based on the dimensional information labeled in each of the N sets of local images and N sets of audio segments, a relationship can be established between a set of local images and a set of audio segments with the same or similar dimensional information, thereby achieving a one-to-one mapping between the N sets of local images and N sets of audio segments.
[0317] It should be understood that the dimensional information identified by the method of this application embodiment for N sets of local images and N sets of audio segments is not limited to the gender, age type, and object type mentioned above. It may also include other information that can associate the audio of the object with the image of the object, including the mouth region (e.g., mouth image or face image), and is not limited here.
[0318] The implementation scheme of this application will be described below for each of the various dimensions mentioned above:
[0319] Gender dimension:
[0320] The method in this application embodiment can perform male and female gender recognition on two sets of face images obtained through S402a to obtain gender labels for the two sets of face images. For example, the label added to one set of face images A as shown in FIG5b is female, and the label added to another set of face images B is male.
[0321] The method in this embodiment identifies audio features such as timbre and pitch in the two sets of audio segments obtained through S402b to determine whether each set of audio segments is male or female audio. For example, the label added to one set of audio segments A as shown in FIG5c is female, and the label added to another set of audio segments B is male.
[0322] Thus, the method of this application embodiment can establish a mapping between a group of face images A as shown in FIG5b and a group of audio segments A as shown in FIG5c, which have the same label; and establish a mapping between a group of face images B as shown in FIG5b and a group of audio segments B as shown in FIG5c.
[0323] Age type dimension:
[0324] Age categories can include, but are not limited to: infants, toddlers, teenagers, adults, and the elderly.
[0325] Because the facial features of people of different age types differ significantly, the method in this application embodiment can identify the age type of at least two of the obtained N sets of facial images and add a label for that age type.
[0326] Furthermore, the audio characteristics of people of different ages vary considerably. For example, infants' speech is mainly characterized by babbling, while toddlers' speech is characterized by frequent pauses, repetitions, and unclear pronunciation. Teenagers' audio is characterized by higher pitch, elderly people's speech is characterized by slowness, and adults' speech is characterized by fluency. Thus, the method of this application embodiment can identify the audio characteristics of N groups of audio segments to determine the age type of each group of audio segments.
[0327] Thus, the method of this application embodiment can combine the age types of N sets of audio segments and N sets of local images to establish the above-mentioned association, so as to associate audio segments with the same or similar age type semantics with local images.
[0328] For example, in the example of the gender dimension mentioned above, N is an integer greater than 2, and among the N characters appearing in the video data, there are at least two characters of the same gender. For example, as shown in Figure 5a, characters A and C are both female. Therefore, not only are the labels of the set of face images A shown in Figure 5b and the set of audio segments A shown in Figure 5c both female, but the labels of the set of face images C shown in Figure 5b and the set of audio segments C shown in Figure 5c are also both female. Using only the gender dimension makes it difficult to establish a mapping between audio segments and face images. Therefore, the method of this application embodiment can also identify the age type of a set of face images B, a set of face images C, a set of audio segments B, and a set of audio segments C respectively. For example, the age type of face image B and audio segment B are both adults, and the age type of face image C and audio segment C are both teenagers.
[0329] Thus, although it is difficult to establish a connection between audio segments and facial images through gender-based labels, the method of this application embodiment can further utilize age-type labels to establish a connection between a set of facial images B and a set of audio segments B, as well as a connection between a set of facial images C and a set of audio segments C.
[0330] Object type:
[0331] Object types can include, but are not limited to: humans, animals, plants, cartoon characters, etc.
[0332] Furthermore, the object type can also include the refined types of the above-mentioned objects. For example, when the object type is animal, the object type can be subdivided into various animals such as lion, tiger, rabbit, cat, and dog; when the object type is plant, the object type can be subdivided into various plants such as flower, grass, and tree; when the object type is cartoon character, the object type can be subdivided into Spider-Man, Iron Man, Sun Wukong, Zhu Bajie, and White Bone Demon.
[0333] Thus, the method of this application embodiment can identify the object type of N sets of local images and N sets of audio segments, add object type labels to them, and use the object type labels as a reference to establish the association between the N sets of local images and N sets of audio segments.
[0334] In any of the above implementation methods and examples, when establishing the association between N sets of audio segments and N sets of local images, the method of this application embodiment can use AI algorithms or AI models to determine the features, labels, etc. of each of the N sets of audio segments and N sets of local images, so as to establish a mapping relationship between audio segments and local images corresponding to the same object.
[0335] As a specific example, referring to Figures 5a to 5c and referring to Figure 5d, the method of this embodiment of the application performs feature recognition on three sets of audio segments and three sets of face images, and associates a set of audio segments A and a set of face images A' corresponding to the same person based on the alignment information between the video file and the audio file 1.
[0336] In this example, as shown in Figures 5b and 5d, the difference between a set of face images A' and a set of face images A is that a set of face images A' does not have frame P103', where frame P103' is the face image of person A in the video frame at time t3' of the video file. However, as shown in Figures 5a and 5c, the video frame at time t3' in the video file is not aligned with audio file 1 (e.g., a set of audio files A), but is aligned with audio frames in a set of audio files C in audio file 2. Therefore, when establishing the association relationship, the method of this embodiment establishes the association relationship between a set of face images A' and a set of audio segments A, as shown in Figure 5d, based on the alignment relationship between audio frames and video frames.
[0337] Of course, in other embodiments, when the video frame at time t3' in the video file shown in FIG5a is also aligned with the timestamp of audio file 1, then the association between the group of face images A and audio file 1, as shown in FIG5b, is established. In this way, the audio data of audio file 1 can be used to drive and synchronize the lip movements of person A in the video frames of time periods t1 to t2 and time t3' in the video file.
[0338] As shown in Figure 5d, the double-headed dashed arrow indicates that the data on both sides of the arrow is timestamp aligned. The timestamp-aligned audio frames are used to synchronize the lip movements of the corresponding images.
[0339] Similarly, the method in this application embodiment can also establish a relationship between a set of face images B as shown in FIG5b and a set of audio segments B as shown in FIG5c.
[0340] Similarly, the method in this application embodiment can also establish a relationship between a set of face images C as shown in FIG5b and a set of audio segments C as shown in FIG5c.
[0341] Please refer to Figure 5a. The video frame at time t3' of the video file contains not only the face image of person A, but also the face image of person C. According to the lip-sync scheme in the related technology, as shown in Figure 5c, a set of audio segments C is aligned with the video segments from time t3 to t4 in the video file. According to the alignment sequence, the lip-sync scheme in the related technology cannot determine whether to use the audio frame corresponding to time t3' in the audio segment C of the audio file 2 to lip-sync the face image of person A or the face image of person C in the video frame at time t3. It is possible that the audio frame at time t3' is used to lip-sync the face image of person A in the video frame at time t3, which will cause person A to speak the lines of person C with the driven lip movements at time t3', causing lip-sync errors and confusion.
[0342] However, in the method of this application embodiment, an association is established between a set of face images C as shown in FIG5b and a set of audio segments C as shown in FIG5c, so that the lip movements of the face image of person C in the video frame at time t3' shown in FIG5a are driven by the audio of person C in the set of audio files C mapped to it. In this way, the method of this application embodiment associates audio segments and face images according to person, while there is no association relationship between audio segments and face images corresponding to different people (e.g., no association relationship). Therefore, the problem of lip movement synchronization failure and lip movement synchronization chaos in multi-person dialogue scenarios will not occur.
[0343] In one possible implementation 2, when implementing S403, the method of this application embodiment can also obtain the above-mentioned association relationship based on user input (e.g., information input by the user through the above-mentioned user interface).
[0344] As an example of implementation A, the method of this embodiment can output N sets of local images corresponding to different objects (or thumbnails of the N sets of local images, or other information indicating the N sets of local images) obtained in S402a. Thus, the user can browse the N sets of local images grouped by object through the system interface. The method of this embodiment can also output N sets of audio segments corresponding to different objects (or partial audio information of the N sets of audio segments) obtained in S402b. The user can play the N sets of audio segments grouped by object through the system interface. Finally, the user can operate the interface to instruct the association of a set of local images and a set of audio segments, so that the system of this embodiment can establish the association between the N sets of local images and the N sets of audio segments in response to the user's operation.
[0345] As another implementation example B, its implementation process is basically the same as that of example A. The difference is that the method of this application embodiment can divide the video data into N groups of video segments before lip-syncing based on the N groups of local images corresponding to different objects identified in S402a, and replace the N groups of local images output in example A with the N groups of video segments corresponding to different objects here. In this way, the user can determine which group of audio segments and which group of video segments are the audio and video of the same object by browsing and playing the N groups of video segments and N groups of audio segments grouped by object, so as to trigger the user operation. This operation can instruct to associate a group of local images and a group of audio segments, so that the system of this application embodiment can establish the association relationship between the N groups of local images corresponding to the N groups of video segments before lip-syncing and the N groups of audio segments in response to the user operation.
[0346] Alternatively, in a possible implementation 3, when implementing S403, the method of this application embodiment can adjust the above-mentioned association relationship initially established through the above implementation 1 based on user input to obtain the association relationship between N sets of audio segments and N sets of local images.
[0347] For example, the system in this application embodiment can output the initially established association relationship between N sets of audio segments and N sets of local images to the user through an interface or window, so that the user can adjust the initial association relationship through the aforementioned user interface or window to obtain the association relationship between N sets of audio segments and N sets of local images.
[0348] Referring to Figures 5a to 5d above, Figure 6 illustrates a schematic diagram of an application scenario of an embodiment of this application.
[0349] As shown in Figure 6(1), it illustrates a schematic diagram of the interface 100 of the lip-sync system 10 of this application. The system 10 of this embodiment can display the interface 100 shown in Figure 6(1) based on the association relationship initially established in the above-described implementation method 1. The interface 100 shows schematic information of the initially established association relationship.
[0350] Specifically, as shown in Figure 6(1), the interface 100 may include window 101a and window 102a. Window 101a displays a thumbnail of a face image from a set of face images A' shown in Figure 5d and control 22. Window 102a displays audio clip A from audio file 1 shown in Figure 5a, which may be a set of audio clips A as shown in Figure 5c. In other embodiments, the audio clip information displayed in window 102a may also be a portion of audio frames from a set of audio clips A. In addition, window 102a also displays control 21 for playing and pausing playback of a set of audio clips A.
[0351] The double-headed thick arrow between windows 101a and 102a is used to indicate the mapping relationship between a set of audio clips A and a set of face images A'. In other words, audio clips A are used to dub the person A shown in the face images in window 101a.
[0352] The information described for windows 101b and 101c as shown in Figure 6(1) is similar to that described for window 101a, and the information described for windows 102b and 102c is similar to that described for window 102a. Thus, the interface 100 shown in Figure 6(1) illustrates the mapping relationship between a set of face images B shown in Figure 5b and a set of audio segments B shown in Figure 5c through windows 101b and 102b indicated by double-headed thick arrows; and the interface 100 shown in Figure 6(1) illustrates the mapping relationship between a set of face images C shown in Figure 5b and a set of audio segments C shown in Figure 5c through windows 101c and 102c indicated by double-headed thick arrows.
[0353] As shown in Figure 6(1), the user can operate the control 21, and the system 10 can respond to the user's operation to play the audio clip A. For example, in the cloud scenario shown in Figure 3, the cloud management platform 201 can send the audio clip A to the client 300 so that the client 300 can play the audio clip A. In this way, the user can determine whether the audio clip A is the voice-over audio of the character A shown in window 101a. Similarly, the user can operate the control 21 in window 102b and the control 21 in window 102c to play audio clips B and C.
[0354] For example, after the user listens to the audio segment A, it is determined that the mapping relationship established by the system 10 is incorrect. The audio segment A is used to dub and synchronize the lip movements of the character C shown in window 101c, rather than to dub and synchronize the lip movements of the character A. Then, as shown in Figure 6(2), the user can operate the control 22 in window 101a, and the system 10 can respond to the user's operation and display window 11.
[0355] Window 11 may include facial images of two other persons besides the facial image of person A displayed in window 101a, as well as two option boxes. Specifically, the two facial images are thumbnails of the facial images of person B and person C. The facial image of person B can be any one of the facial images in set B shown in Figure 5b. The facial image of person C can be any one of the facial images in set C shown in Figure 5b.
[0356] As shown in Figure 6(3), when the user selects the option box 13 corresponding to the person C in window 11 (for example, from white to black), the system 10 can respond to the user's operation by adjusting the initially established association between a set of audio segments A and a set of face images A', and the association between a set of audio segments C and a set of face images C, to: the association between a set of audio segments A and a set of face images C, and the association between a set of audio segments C and a set of face images A', and switch the interface 100 from the interface 100 shown in Figure 6(3) to the interface 100 shown in Figure 6(4).
[0357] Compared to the interface 100 shown in Figure 6(1), as shown in the interface 100 in Figure 6(4), the face image in window 101a is switched from the face image of person A to the face image of person C, and the face image in window 101c is switched from the face image of person C to the face image of person A, so as to represent the mapping between the face image of audio segment A and the face image of person C, and the mapping between the face image of audio segment C and the face image of person A.
[0358] Thus, in the method of this application embodiment, after establishing the association relationship between N sets of audio segments and N sets of facial images using the AI algorithm through the above-described implementation method 1, the method of this application embodiment can display the information of the association relationship through the interface. In the event of an error in the association relationship, the method of this application embodiment can adjust the association relationship based on the user's adjustment operation to ensure that each set of audio segments is used to accurately drive the lip movements of the corresponding person's facial image.
[0359] In the embodiment shown in Figure 6 above, any operation performed by the user on the controls in the interface can be a touch operation (e.g., a swipe operation), a press operation, a voice operation, etc., and this application does not impose any restrictions on this.
[0360] Furthermore, this application does not impose any restrictions on the specific control layout of the display interface that illustrates the relationship.
[0361] Furthermore, in other embodiments, the information representing the facial images of the corresponding person in windows 101a, 101b, and 101c shown in FIG6 is not limited to a thumbnail of a single facial image of the corresponding person. It can also be multiple facial images of the person (for example, window 101a displays multiple facial images in a set of facial images A'), or it can be a set of video segments in a video file corresponding to a set of facial images of the person before lip-syncing. For example, window 101 displays a set of video segments corresponding to a set of facial images A' of person A in a video file as shown in FIG5a. For example, the video segments of time period t1 to t2 in the video file, and supports the playback of the video segments of time period t1 to t2 within window 101. When playing the video segments of time period t1 to t2, the aforementioned set of facial images A' in the video screen is marked.
[0362] In addition, in the embodiment of FIG6, the initially established association relationship can also be adjusted by operating the control 22 in the window 101c shown in FIG6(2) to be: the association relationship between a set of audio segments A and a set of face images C, and the association relationship between a set of audio segments C and a set of face images A'. The process is similar and will not be described in detail here.
[0363] In one possible implementation 4, when a user uploads an audio file to the system of this application through a client, they can also upload the binding relationship between the audio file and the person's identity document (ID).
[0364] Therefore, when implementing S403, the association between the N sets of local images and the N sets of audio segments can be determined based on the binding relationship, the N sets of local images corresponding to different objects, and the N sets of audio segments corresponding to different objects.
[0365] For example, users can pre-specify the person ID information corresponding to each audio file to be uploaded to the cloud management platform 200.
[0366] As an example of implementation, the person ID information can be a partial image of the object, such as a face image, or it can be a label of the object (such as the object label mentioned above).
[0367] The object label can be gender, age type, object type, etc.
[0368] For example, client 300 can upload audio file 1 associated with the facial image 'a' of person A to cloud management platform 200; it can also upload audio file 2 split into two audio files representing two different people, such as audio file 2.1 for person B and audio file 2.2 for person C; and it can upload audio file 2.1 associated with information indicating male gender and audio file 2.2 associated with information indicating female age type to cloud management platform 200.
[0369] In this way, the cloud management platform 200 can receive audio file 1 and the binding relationship between audio file 1 and face image a; and receive audio file 2.1 and the binding relationship between audio file 2.1 and information indicating gender as "male"; and receive audio file 2.2 and the binding relationship between audio file 2.2 and information indicating gender as "female".
[0370] In implementing S403, the face image a that is bound to audio file 1 can be matched with N sets of face images corresponding to different objects obtained based on video data to determine a set of face images that can match face image a. For example, face image a matches a set of face images A as shown in Figure 5b, indicating that a set of face images A and face image a are the same face of the same object. Since face image a is bound to audio file 1, the association between a set of face images A' and a set of audio segments A (e.g., audio file 1) is established as shown in Figure 5d.
[0371] When implementing S403, the object label (e.g., "gender is male") of the audio file 2.1 (e.g., a set of audio segments B shown in Figure 5c) can be semantically matched with the object labels identified in a set of face images B and a set of face images C shown in Figure 5b. For example, the object label of face image B is "gender is male, adult", and the object label of face image C is "gender is female, adult". In this way, the set of face images (e.g., face image B) that best matches the object label of the audio segment B can be selected from a set of face images B and a set of face images C, thus forming an association between a set of audio segments B and a set of face images B.
[0372] The process of establishing the association between a set of face images C and a set of audio segments C (e.g., the example in audio file 2.2) as shown in Figure 5c is similar and will not be repeated here.
[0373] S404, based on the association, the above N sets of audio segments and the above N sets of local images, the video data is lip-sync driven to obtain N sets of video segments of N objects.
[0374] In this application, the method can be based on a set of mutually mapped audio segments and a set of local images to drive the lip movements of a set of video segments in the video data that are located with the set of local images, so as to obtain a set of video segments after the lip movements of the object are driven.
[0375] In one possible implementation, Figure 7 exemplarily illustrates the process of obtaining a set of video clips for one of N objects from N sets of video clips. The process of obtaining the remaining video clips is similar and will not be described again here.
[0376] As shown in Figure 7, this process may include, but is not limited to, the following steps:
[0377] Optional step: S4001, perform time alignment on a set of local images and a set of audio clips that correspond to the same object in a one-to-one mapping.
[0378] Before S402a and S402b as shown in Figure 4, the obtained video data and audio data can be time-stamped first, for example, by aligning them using audio and video tracks. In this case, there is no need to execute S4001, so that the local image and audio segments in the association obtained in S403 above have already been aligned.
[0379] However, if the timestamps of the video and audio data are not aligned before S402a and S402b as shown in Figure 4, then after obtaining the association through S403, the timestamps of a set of face images and a set of audio segments in the association can be aligned to remove face images that do not need to be lip-synced.
[0380] For example, before S402a and S402b as shown in Figure 4, the timestamps of the video data and audio data were not aligned. Referring to Figure 5a, the association between audio file 1 and a set of face images A in the video file as shown in Figure 5b was established through S403. After the timestamp alignment in S404, the user determines that it is not necessary to use audio file 1 to drive the lip movement of the face image of person A at time t3' through the alignment of the audio track and video track. Therefore, the aligned association is as follows: the association between a set of face images A' and a set of audio segments A as shown in Figure 5d.
[0381] Users can set which video frames the audio data is lip-sync driven by through the user interface according to the dubbing requirements, thereby obtaining the alignment information between the audio data and the video data. Then, the method of this application embodiment can adjust the correlation between the generated N sets of local images and N sets of audio data based on the alignment information input by the user, so as to achieve time alignment between the N sets of local images and N sets of audio data.
[0382] Following S4001, time-aligned, one-to-one mapped local image sequences and audio segment sequences are obtained. Taking a set of mutually mapped audio segments A and a set of face images A' as shown in Figure 5d as an example, the local image sequence may include 50 face images of person A represented by frames P1 to P50 in the set of face images A', and the audio segment sequence may include 50 audio frames of person A represented by frames A1 to A50 in the set of audio segments A.
[0383] Following S4001, as shown in Figure 7, S4002 may also be included.
[0384] S4002, based on the audio segment sequence, perform lip-reading inference to obtain a lip-reading image sequence that is aligned with the aforementioned local image sequence.
[0385] As an implementation example, please continue to refer to Figure 5d. The method of this embodiment can perform lip-reading inference on a set of audio segments A (here described as an audio segment sequence) based on the principle of lip-reading inference as shown in Figure 1b and related embodiments to obtain a set of corresponding lip-reading images A, represented by frames m1 to m50. Here, one audio frame is used for lip-reading inference to obtain one lip-reading image. For example, lip-reading inference can be performed based on frame A1 to obtain frame m1, and so on, without further details. In other embodiments, multiple audio frames can also be used to infer a lip-reading image, which is not limited here.
[0386] For example, a trained audio feature extraction model (an AI model) can be used to extract audio features from a set of audio segments A, as shown in Figure 5d, to obtain an audio feature sequence. Then, a trained lip-sync generation model (e.g., a two-dimensional lip-sync generation model as shown in Figure 1b, or a three-dimensional lip-sync generation model not shown above) can be used to perform two-dimensional or three-dimensional lip-sync reasoning on the audio feature sequence, directly obtaining a two-dimensional lip-sync image sequence or a sequence of three-dimensional lip-sync assets (which can be rendered as two-dimensional lip-sync images). The two-dimensional lip-sync image sequence can be obtained by rendering the three-dimensional lip-sync asset sequence. In this way, a lip-sync image sequence aligned with the time period t1 to t2 of the video file, as shown in Figure 5d, can be obtained, from frames m1 to m50.
[0387] Please continue to refer to Figure 7. After S4002, S4003 may also be included.
[0388] S4003, replace the lip-sync images in the aligned local image sequence with the lip-sync image sequence obtained in S4002, to obtain the local image sequence after lip-sync synchronization.
[0389] As an implementation example, please refer to Figure 5d. The method of this application embodiment can replace the lip-shape region in each frame of a set of face images A' with the corresponding lip-shape image in a set of lip-shape images A, to obtain the face image sequence after synchronized lip-shape as shown in Figure 5e.
[0390] Specifically, referring to Figures 5d and 5e, the lip-shape region of frame P1 in a set of face images A' can be replaced with frame m1 (which is a lip-shape image obtained using audio inference) aligned with frame P1. Optionally, the image regions in frame P1 other than the lip-shape region can be fused with frame m1 (any image fusion algorithm can be used, without restriction) to obtain frame P1' as shown in Figure 5e.
[0391] Similarly, the lip-sync region in frame P2 can be replaced with frame m2 to obtain frame P2', and so on, to obtain the lip-sync synchronized face image sequence shown in Figure 5e, represented by frames P1' to P50'. Since the set of face images A' before the lip-sync region replacement is aligned with the set of audio segments A, the lip-sync synchronized face image sequence is also aligned with the set of audio segments A.
[0392] Please continue to refer to Figure 7. After S4003, S4004 may also be included.
[0393] S4004 can generate a set of video clips for an object based on an audio clip sequence and a local image sequence after synchronized lip movements.
[0394] Referring to Figure 5e, a video clip of person A can be generated using the audio clip A and the sequence of face images of person A (frames P1' to P50') that have been lip-synced based on the audio clip A. Because the audio clip A is aligned with the time interval t1 to t2 of the video file, the video clip A is also aligned with the time interval t1 to t2 of the video file.
[0395] Returning to Figure 4, after implementing S404 through the above steps S4001 to S4004, S405 can then be executed.
[0396] S405, paste the N sets of video clips back into the video data to obtain the target video with lip-sync.
[0397] Let's continue with the example of a set of video clips corresponding to a single object.
[0398] Following the process shown in Figure 7, we can obtain a set of video clips A (aligned to the time intervals t1 to t2 of the video file) after lip-syncing of person A, a set of video clips B (aligned to the time intervals t2 to t3 and t4 to t5 of the video file) after lip-syncing of person B, and a set of video clips C (aligned to the time intervals t3 to t4 of the video file) after lip-syncing of person C.
[0399] Taking video clip A as an example, Figure 5e shows the generation process of video clip A after lip-syncing. In this step S405, the method of this application embodiment can paste the video clip A after lip-syncing shown in Figure 5e back to the time period t1 to t2 in the video data (e.g., the video file shown in Figure 5a), so that the audio of character A in the video clip of time period t1 to t2 in the video file is the dubbing file of character A, specifically the audio clip A, and the lip movements of character A are also driven by the audio clip A.
[0400] The principle behind replying to video clips B and C is similar to that of replying to video clip A in a video file, so it will not be elaborated here.
[0401] The double-headed dashed arrows in Figures 5d, 5e, and 7 indicate that the two types of data pointed to by the arrows have their timestamps aligned.
[0402] Figure 8 illustrates a process diagram of a method according to an embodiment of this application. The process shown in Figure 8 can be implemented based on the cloud architecture shown in Figure 3, but is not limited to the architecture shown in Figure 3. In addition, the process shown in Figure 8 can be combined with any one of the embodiments in Figures 4, 5a to 5e, 6, and 7, but is not limited to combining any one of the above embodiments.
[0403] The implementation process in Figure 8 is largely the same as that in Figure 4. The main difference is that, in the process shown in Figure 8, the HDR video data to be lip-synced can be converted to LDR video data before lip-syncing, and then lip-syncing can be performed on the LDR video data; after lip-syncing, the LDR video data can be converted back to HDR video data. In this way, the method, apparatus and system of this application embodiment can support lip-syncing of HDR images or HDR video data using audio data in HDR space, so that the application scenario of the method of this application embodiment is not limited to lip-syncing of LDR images or videos, but can also be used for lip-syncing of HDR images or videos, solving the need for lip-syncing of a large number of HDR videos in existing scenarios.
[0404] Please refer to Figure 8. This process may include, but is not limited to, the following steps:
[0405] S500a, acquires HDR video data.
[0406] The implementation principle of S500a is the same as that of S401a shown in Figure 4. The HDR video data can be video data uploaded by the user to the system 10 or pre-existing in the system 10 for lip-syncing. Each video frame in the video data is an HDR image, not an LDR image.
[0407] S500b, acquires audio data.
[0408] The implementation principle of S500b is the same as that of S401b shown in Figure 4, and will not be repeated here.
[0409] S501a can be executed after S500a.
[0410] S501a converts HDR video data into LDR video data.
[0411] As an example of implementation, the method of this application embodiment can utilize an AI model to pre-learn the transformation relationship from HDR image to LDR image, and use this transformation relationship to convert HDR image in HDR video data into LDR image in LDR video data, thereby realizing the conversion from HDR video data to LDR video data.
[0412] Taking the CCN model as an example, the training dataset can include sample pairs of HDR and LDR images from the same captured scene. When the CNN model learns the transformation relationship from HDR to LDR, the LDR image serves as the reference image. Each sample pair from the training dataset is input into the CNN model to learn the transformation relationship, allowing it to convert the HDR image in the sample pair into a predicted LDR image. Then, the loss between the LDR image in the sample pair and the predicted LDR image output by the CNN model is calculated. Finally, this loss is used for backpropagation of the CNN model to optimize the learned transformation relationship. By training the CNN model multiple times using this training dataset, it can learn the complex transformation relationship between HDR and LDR images.
[0413] For example, this transformation relationship is the mapping between the pixel values (e.g., RGB values) of a pixel in an HDR image and the pixel values of a pixel in an LDR image that is located in the same display position as the aforementioned pixel. Thus, the image domain transformation relationship between HDR and LDR images can be simulated using the network parameters of a CNN model.
[0414] As another implementation example, the method in this application embodiment can also utilize tone mapping algorithms or other image conversion algorithms from HDR images to LDR images to convert HDR images into LDR images, thereby realizing the conversion of HDR video data to LDR video data. The tone mapping algorithm may include, but is not limited to, the Reinhard algorithm, the Durand algorithm, etc., and is not limited here.
[0415] S502a can be executed after S501a.
[0416] S502a, based on LDR video data, identifies N sets of LDR local images corresponding to different objects.
[0417] The implementation principle of S502a is the same as that of S402a shown in Figure 4, except that S502 is specifically limited to each local image in the N groups of local images being an LDR image.
[0418] S501b can be optionally executed after S500b.
[0419] S501b removes background noise from the audio data to obtain the noise-reduced audio data.
[0420] Taking the production of film and television resources as an example, the dubbing files used for dubbing and lip-syncing of film and television resources may contain a lot of noisy background noise in addition to the dubbing audio.
[0421] In lip-sync schemes in related technologies, background noise in audio data can also be used for lip-sync inference to obtain the lip movements corresponding to the background noise. These lip movements can then be used to synchronize video or image data with lip movements. This results in these background noises also driving the lip movements in the video or image, leading to lower quality output lip-synced video or image and the presence of lip-sync errors.
[0422] In this embodiment of the application, before using audio data to infer the corresponding lip movements, background noise can be removed from the audio data to obtain noise-reduced audio data.
[0423] When removing background noise from audio data, it can be done using traditional audio noise reduction methods or AI algorithms.
[0424] Taking the use of AI algorithms to remove background noise as an example, the AI algorithm can learn the audio features of background noise (such as car start-up sounds, horn sounds, noisy human voices, etc.) and the audio features of available audio without background noise. Based on the learned audio features of the two types of audio, the background noise can be identified from the audio data to be processed and then removed.
[0425] The usable audio without background noise may vary depending on the application scenario, and no restrictions are imposed here. For example, in the production of film and television resources, the usable audio could be dubbing audio collected in a professional, silent dubbing venue, such as human voice.
[0426] In this way, by using the audio data after removing background noise for lip-syncing, the accurate lip-syncing image obtained by the inference can be synchronized to the lip-syncing area in the video data, thereby improving the accuracy of lip-syncing.
[0427] S502b can be executed after S501b.
[0428] S502b, based on the noise-reduced audio data, cuts into N groups of audio segments corresponding to different objects.
[0429] The implementation principle of S502b is the same as that of S402b shown in Figure 4, and will not be repeated here.
[0430] This application does not restrict the execution order between S500a and S500b; they can be executed serially or in parallel.
[0431] S503 can be executed after S502a and S502b.
[0432] S503, based on N sets of LDR local images corresponding to different objects and N sets of audio segments corresponding to different objects, determine the association relationship between the N sets of LDR local images and the N sets of audio segments.
[0433] The implementation principle of S503 is the same as that of S403 shown in Figure 4, and will not be repeated here.
[0434] S504. Based on the association, the above N sets of audio segments, and the above N sets of LDR local images, lip-syncing is applied to the video data to obtain N sets of LDR video segments of N objects (which are N sets of video segments after lip-syncing).
[0435] An LDR video clip refers to a video clip in which the images of the video frames are LDR images.
[0436] The implementation principle of S504 is the same as that of S404 shown in Figure 4, and will not be repeated here.
[0437] Optionally, S505 may also be included after S504.
[0438] S505, perform super-resolution processing on the images in N groups of LDR video clips from N objects to obtain N groups of super-resolution LDR video clips.
[0439] In this embodiment, after using a set of audio segments to drive the lip movements in a set of LDR video segments, super-resolution processing can be performed on the image sequence in the lip-synced LDR video segments to improve the image clarity of the lip-synced LDR video segments. For example, the image sequence can be a lip-synced local image sequence (e.g., a face image sequence) in the set of LDR video segments, thereby improving the clarity of local images (e.g., faces) in the lip-synced LDR video segments.
[0440] In this embodiment, after lip-syncing is performed on N sets of LDR local images in the video data, N sets of LDR video segments of N objects with lip-synced results are obtained. Then, super-resolution processing is performed on the images in the N sets of LDR video segments to achieve super-resolution of the complete images in the LDR video segments (e.g., including the lip-synced images) to improve the resolution of the images in the lip-synced video.
[0441] Referring to the embodiment in Figure 7, super-resolution processing can be performed on the local image sequence obtained after lip-syncing in S4003, after S4003 and before S4004. Alternatively, super-resolution processing of the images in the video segments can be performed on a group of video segments generated after lip-syncing for each object, without limitation.
[0442] In other embodiments, super-resolution processing can also be performed on the N sets of LDR local images obtained in S502a before the lip-sync driving in S504, so as to realize super-resolution processing of the local images before lip-sync synchronization.
[0443] S506 re-attaches the N sets of LDR video clips after super-resolution to the video data to obtain the LDR video data after lip-syncing.
[0444] The implementation principle of S506 is the same as that of S405 in Figure 4, and will not be repeated here.
[0445] S506 may also be included after S507.
[0446] S507 performs an inverse conversion on the lip-synced LDR video data (the image is converted from LDR space to HDR space) to obtain the lip-synced HDR video data.
[0447] The "inverse conversion" operation in this step is the reverse of the "conversion" operation in S501a above. Through this "inverse conversion" operation, the LDR image in the LDR video data can be converted into an HDR image to obtain HDR video data.
[0448] As an example of implementation, the method of this application embodiment can utilize an AI model to pre-learn the inverse transformation relationship from LDR image to HDR image, and use this inverse transformation relationship to convert the LDR image in LDR video data into the HDR image in HDR video data, thereby realizing the inverse transformation from LDR video data to HDR video data.
[0449] When using an AI model to learn the inverse transformation relationship, the implementation principle is similar to that of using an AI model to learn the transformation relationship. The difference is that the reference data in each sample pair is an HDR image, and the LDR image that together with the HDR image constitutes the sample pair is used as input to the AI model so that the AI model can learn the inverse transformation relationship.
[0450] In other implementation examples, the method of this application embodiment can also utilize tone mapping algorithms, etc., to achieve the inverse conversion from LDR image to HDR image, which is not limited here.
[0451] In this embodiment of the application, when implementing lip-syncing of audio data with HDR image or HDR video data in an HDR space scenario, the lip movements obtained by lip-syncing inference using audio data are still LDR lip movements. This ensures that the AI model used for lip-syncing inference (e.g., the two-dimensional lip-syncing generation model shown in Figure 1b) remains unchanged, and the AI model is still used for lip-syncing in the LDR space. Instead, before lip-syncing the video data, the HDR video data is converted to LDR video data, and audio data is used to perform lip-syncing on the LDR video data in the LDR space. Finally, the lip-synced LDR video data is inversely converted back to HDR video data, thereby supporting lip-syncing of video or image in the HDR space.
[0452] In another possible implementation, to enable the method of this application embodiment to support lip-syncing of video or images in HDR space, unlike the embodiment in FIG8, this application embodiment may not execute S501a and S507, but directly perform the recognition of local images (in HDR space) of different objects of video data in HDR space, and the establishment of the association between audio segments and local images. In the lip-syncing driving process of S504, referring to FIG7, the lip-syncing image sequence obtained by inferring each group of audio segments in N groups of audio segments through S4002 is still an LDR lip-syncing image sequence. However, in this embodiment, unlike the embodiment in FIG8, this embodiment can convert the LDR lip-syncing image sequence into an HDR lip-syncing image sequence, that is, convert the lip-syncing image from LDR space to HDR space, and then execute the subsequent S4003 and S4004 processes. Thus, returning to the process shown in FIG8, in this application embodiment, what is obtained by S504 is N groups of HDR video segments of N objects, and what is obtained after posting is also HDR video data.
[0453] In another possible implementation, in order to enable the method of this application embodiment to support lip-syncing of video or images in HDR space, unlike the embodiment in FIG8, this application embodiment may not convert the HDR video data to be synchronized to LDR space. Instead, it may use an AI model to infer lip-syncing actions in HDR space from the audio data to obtain lip-syncing images in HDR space. Then, the lip-syncing images in HDR space obtained by inference are used to paste back onto the corresponding lip-syncing regions in the HDR video data.
[0454] Specifically, the implementation process of this method is largely the same as that in Figure 8. The differences are described below:
[0455] Difference 1: S501a and S507 do not need to be executed, so there is no need to perform HDR space to LDR space conversion and inverse conversion on HDR video, which can reduce the complexity of the lip-sync process and reduce latency.
[0456] Difference 2: In S504, based on the correlation between N sets of HDR local images and N sets of audio segments, the N sets of audio segments corresponding to different objects are used to drive the lip movements of the N sets of HDR local images corresponding to different objects, resulting in N sets of video segments for N objects.
[0457] Regarding this difference 2, please refer to Figure 7. As shown in Figure 7, the time-aligned local image sequence is an HDR local image sequence. In S4002, the lip-sync image sequence obtained by performing lip-sync reasoning on the audio segment sequence is an HDR lip-sync image sequence, not the LDR lip-sync image sequence corresponding to the embodiment in Figure 8. Thus, after S4003 and S4004, the resulting set of video segments for an object is a set of HDR video segments for that object, so there is no need to perform the inverse conversion operation shown in S507 of Figure 8.
[0458] The implementation process of S4002 shown in Figure 7 is described below:
[0459] In S4002, the AI model used for lip-reading reasoning in this application can be a two-dimensional lip-reading generation model for reasoning audio features to obtain a two-dimensional HDR lip-reading image.
[0460] The training process of this AI model follows the principle of the training process shown in Figure 1a. However, the difference lies in the fact that the reference lip-sync image obtained in S102 corresponding to the audio features is not a reference lip-sync image in the LDR space of related technologies, but rather a reference lip-sync image in the HDR space (e.g., an HDR image of the lip shape described by the audio features). Therefore, the predicted lip-sync image output by the AI model in S103 shown in Figure 1a is also a predicted lip-sync image in the HDR space, rather than a predicted lip-sync image in the LDR space of related technologies. Thus, the AI model trained according to this process can generate corresponding lip-sync images in the HDR space based on the audio features.
[0461] The reference lip-sync images in the HDR space used in the training data for training the AI model can be HDR lip-sync images acquired using specialized equipment capable of capturing images in the HDR space.
[0462] After training the two-dimensional lip-sync generation model for generating HDR lip-sync images, when implementing S4002 as shown in FIG7 in this embodiment, the process shown in FIG1b can be referred to. However, the lip-sync images generated by the two-dimensional lip-sync generation model used in FIG1b are HDR lip-sync images, not LDR lip-sync images.
[0463] Figure 9 is a schematic diagram of an exemplary lip-sync device 800. Referring to Figure 9, the lip-sync device 800 includes: an acquisition module 801, used to acquire audio data and media data, wherein the audio data includes a first audio segment corresponding to a first object and a second audio segment corresponding to a second object, and the media data includes a first target image corresponding to the first object and a second target image corresponding to the second object, wherein the first target image is an image including the mouth region of the first object, and the second target image is an image including the mouth region of the second object; and an association module 802, used to associate the first audio segment corresponding to the first object and the first target image based on the audio data and the media data, and to associate the second audio segment corresponding to the second object. The first audio segment is associated with the second target image to obtain a first association relationship between the first audio segment and the first target image, and a second association relationship between the second audio segment and the second target image; the lip-sync module 803 is used to perform lip-sync on the first target image based on the first audio segment according to the first association relationship, and to perform lip-sync on the second target image based on the second audio segment according to the second association relationship, to obtain target media data, wherein the first target image of the first object in the target media data includes a lip-sync image formed by the mouth region of the first object based on the first audio segment, and the second target image of the second object in the target media data includes a lip-sync image formed by the mouth region of the second object based on the second audio segment.
[0464] In one possible implementation, the association module 802 is specifically configured to: obtain first feature information of the first object corresponding to the first audio segment in the audio data; obtain first feature information of the second object corresponding to the second audio segment in the audio data; obtain second feature information of the first object corresponding to the first target image in the media data; obtain second feature information of the second object corresponding to the second target image in the media data; associate the first audio segment and the first target image whose first feature information and second feature information match to obtain a first association relationship between the first audio segment and the first target image; associate the second audio segment and the second target image whose first feature information and second feature information match to obtain a first association relationship between the second audio segment and the second target image.
[0465] In one possible implementation, both the first feature information and the second feature information include at least one of the following: facial feature information, gender feature information, age feature information, and object type feature information.
[0466] In one possible implementation, the association module 802 is specifically used to: in response to a received first user operation, obtain a first tag of the first object corresponding to the first audio segment, wherein the first user operation indicates the first tag, and the first tag of the first object indicates the feature information of the first object corresponding to the first audio segment.
[0467] In one possible implementation, the association module 802 is specifically used to: use an artificial intelligence (AI) algorithm to identify the audio features of the second audio segment; and based on the audio features, determine a first tag of the second object corresponding to the second audio segment, wherein the first tag of the second object indicates the feature information of the second object corresponding to the second audio segment.
[0468] In one possible implementation, the association module 802 is specifically used to: in response to a received second user operation, obtain a second tag of the first object corresponding to the first target image, wherein the second user operation indicates the second tag, and the second tag of the first object indicates feature information of the first object corresponding to the first target image.
[0469] In one possible implementation, the association module 802 is specifically used to: use an AI algorithm to identify image features of the second target image; and based on the image features, determine a second label of the second object corresponding to the image features, wherein the second label of the second object indicates feature information of the second object corresponding to the second target image.
[0470] In one possible implementation, the apparatus further includes: an output module 804, configured to output information indicating the first association relationship and the second association relationship; the association module 802 is further configured to adjust the first association relationship and the second association relationship in response to a received third user operation, to obtain an adjusted first association relationship and an adjusted second association relationship, wherein the third user operation instructs to adjust a first audio segment in the first association relationship to a second audio segment in the second association relationship, or instructs to adjust a first target image in the first association relationship to a second target image in the second association relationship.
[0471] In one possible implementation, the first target image is a high dynamic range (HDR) image, and the lip-sync module 803 is specifically used to: obtain a first low dynamic range (LDR) lip-sync image based on the first audio segment, wherein the first LDR lip-sync image is a lip-sync image of the first object driven by the first audio segment; convert the first LDR lip-sync image into a first HDR lip-sync image; and, according to the first association relationship, perform lip-sync on the lip-sync image of the first object in the first target image that is associated with the first audio segment based on the first HDR lip-sync image to obtain target media data.
[0472] In one possible implementation, the media data is HDR media data, and the association module 802 is specifically configured to: convert the brightness range of the HDR media data from HDR to LDR to obtain LDR media data, wherein the LDR media data includes a first LDR target image corresponding to the first object and a second LDR target image corresponding to the second object; based on the audio data and the LDR media data, associate the first audio segment corresponding to the first object with the first LDR target image, and associate the second audio segment corresponding to the second object with the second LDR target image to obtain a first association relationship between the first audio segment and the first LDR target image, and a second association relationship between the second audio segment and the second LDR target image.
[0473] In one possible implementation, the lip-sync module 803 is specifically configured to: perform lip-sync on the first LDR target image based on the first audio segment according to the first association relationship, and perform lip-sync on the second LDR target image based on the second audio segment according to the second association relationship to obtain target LDR media data; and convert the brightness range of the target LDR media data from LDR to HDR to obtain target HDR media data.
[0474] In one possible implementation, the media data is HDR media data, the first target image in the HDR media data is a first HDR target image, and the second target image in the HDR media data is a second HDR target image; the lip-sync module 803 is specifically used for: performing lip-sync inference on the first audio segment using an AI model, and performing lip-sync inference on the second audio segment to obtain a first HDR lip-sync image corresponding to the first audio segment and a second HDR lip-sync image corresponding to the second audio segment; according to the first association relationship, based on the first HDR lip-sync image, performing lip-sync on the lip-sync image of the first object in the first HDR target image that is associated with the first audio segment, and according to the second association relationship, based on the second HDR lip-sync image, performing lip-sync on the lip-sync image of the first object in the second HDR target image that is associated with the second audio segment, to obtain target HDR media data.
[0475] In one possible implementation, the acquisition module 801 is specifically used to perform audio segmentation on the audio data for different objects, to obtain a first audio segment corresponding to a first object and a second audio segment corresponding to a second object.
[0476] In one possible implementation, the acquisition module 801 is specifically used to identify different objects in the media data and obtain a first target image corresponding to the first object and a second target image corresponding to the second object.
[0477] All of the above modules can be implemented in software or hardware. As an example of a software functional unit, a module may include code running on a computing instance. A computing instance may include at least one of a physical host (computing device), a virtual machine, or a container. Furthermore, there may be one or more computing instances. For example, associated module 802 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Further, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one or more geographically proximate data centers. Typically, a region may include multiple AZs.
[0478] Similarly, multiple hosts / virtual machines / containers used to run this code can be distributed within the same Virtual Private Cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Communication between two VPCs within the same region, as well as between VPCs in different regions, requires a communication gateway to be set up within each VPC to enable interconnection between VPCs.
[0479] As an example of a hardware functional unit, the module mentioned above may include at least one computing device, such as a server. Alternatively, the module may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be implemented using a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), generic array logic (GAL), or any combination thereof.
[0480] The aforementioned lip-syncing device includes multiple computing devices that can be distributed within the same region or in different regions. Similarly, the multiple computing devices can be distributed within the same Availability Zone (AZ) or in different AZs. Likewise, the multiple computing devices can be distributed within the same Virtual Private Cloud (VPC) or in multiple VPCs. These multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.
[0481] It should be noted that, in other embodiments, the above-described module can be used to perform the corresponding steps in the above-described method to realize all the functions of the lip-sync device.
[0482] This application also provides a computing device 900. As shown in FIG10, the computing device 900 includes: a bus 902, a processor 904, a memory 906, and a communication interface 909. The processor 904, the memory 906, and the communication interface 909 communicate with each other via the bus 902. The computing device 900 may be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computing device 900.
[0483] Bus 902 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, only one line is used in Figure 10, but this does not imply that there is only one bus or one type of bus. Bus 902 can include pathways for transmitting information between various components of computing device 900 (e.g., memory 906, processor 904, communication interface 909).
[0484] Processor 904 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0485] Memory 906 may include volatile memory, such as random access memory (RAM). Memory 906 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).
[0486] The memory 906 stores executable program code, which the processor 904 executes to implement the functions of the aforementioned acquisition module 801, association module 802, lip-sync module 803, and output module 804, thereby realizing the lip-sync method. In other words, the memory 906 stores instructions for executing the lip-sync method.
[0487] The communication interface 909 uses transceiver modules, such as, but not limited to, network interface cards and transceivers, to enable communication between the computing device 900 and other devices or communication networks.
[0488] This application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.
[0489] As shown in Figure 11, the computing device cluster includes at least one computing device 1000. The memory 1006 of one or more computing devices 1000 in the computing device cluster may store the same instructions for executing the lip-sync method.
[0490] The computing device 1000 includes a bus 1002, a processor 1004, a memory 1006, and a communication interface 1008. The processor 1004, the memory 1006, and the communication interface 1008 communicate with each other via the bus 1002.
[0491] In some possible implementations, the memory 1006 of one or more computing devices 1000 in the computing device cluster may also store partial instructions for executing the lip-sync method. In other words, a combination of one or more computing devices 1000 can jointly execute the instructions for executing the lip-sync method.
[0492] It should be noted that the memory 1006 in different computing devices 1000 within the computing device cluster can store different instructions, each used to execute a portion of the functions of the lip-syncing device. That is, the instructions stored in the memory 1006 of different computing devices 1000 can implement the functions of one or more of the aforementioned acquisition module 801, association module 802, lip-syncing module 803, and output module 804.
[0493] In some possible implementations, one or more computing devices in a computing device cluster can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc. Figure 12 illustrates one possible implementation. As shown in Figure 12, two computing devices 1100A and 1100B are connected via a network. Computing device 1100A includes a bus 1102, a processor 1104, a memory 1106, and a communication interface 1108. The processor 1104, memory 1106, and communication interface 1108 communicate via the bus 1102. Computing device 1100B includes a bus 1102, a processor 1104, a memory 1106, and a communication interface 1108. The processor 1104, memory 1106, and communication interface 1108 communicate via the bus 1102. Specifically, the communication interface 1108 in each computing device is used to connect to the network. In this type of possible implementation, the memory 1106 in computing device 1100A stores instructions for executing the functions of the acquisition module 801, the association module 802, and the lip-sync module 803. Meanwhile, the memory 1106 in computing device 1100B stores instructions for executing the functions of the output module 804.
[0494] It should be understood that the functions of computing device 1100A shown in Figure 12 can also be performed by multiple computing devices 1100. Similarly, the functions of computing device 1100B can also be performed by multiple computing devices 1100.
[0495] This application also provides another computing device cluster. The connection relationship between the computing devices in this computing device cluster can be similarly referred to the connection method of the computing device clusters shown in Figures 10 and 12. The difference is that the memory 1106 of one or more computing devices 1100 in this computing device cluster can store the same instructions for executing the lip-sync method.
[0496] In some possible implementations, the memory 1106 of one or more computing devices 1100 in the computing device cluster may also store partial instructions for executing the lip-sync method. In other words, a combination of one or more computing devices 1100 can jointly execute the instructions for executing the lip-sync method.
[0497] It should be noted that the memory 1106 in different computing devices 1100 within the computing device cluster can store different instructions for executing some functions of the lip-syncing device. That is, the instructions stored in the memory 1106 of different computing devices 1100 can implement the functions of one or more devices in the lip-syncing device.
[0498] This application also provides a computer program product containing instructions. The computer program product may be a software or program product containing instructions, capable of running on a computing device or stored on any usable medium. When the computer program product is run on at least one computing device, it causes the at least one computing device to perform the lip-syncing method described in the above embodiments.
[0499] This application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store, or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to perform the lip-sync method described in the above embodiments.
[0500] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of this application.
Claims
1. A method of mouth synchronisation, characterized in that, The method comprises: acquiring audio data and media data, the audio data comprising a first audio segment corresponding to a first object and a second audio segment corresponding to a second object, and the media data comprising a first target image corresponding to the first object and a second target image corresponding to the second object, the first target image being an image comprising a mouth region of the first object, and the second target image being an image comprising a mouth region of the second object; based on the audio data and the media data, associating the first audio segment corresponding to the first object with the first target image, and associating the second audio segment corresponding to the second object with the second target image, to obtain a first association relationship between the first audio segment and the first target image, and a second association relationship between the second audio segment and the second target image; based on the first audio segment, performing lip synchronization on the first target image according to the first association relationship, and based on the second audio segment, performing lip synchronization on the second target image according to the second association relationship, to obtain target media data, wherein the first target image of the first object in the target media data comprises a lip image of the mouth region of the first object driven based on the first audio segment, and the second target image of the second object in the target media data comprises a lip image of the mouth region of the second object driven based on the second audio segment.
2. The method of claim 1, wherein, The association of the first audio segment corresponding to the first object with the first target image, and the association of the second audio segment corresponding to the second object with the second target image, based on the audio data and the media data, to obtain a first association relationship between the first audio segment and the first target image, and a second association relationship between the second audio segment and the second target image, comprises: based on the first audio segment in the audio data, acquiring first feature information of the first object corresponding to the first audio segment, and based on the second audio segment in the audio data, acquiring first feature information of the second object corresponding to the second audio segment; based on the first target image in the media data, acquiring second feature information of the first object corresponding to the first target image, and based on the second target image in the media data, acquiring second feature information of the second object corresponding to the second target image; associating the first audio segment and the first target image, whose first feature information and second feature information match each other, to obtain a first association relationship between the first audio segment and the first target image; associating the second audio segment and the second target image, whose first feature information and second feature information match each other, to obtain a first association relationship between the second audio segment and the second target image.
3. The method of claim 2, wherein, The first feature information and the second feature information each include at least one of the following: face feature information, gender feature information, age feature information, and object type feature information.
4. The method according to claim 2 or 3, characterized in that, The first feature information of the first object corresponding to the first audio segment is obtained based on the first audio segment in the audio data, including: In response to a received first user operation, a first label of the first object corresponding to the first audio segment is obtained, wherein the first user operation indicates the first label, and the first label of the first object indicates first feature information of the first object corresponding to the first audio segment.
5. The method according to any one of claims 2 to 4, characterized in that, The first feature information of the second object corresponding to the second audio segment is obtained based on the second audio segment in the audio data, including: An audio feature of the second audio segment is identified by using an artificial intelligence (AI) algorithm. Based on the audio feature, a first label of the second object corresponding to the second audio segment is determined, wherein the first label of the second object indicates first feature information of the second object corresponding to the second audio segment.
6. The method according to any one of claims 2 to 5, characterized in that, The second feature information of the first object corresponding to the first target image is obtained based on the first target image in the media data, including: In response to a received second user operation, a second label of the first object corresponding to the first target image is obtained, wherein the second user operation indicates the second label, and the second label of the first object indicates second feature information of the first object corresponding to the first target image.
7. The method according to any one of claims 2 to 6, characterized in that, The second feature information of the second object corresponding to the second target image is obtained based on the second target image in the media data, including: An image feature of the second target image is identified by using an AI algorithm. Based on the image feature, a second label of the second object corresponding to the image feature is determined, wherein the second label of the second object indicates second feature information of the second object corresponding to the second target image.
8. The method according to any one of claims 1 to 7, characterized in that, Before the target media data is obtained by performing mouth shape synchronization on the first target image based on the first audio segment according to the first association relationship and performing mouth shape synchronization on the second target image based on the second audio segment according to the second association relationship, the method further includes: Outputting information indicating the first association relationship and the second association relationship; In response to a received third user operation, the first association relationship and the second association relationship are adjusted to obtain an adjusted first association relationship and an adjusted second association relationship, wherein the third user operation indicates that a first audio segment in the first association relationship is adjusted to a second audio segment in the second association relationship, or indicates that a first target image in the first association relationship is adjusted to a second target image in the second association relationship.
9. The method according to any one of claims 1 to 8, characterized in that, The first target image is a high dynamic range (HDR) image, and the mouth shape synchronization on the first target image based on the first audio segment according to the first association relationship includes: obtaining a first low dynamic range (LDR) lip image based on the first audio clip, the first LDR lip image being a lip image of the first object formed based on the first audio clip; converting the first LDR lip image into a first high dynamic range (HDR) lip image; performing lip synchronization on the first target image based on the first audio clip according to the first association relationship, to obtain target media data.
10. The method according to any one of claims 1 to 7, characterized in that, The media data is HDR media data, and the association between the first audio clip corresponding to the first object and the first target image and the association between the second audio clip corresponding to the second object and the second target image are obtained based on the audio data and the media data, to obtain a first association relationship between the first audio clip and the first target image and a second association relationship between the second audio clip and the second target image, which includes: converting the brightness range of the HDR media data from HDR to LDR to obtain LDR media data, the LDR media data including a first LDR target image corresponding to the first object and a second LDR target image corresponding to the second object; associating the first audio clip corresponding to the first object with the first LDR target image and associating the second audio clip corresponding to the second object with the second LDR target image based on the audio data and the LDR media data, to obtain a first association relationship between the first audio clip and the first LDR target image and a second association relationship between the second audio clip and the second LDR target image.
11. The method of claim 10, wherein, The lip synchronization of the first target image based on the first audio clip according to the first association relationship and the lip synchronization of the second target image based on the second audio clip according to the second association relationship to obtain target media data includes: performing lip synchronization on the first LDR target image based on the first audio clip according to the first association relationship and performing lip synchronization on the second LDR target image based on the second audio clip according to the second association relationship to obtain target LDR media data; converting the brightness range of the target LDR media data from LDR to HDR to obtain target HDR media data.
12. The method according to any one of claims 1 to 7, characterized in that, The media data is HDR media data, the first target image in the HDR media data is a first HDR target image, and the second target image in the HDR media data is a second HDR target image; and the lip synchronization of the first target image based on the first audio clip according to the first association relationship and the lip synchronization of the second target image based on the second audio clip according to the second association relationship to obtain target media data includes: performing lip inference on the first audio segment and the second audio segment by using an AI model to obtain a first HDR lip image corresponding to the first audio segment and a second HDR lip image corresponding to the second audio segment; performing lip synchronization on a lip image of the first object in the first HDR target image that has a correlation relationship with the first audio segment based on the first HDR lip image according to the first correlation relationship, and performing lip synchronization on a lip image of the first object in the second HDR target image that has a correlation relationship with the second audio segment based on the second HDR lip image according to the second correlation relationship, to obtain target HDR media data.
13. The method according to any one of claims 1 to 12, characterized in that, The method further comprises: performing audio segmentation on the audio data to obtain a first audio segment corresponding to a first object and a second audio segment corresponding to a second object.
14. The method according to any one of claims 1 to 13, characterized in that, The method further comprises: performing object recognition on the media data to obtain a first target image corresponding to the first object and a second target image corresponding to the second object.
15. A lip sync apparatus, characterized by: The apparatus comprises: an acquisition module configured to acquire audio data and media data, the audio data comprising a first audio segment corresponding to a first object and a second audio segment corresponding to a second object, and the media data comprising a first target image corresponding to the first object and a second target image corresponding to the second object, the first target image being an image comprising a mouth region of the first object, and the second target image being an image comprising a mouth region of the second object; a correlation module configured to correlate the first audio segment corresponding to the first object and the first target image based on the audio data and the media data, and correlate the second audio segment corresponding to the second object and the second target image, to obtain a first correlation relationship between the first audio segment and the first target image and a second correlation relationship between the second audio segment and the second target image; a lip synchronization module configured to perform lip synchronization on the first target image based on the first audio segment according to the first correlation relationship, and perform lip synchronization on the second target image based on the second audio segment according to the second correlation relationship, to obtain target media data, wherein a first target image of the first object in the target media data comprises a lip image of a mouth region of the first object driven based on the first audio segment, and a second target image of the second object in the target media data comprises a lip image of a mouth region of the second object driven based on the second audio segment.
16. The apparatus of claim 15, wherein, The correlation module is specifically configured to: acquire first feature information of the first object corresponding to the first audio segment based on the first audio segment in the audio data, and acquire first feature information of the second object corresponding to the second audio segment based on the second audio segment in the audio data. obtaining second feature information of the first object corresponding to the first target image based on the first target image in the media data, and obtaining second feature information of the second object corresponding to the second target image based on the second target image in the media data; associating the first audio segment and the first target image with each other according to the first feature information and the second feature information, to obtain a first association relationship between the first audio segment and the first target image; associating the second audio segment and the second target image with each other according to the first feature information and the second feature information, to obtain a first association relationship between the second audio segment and the second target image.
17. The apparatus of claim 16, wherein, The first feature information and the second feature information each include at least one of the following: face feature information, gender feature information, age feature information, and object type feature information.
18. The apparatus of claim 16 or 17, wherein, The association module is specifically configured to: in response to a received first user operation, obtaining a first label of the first object corresponding to the first audio segment, wherein the first user operation indicates the first label, and the first label of the first object indicates feature information of the first object corresponding to the first audio segment.
19. The apparatus of any one of claims 16-18, wherein, The association module is specifically configured to: using an artificial intelligence (AI) algorithm, identifying audio features of the second audio segment; based on the audio features, determining a first label of the second object corresponding to the second audio segment, wherein the first label of the second object indicates feature information of the second object corresponding to the second audio segment.
20. The apparatus of any one of claims 16-19, wherein, The association module is specifically configured to: in response to a received second user operation, obtaining a second label of the first object corresponding to the first target image, wherein the second user operation indicates the second label, and the second label of the first object indicates feature information of the first object corresponding to the first target image.
21. The apparatus of any one of claims 16-20, wherein, The association module is specifically configured to: using an AI algorithm, identifying image features of the second target image; based on the image features, determining a second label of the second object corresponding to the image features, wherein the second label of the second object indicates feature information of the second object corresponding to the second target image.
22. The apparatus of any one of claims 15 to 21, wherein, The device further includes: an output module configured to output information indicating the first association relationship and the second association relationship; The association module is further configured to, in response to a received third user operation, adjust the first association relationship and the second association relationship to obtain an adjusted first association relationship and an adjusted second association relationship, wherein the third user operation indicates adjusting a first audio segment in the first association relationship to a second audio segment in the second association relationship, or adjusting a first target image in the first association relationship to a second target image in the second association relationship.
23. The apparatus of any of claims 15 to 22, wherein, The first target image is a high dynamic range (HDR) image, and the lip synchronization module is specifically configured to: based on the first audio segment, obtaining a first low dynamic range (LDR) lip image, wherein the first LDR lip image is a lip image of the first object formed based on the first audio segment. convert the first LDR lip shape image into a first HDR lip shape image; synchronize the lip shape of the first object in the first target image associated with the first audio segment based on the first HDR lip shape image according to the first association relationship, to obtain target media data.
24. The apparatus of any one of claims 15 to 21, wherein, The media data is HDR media data, and the association module is specifically configured to: convert the brightness range of the HDR media data from HDR to LDR to obtain LDR media data, wherein the LDR media data includes a first LDR target image corresponding to the first object and a second LDR target image corresponding to the second object; associate the first audio segment corresponding to the first object with the first LDR target image and associate the second audio segment corresponding to the second object with the second LDR target image based on the audio data and the LDR media data, to obtain a first association relationship between the first audio segment and the first LDR target image and a second association relationship between the second audio segment and the second LDR target image.
25. The apparatus of claim 24, wherein, The lip shape synchronization module is specifically configured to: synchronize the lip shape of the first LDR target image based on the first audio segment according to the first association relationship, and synchronize the lip shape of the second LDR target image based on the second audio segment according to the second association relationship, to obtain target LDR media data; convert the brightness range of the target LDR media data from LDR to HDR to obtain target HDR media data.
26. The apparatus of any one of claims 15 to 21, wherein, The media data is HDR media data, the first target image in the HDR media data is a first HDR target image, and the second target image in the HDR media data is a second HDR target image; and the lip shape synchronization module is specifically configured to: perform lip reasoning on the first audio segment and the second audio segment by using an AI model, to obtain a first HDR lip shape image corresponding to the first audio segment and a second HDR lip shape image corresponding to the second audio segment; synchronize the lip shape of the first object in the first HDR target image associated with the first audio segment based on the first HDR lip shape image according to the first association relationship, and synchronize the lip shape of the first object in the second HDR target image associated with the second audio segment based on the second HDR lip shape image according to the second association relationship, to obtain target HDR media data.
27. The apparatus of any of claims 15 to 26, wherein, The acquisition module is specifically configured to perform audio segmentation of different objects on the audio data, to obtain a first audio segment corresponding to a first object and a second audio segment corresponding to a second object.
28. The apparatus of any of claims 15 to 27, wherein, The acquisition module is specifically configured to perform object recognition on the media data, to obtain a first target image corresponding to the first object and a second target image corresponding to the second object.
29. A cluster of computing devices, characterized in that, comprising at least one computing device, each computing device comprising a processor and a memory; the processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the cluster of computing devices to perform the method of any one of claims 1 to 14.
30. A computer program product comprising instructions, wherein: the instructions, when executed by the cluster of computing devices, cause the cluster of computing devices to perform the method of any one of claims 1 to 14.
31. A computer readable storage medium, characterized in that, computer program instructions, which, when executed by a cluster of computing devices, cause the cluster of computing devices to perform the method of any one of claims 1 to 14.
Citation Information
Patent Citations
Video speaker identification method and device, computer equipment and storage medium
CN111785279A
Video stitching method and device, equipment and medium
CN115766973A
Image processing method and related equipment thereof
CN116309226A
Mouth shape driving method and device, mouth shape driving model training method and device, equipment and medium
CN117079664A
Sound and picture synchronization method, sound and picture synchronization device, electronic equipment and storage medium
CN117793430A