Virtual image video generation method and device, computer device, and storage medium

By extracting and processing facial videos, background videos and voice information, lip movement videos and target voices are generated, which solves the problem of separation between lip movements and virtual voices in virtual image videos and achieves high-quality restoration of the target object's lip movements.

CN115147516BActive Publication Date: 2025-10-17CHINA PING AN LIFE INSURANCE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210744674.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-28
Publication Date
2025-10-17
Estimated Expiration
2042-06-28

AI Technical Summary

Technical Problem

In the prior art, when generating a virtual image video, the lip movements are separated from the virtual voice, resulting in poor restoration of the facial movements in the virtual image video.

Method used

By extracting the target object's facial video, background video and voice information, the lip feature information is determined, and the lip movement video and target voice are generated. The facial video is combined to generate a facial fusion video, and finally a virtual image video is synthesized to make the lip movement match the virtual voice.

Benefits of technology

The accuracy of the target object's lip movements in the virtual image video is improved, ensuring the synchronization of lip movements and virtual voice, and enhancing the authenticity of the virtual image video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115147516B_ABST
    Figure CN115147516B_ABST
Patent Text Reader

Abstract

The application relates to the field of artificial intelligence and discloses a virtual image video generation method, device, equipment and medium. The method comprises the following steps: obtaining a to-be-processed video, and extracting a pre-processed video based on the facial features of a target object in the to-be-processed video; extracting a facial video, a background video and voice information of the target object in the pre-processed video; determining the lip feature information of the target object based on the facial video, and generating a lip action video and target voice corresponding to the lip action video according to the voice information and the lip feature information; generating a facial fusion video of the target object according to the lip action video and the facial video; and synthesizing a virtual image video according to the facial fusion video, the background video and the target voice, so that the lip action of the target object in the virtual image video is consistent with the virtual voice, and the obtained virtual image video can restore the actual action of the lips of the target object.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of big data, and in particular to a virtual image video generation method and device, computer equipment and a storage medium. BACKGROUND

[0002] At present, people can generate personalized virtual images through their own photos or videos, and generate corresponding virtual image videos based on the virtual images, and when sharing and communicating on social platforms, people can use their personalized virtual image videos to replace traditional real videos, which can effectively protect their privacy.

[0003] However, when generating virtual image videos directly through photos or videos, the generation effect of the head picture and the action of the virtual image cannot meet the actual use requirements, especially the problem of the split between the lip action of the target object in the virtual image video and the virtual voice, which will result in a relatively poor restoration of the actual facial action of the target object in the obtained virtual image video. SUMMARY

[0004] The embodiments of the present application provide a virtual image video generation method, device, computer equipment and medium, aiming to make the lip action of the target object in the virtual image video consistent with the virtual voice, and make the obtained virtual image video better restore the actual action of the target object's lips.

[0005] In a first aspect, the embodiments of the present application provide a virtual image video generation method, comprising:

[0006] Obtaining a to-be-processed video, and extracting a pre-processed video based on the facial features of a target object in the to-be-processed video;

[0007] Extracting a facial video, a background video and voice information of the target object in the pre-processed video;

[0008] Determining the lip feature information of the target object based on the facial video, and generating a lip action video and a target voice corresponding to the lip action video according to the voice information and the lip feature information;

[0009] Generating a facial fusion video of the target object according to the lip action video and the facial video;

[0010] Synthesizing a virtual image video according to the facial fusion video, the background video and the target voice.

[0011] In some embodiments, extracting the facial video and the background video of the target object in the pre-processed video comprises:

[0012] Extracting a plurality of single-frame images in the pre-processed video;

[0013] performing semantic segmentation on the single-frame image to obtain at least one face connected domain corresponding to a face of the target object;

[0014] generating a dynamic face mask model of the target object according to the face connected domain in the plurality of single-frame images;

[0015] extracting a face video and a background video in the preprocessed video based on the dynamic face mask model.

[0016] In some embodiments, the lip feature information of the target object is determined based on the face video, including:

[0017] extracting a plurality of face image frames in the face video;

[0018] performing difference processing on the plurality of face image frames to obtain a lip connected domain of the lips of the target object in the face image frames;

[0019] extracting a lip image in the face image frames according to the lip connected domain;

[0020] inputting the lip image into a pre-set feature extraction model to extract the lip feature information of the target object, wherein the lip feature information includes at least one of lip shape feature information, lip color feature information, and lip movement feature information.

[0021] In some embodiments, the lip connected domain of the lips of the target object in the face image frames is obtained by performing difference processing on the plurality of face image frames, including:

[0022] performing difference processing on the time-adjacent face image frames to obtain a difference image frame;

[0023] dividing the face image frames into a plurality of sub-regions;

[0024] determining an average difference value of the plurality of sub-regions in the face image frames according to the difference image frame;

[0025] merging the position-adjacent sub-regions according to the average difference value to obtain at least one pending connected domain, and determining the lip connected domain in the pending connected domain.

[0026] In some embodiments, the lip action video and the target speech corresponding to the lip action video are generated according to the speech information and the lip feature information, including:

[0027] converting the speech information into corresponding speech text;

[0028] determining a target lip action model from a plurality of pre-set candidate lip action models according to the lip feature information;

[0029] generating a target lip action video according to the target lip action model, the speech information, and the speech text.

[0030] generating a target voice corresponding to the target lip movement video according to the voice information and the voice text.

[0031] In some embodiments, generating a target voice corresponding to the target lip movement video according to the voice information and the voice text comprises:

[0032] generating voice pause information corresponding to the voice text according to the voice information;

[0033] generating a voice text sequence according to the voice pause information and the voice text;

[0034] inputting the voice text sequence into a preset voice synthesis model to synthesize the target voice.

[0035] In some embodiments, generating a face fusion video of a target object according to a lip movement video and a face video comprises:

[0036] generating a virtual face video according to the face video;

[0037] identifying a pixel partition corresponding to a lip in the virtual face video, and performing elimination processing on the pixel partition to obtain a preliminary face video;

[0038] obtaining position information of the pixel partition in the virtual face video, and fusing the lip movement video and the preliminary face video according to the position information to obtain the face fusion video.

[0039] In a second aspect, the embodiments of the present application further provide a virtual image video generation device, comprising:

[0040] a preprocessing module configured to obtain a to-be-processed video, and extract a preprocessed video based on a face feature of a target object in the to-be-processed video;

[0041] a target extraction module configured to extract a face video, a background video and voice information of the target object in the preprocessed video;

[0042] a lip processing module configured to determine lip feature information of the target object based on the face video, and generate a lip movement video and a target voice corresponding to the lip movement video according to the voice information and the lip feature information;

[0043] a face processing module configured to generate a face fusion video of the target object according to the lip movement video and the face video;

[0044] a video synthesis module configured to synthesize a virtual image video according to the face fusion video, the background video and the target voice.

[0045] In a third aspect, the embodiments of the present application further provide a computer device, which comprises a memory and a processor.

[0046] a memory for storing the computer program;

[0047] a processor for executing the computer program and implementing the virtual figure video generation method provided by any of the embodiments of the present application when the computer program is executed.

[0048] In a fourth aspect, the embodiments of the present application further provide a computer readable storage medium, which stores a computer program, and the computer program, when executed by a processor, causes the processor to implement the virtual figure video generation method provided by any of the embodiments of the present application.

[0049] The embodiments of the present application provide a virtual figure video generation method, device, computer device and medium, wherein the virtual figure video generation method comprises: obtaining a to-be-processed video, and extracting a pre-processed video based on a face feature of a target object in the to-be-processed video; extracting a face video, a background video and voice information of the target object in the pre-processed video; determining lip feature information of the target object based on the face video, and generating a lip action video and a target voice corresponding to the lip action video according to the voice information and the lip feature information; generating a face fusion video of the target object according to the lip action video and the face video; and synthesizing a virtual figure video according to the face fusion video, the background video and the target voice, so that the lip action of the target object in the virtual figure video is consistent with the virtual voice, and the obtained virtual figure video can better restore the actual action of the lip of the target object. BRIEF DESCRIPTION OF DRAWINGS

[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are only some of the embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0051] Figure 1 is a step flowchart of a virtual figure video generation method provided by the embodiments of the present application;

[0052] Figure 2 is Figure 1 is a flowchart of a lip feature information determination step in the virtual figure video generation method;

[0053] Figure 3 is Figure 1 is a flowchart of a target voice generation step in the virtual figure video generation method;

[0054] Figure 4 is Figure 3A flowchart of a target speech generation step in a target speech synthesis step;

[0055] Figure 5 A module structure schematic diagram of a virtual image video generation device provided by an embodiment of the present application;

[0056] Figure 6 A structure schematic block diagram of a computer device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0057] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.

[0058] The flowchart shown in the drawings is only an example and does not necessarily include all the contents and operations / steps, nor does it necessarily execute in the order described. For example, some operations / steps can be further decomposed, combined or partially merged, so the actual execution order may be changed according to the actual situation.

[0059] At present, people can generate personalized virtual images through their own photos or videos, and generate corresponding virtual image videos based on the virtual images, and when sharing and communicating on social platforms, people can use their personalized virtual image videos to replace traditional real videos, which can effectively protect their privacy. However, when generating virtual image videos directly through photos or videos, the generation effect of the head picture and the action of the virtual image cannot meet the actual use requirements, especially the problem of splitting the lip action of the target object in the virtual image video and the virtual voice, which will lead to the virtual image video obtained to be relatively poor in restoring the actual facial action of the target object.

[0060] Based on this, the embodiments of the present application provide a virtual image video generation method, device, equipment and medium, aiming to make the lip action of the target object in the generated virtual image video consistent with the virtual voice, and make the virtual image video obtained can better restore the actual action of the lips of the target object. Among them, the virtual image video generation method can be applied to computer, intelligent robot, independent server or server cluster and other electronic equipment, which is not limited here.

[0061] In this embodiment, the virtual image video generation method is applied to a computer as an example, but it is not limited to that the virtual image video generation method can only be used for computers.

[0062] Some embodiments of the present application will be described in detail with reference to the drawings, which are shown schematically. The following embodiments and features can be combined with each other, without conflict.

[0063] Please refer to Figure 1 , Figure 1 A step schematic flow chart of a virtual image video generation method provided by an embodiment of the present application is shown in the figure. The method specifically includes the following steps S1-S6.

[0064] In step S1, a to-be-processed video is obtained, and a pre-processed video is extracted from the to-be-processed video based on a facial feature of a target object.

[0065] The computer executing the method can download the to-be-processed video according to the input instruction or link, or call the to-be-processed video through the input video storage location. For example, the to-be-processed video can be obtained by receiving a video link of the to-be-processed video by the computer executing the method, and downloading the to-be-processed video based on the video link. Alternatively, the to-be-processed video can be obtained by directly receiving the to-be-processed video by the computer executing the method.

[0066] Specifically, the to-be-processed video contains the face of at least one target object. It can be understood that the target object can be a target person or a target animal, and the present application does not limit the target object.

[0067] For the sake of convenience, the following embodiments will be described by taking the target object as a target person as an example.

[0068] After obtaining the to-be-processed video, the computer executing the method extracts a pre-processed video from the to-be-processed video based on the facial feature of the target object. Specifically, extracting the pre-processed video from the to-be-processed video based on the facial feature of the target object includes: extracting a plurality of image frames from the to-be-processed video, performing face recognition on the image frames according to the facial feature, and taking the image frames containing the facial feature as target image frames; when both adjacent image frames are target image frames, extracting a video segment between the adjacent image frames as a unit video, and merging the unit videos to obtain a pre-processed video corresponding to the target object.

[0069] It should be understood that the to-be-processed video can be a plurality of discrete to-be-processed videos or a continuous to-be-processed video. When the to-be-processed video is a plurality of discrete to-be-processed videos, extracting the pre-processed video from the to-be-processed video based on the facial feature of the target object can include: first determining the unit videos in each to-be-processed video, and then merging the unit videos in the plurality of to-be-processed videos to obtain the pre-processed video.

[0070] In some embodiments, before the pre-processing video is extracted from the to-be-processed video based on the facial feature of the target object, the method further includes: obtaining the facial feature of the target object. Specifically, the facial feature of the target object can be obtained by receiving a processing instruction input by a user and parsing the processing instruction to obtain the facial feature of the target object carried in the processing instruction.

[0071] By determining the target image frame in the to-be-processed video based on the facial feature of the target object and extracting the pre-processing video according to the target image frame, it can be ensured that the pre-processing video obtained contains the face of the target object, and the part of the to-be-processed video containing the face of the target object is removed, thereby further improving the work efficiency of synthesizing the virtual image video based on the pre-processing video.

[0072] Step S2: extracting the face video, the background video, and the voice information of the target object from the pre-processing video.

[0073] It should be understood that the face video of the target object is a video within a region corresponding to the face of the target object, the background video is a video within a region corresponding to the face of the target object, and the voice information can be an audio track extracted from the pre-processing video or text information obtained by performing voice recognition on the audio track extracted from the pre-processing video.

[0074] In some embodiments, the face video and the background video of the target object are extracted from the pre-processing video, including:

[0075] extracting a plurality of single-frame images from the pre-processing video;

[0076] performing semantic segmentation processing on the single-frame images to obtain at least one face connected domain corresponding to the face of the target object;

[0077] generating a dynamic face mask model corresponding to the target object according to the face connected domain in the plurality of single-frame images;

[0078] extracting the face video and the background video from the pre-processing video based on the dynamic face mask model.

[0079] Specifically, the computer executing the method extracts a plurality of single-frame images from the pre-processing video, inputs the single-frame images into a preset semantic segmentation model, and performs semantic segmentation processing on the single-frame images to obtain at least one face connected domain corresponding to the face of the target object.

[0080] For example, the preset semantic segmentation model can adopt a Deeplab model or a Mask R-CNN model. Taking the semantic segmentation of a single frame image by using the Deeplab model as an example, the Deeplab model is a semantic segmentation model that can determine the position, object category and contour of each element in a single frame image, and extract multiple objects in the image according to the recognition result to obtain multiple connected domains corresponding to different object categories. In this embodiment, the single frame image is segmented based on the semantic segmentation model, and the regions corresponding to different semantics in the image to be recognized, such as the face connected domain corresponding to the face of the target object, can be determined.

[0081] After obtaining the face connected domain, a dynamic face mask model corresponding to the target object is generated according to the face connected domain. It should be noted that the face connected domains corresponding to multiple different single frame images in the preprocessed video are dynamic connected domains, that is, the boundaries of the face connected domains change over time. The dynamic face mask model corresponding to the target object can be generated according to the face connected domains in multiple single frame images.

[0082] It should be understood that the dynamic face mask model matches the pixel size of the single frame image. Generating the dynamic face mask model corresponding to the target object according to the face connected domain specifically includes: merging the face connected domains corresponding to the face of the semantic target object, taking the pixels corresponding to the face connected domains as foreground and marking them as "1", taking the pixels in the single frame image other than the face connected domains as background and marking them as "0", thereby obtaining a binary mask image of the face of the semantic target object, and then generating a dynamic face mask model based on the binary mask images corresponding to multiple single frame images at different time points. It should be noted that the dynamic face mask model includes the binary mask images corresponding to different single frame images.

[0083] After obtaining the dynamic face mask model, the face video and the background video are extracted in the preprocessed video based on the dynamic face mask model. Specifically, the binary mask image corresponding to the single frame image is obtained based on the dynamic face mask model, the face video can be extracted from multiple single frame images according to the pixels marked as "1" in the binary mask image, and the background video can be extracted from multiple single frame images according to the pixels marked as "0" in the binary mask image.

[0084] In step S3, the lip feature information of the target object is determined based on the face video, and the lip action video and the target voice corresponding to the lip action video are generated according to the voice information and the lip feature information.

[0085] After extracting the facial video, the computer executing the present method determines the lip feature information of the target object based on the facial video. It should be understood that the facial video is dynamic, so the lip shape feature information or lip color feature information of the target object in static state can be determined based on the facial video, and the lip movement feature information of the target object in dynamic state can also be determined.

[0086] It should also be understood that the lip shape feature information is used to characterize the outer contour shape of the target object's lips, the lip color feature information is used to characterize the lip color of the target object, and the lip motion feature information is used to characterize the motion characteristics of the target object's lips when speaking, for example, the movement position and movement trend of the target object's lips when speaking the target word.

[0087] like Figure 2 As shown, in some embodiments, determining the lip feature information of the target object based on the facial video in step S3 specifically includes steps S310 to S340:

[0088] Step S310: extracting multiple facial image frames from the facial video;

[0089] Step S320: performing differential processing on the multiple facial image frames to obtain a lip connected domain corresponding to the lips of the target object in the facial image frames;

[0090] Step S330: extracting a lip image from the facial image frame according to the lip connected domain;

[0091] Step S340: Input the lip image into a preset feature extraction model to extract lip feature information of the target object, wherein the lip feature information includes at least one of lip shape feature information, lip color feature information and lip movement feature information.

[0092] Specifically, a plurality of time-discrete facial image frames are first extracted from the facial video. For example, a plurality of facial image frames uniformly distributed at corresponding times can be extracted from the facial video. Then, differential processing is performed on the plurality of facial image frames to obtain a lip connected domain corresponding to the lips of the target object in the facial image frame. A lip image is extracted from the facial image frame according to the lip connected domain, and the lip image is input into a preset feature extraction model to extract lip feature information of the target object, wherein the lip feature information includes at least one of lip shape feature information, lip color feature information and lip movement feature information.

[0093] It should be understood that, taking the target object as the target person as an example, the target person's mouth will move when speaking. By performing differential processing on multiple facial image frames, the area corresponding to the moving elements in the facial video can be quickly determined, and this area can be used as the lip connected domain corresponding to the target object's lips, which greatly speeds up the speed of extracting the lip connected domain.

[0094] In some embodiments, the plurality of facial image frames are differentially processed in step S340 to obtain a lip connected domain of a corresponding target object lip in the facial image frames, specifically comprising:

[0095] differentially processing the facial image frames adjacent in time to obtain a differential image frame;

[0096] dividing the facial image frame into a plurality of sub-regions;

[0097] determining an average differential value of the plurality of sub-regions in the facial image frame according to the differential image frame;

[0098] merging the sub-regions adjacent in position according to the average differential value to obtain at least one pending connected domain, and determining the lip connected domain in the pending connected domain.

[0099] Specifically, when determining the lip connected domain, the computer executing the method first needs to differentially process the facial image frames adjacent in time to obtain a differential image frame. It should be understood that the differential image frame can be obtained by differentially processing at least two facial image frames adjacent in time. Taking the example of differentially processing two facial image frames adjacent in time, each pixel in the differential image frame corresponds to the pixels in the opposite positions in the two facial image frames adjacent in time, and the pixel value of each pixel in the differential image frame is the difference between the pixel values of the pixels in the opposite positions in the two facial image frames adjacent in time.

[0100] After obtaining the differential image frame, the facial image frame is divided into a plurality of sub-regions, the average differential value of the plurality of sub-regions in the facial image frame is determined according to the differential image frame, and then the sub-regions adjacent in position are merged according to the average differential value to obtain at least one pending connected domain, and the lip connected domain is determined in the pending connected domain.

[0101] Specifically, the sub-region includes at least one pixel, and the pixels in the facial image frame are divided according to a preset rule, so that the facial image frame can be divided into a plurality of sub-regions. For example, the facial image frame can be divided into a plurality of sub-regions with rectangular boundary shapes according to the two sets of mutually perpendicular auxiliary lines.

[0102] The mapping pixels corresponding to the pixels in the sub-region of the facial image frame on the differential image frame are determined, and the average value of the pixel values of the plurality of mapping pixels corresponding to the sub-region is calculated, so as to obtain the average differential value of the plurality of sub-regions in the facial image frame.

[0103] After that, the sub-regions with similar and adjacent average difference values are merged to obtain at least one pending connected domain in the face image frame, and a lip connected domain is determined in the pending connected domain. Wherein, the lip connected domain in the pending connected domain is determined specifically by determining the average difference values of the plurality of sub-regions in the pending connected domain, and calculating the connected domain difference value corresponding to the pending connected domain based on the average difference values of the plurality of sub-regions in the pending connected domain. When the connected domain difference value is greater than a preset threshold, the corresponding pending connected domain is taken as the lip connected domain, thereby greatly accelerating the speed of extracting the lip connected domain.

[0104] Based on the obtained lip connected domain, the lip image is extracted in the face image frame, and the lip image is input into a preset feature extraction model to extract the lip feature information of the target object, wherein the lip feature information includes at least one of lip shape feature information, lip color feature information and lip movement feature information, and then the lip action video and the target voice corresponding to the lip action video are generated according to the voice information and the lip feature information.

[0105] As shown in Figure 3 In some embodiments, the step S3 of generating the lip action video and the target voice corresponding to the lip action video according to the voice information and the lip feature information specifically includes steps S350-S380:

[0106] Step S350: converting the voice information into corresponding voice text;

[0107] Step S360: determining a target lip action model from a plurality of preset candidate lip action models according to the lip feature information;

[0108] Step S370: generating a target lip action video according to the target lip action model, the voice information and the voice text;

[0109] Step S380: generating a target voice corresponding to the target lip action video according to the voice information and the voice text.

[0110] Specifically, the voice information can be an audio track extracted from the preprocessed video, or text information obtained by performing voice recognition on the audio track extracted from the preprocessed video. When the voice information is an audio track extracted from the preprocessed video, the voice information is subjected to voice-to-text recognition to convert the voice information into corresponding voice text, a target lip action model is determined from a plurality of preset candidate lip action models according to the lip feature information, and then a target lip action video is generated according to the target lip action model, the voice information and the voice text. The target voice corresponding to the target lip action video is generated according to the voice information and the voice text.

[0111] In some embodiments, a computer executing the present method is pre-configured with several candidate lip movement models, the candidate lip movement models including corresponding virtual lip maps and lip drives for driving the virtual lips to move, and determining the target lip movement model from several preset candidate lip movement models according to the lip feature information specifically includes: determining the corresponding virtual lip map according to at least one of the lip shape feature information and the lip color feature information, determining the corresponding lip drive according to the lip movement feature information, and then determining the target lip movement model from several preset candidate lip movement models based on the corresponding virtual lip map and the lip drive.

[0112] By combining static lip shape feature information, lip color feature information and dynamic lip movement feature information to determine the target lip action model, the subsequent target lip action video and virtual image video can better restore the actual lip movements of the target object.

[0113] It should be understood that the target subject often pauses at certain positions when speaking. In some embodiments, step S370 specifically includes: generating speech pause information corresponding to the speech text based on the speech information, wherein the speech pause information is used to represent the first pause time length of the pause between adjacent characters or words in the speech text. Then, a speech text timeline is established, and the characters or words in the speech text are marked on the speech text timeline based on the speech pause information to obtain a speech text sequence carrying the speech pause information. The speech text sequence carrying the speech pause information is input into the target lip movement model so that the target lip movement model generates a target lip movement video based on the characters or words in the speech text sequence, the first pause time length of the pause between adjacent characters or words, the corresponding virtual lip map and the lip drive, wherein the target lip movement video includes the movement of the virtual lips when speaking according to the speech text sequence.

[0114] like Figure 4 As shown, in some embodiments, step S380 specifically includes steps S381 to S383:

[0115] Step S381: generating speech pause information corresponding to the speech text according to the speech information;

[0116] Step S382: generating a speech text sequence according to the speech pause information and the speech text;

[0117] Step S383: Input the speech text sequence into a preset speech synthesis model to synthesize the target speech.

[0118] The computer executing the method first generates speech pause information corresponding to the speech text according to the speech information, the speech pause information being used to represent a first pause time length of pauses between adjacent words or phrases in the speech text, and the computer executing the method further generates the speech pause information according to a plurality of first pause time lengths.

[0119] The speech text sequence is generated according to the speech pause information and the speech text, specifically including: establishing a speech text time axis, and labeling words or phrases in the speech text on the speech text time axis according to the speech pause information to obtain a speech text sequence carrying the speech pause information, and then inputting the speech text sequence carrying the speech pause information into a preset speech synthesis model to synthesize a virtual target speech.

[0120] The inputting of the speech text sequence into the preset speech synthesis model specifically includes: inputting the words or phrases into the speech synthesis model according to the order in the speech text sequence according to the first pause time length, that is, the target speech output by the speech synthesis model also carries the speech pause information, and the second pause time length of pauses between corresponding words or phrases is the same as the first pause time length of pauses between adjacent words or phrases in the speech text.

[0121] It should be understood that the second pause time length of pauses between adjacent words or phrases in the target speech is synchronized in time with the motion of the virtual lips in the target lip motion video when speaking according to the speech text sequence, so that the target speech and the target lip motion video are mutually consistent, and the problem of splitting of the lip motion and the virtual speech is solved.

[0122] Step S4, generating a face fusion video of the target object according to the lip motion video and the face video.

[0123] In some embodiments, generating the face fusion video of the target object according to the lip motion video and the face video includes:

[0124] Generating a virtual face video according to the face video;

[0125] Identifying a pixel partition corresponding to the lips in the virtual face video, and performing elimination processing on the pixel partition to obtain a preliminary face video;

[0126] Obtaining position information of the pixel partition in the virtual face video, and fusing the lip motion video and the preliminary face video according to the position information to obtain the face fusion video.

[0127] It should be understood that the virtual face video is a virtual video corresponding to the face of the target object, and the video material in the face video can be replaced with cartoon material or other material. It should be noted that the virtual face video includes pixels corresponding to the lips, and the pixel partition corresponding to the lips can conflict with the subsequent fused lip action video, or the subsequent fused lip action video cannot completely cover the pixels corresponding to the lips in the virtual face video, thereby generating an incorrect face fusion video. Therefore, the pixel partition corresponding to the lips needs to be eliminated.

[0128] Specifically, the virtual face video is generated according to the face video, the pixel partition corresponding to the lips in the virtual face video is identified, and the pixel partition is eliminated to obtain a preliminary face video. For example, the face material of the pixel partition except the pixels corresponding to the lips in the virtual face video can be sampled, and the pixel partition can be filled based on the face material to eliminate the pixel partition to obtain the preliminary face video. Then, the position information of the pixel partition in the virtual face video is obtained, and the lip action video and the preliminary face video are fused according to the position information to obtain the face fusion video.

[0129] By eliminating the pixel partition corresponding to the lips first, and then fusing the lip action video to generate the face fusion video, the problem that the pixel partition corresponding to the lips can conflict with the subsequently fused lip action video, or the subsequently fused lip action video cannot completely cover the pixels corresponding to the lips in the virtual face video is avoided, thereby improving the quality of the generated face fusion video.

[0130] Step S5, synthesizing the virtual image video according to the face fusion video, the background video and the target voice.

[0131] After obtaining the face fusion video, the virtual image video is synthesized according to the face fusion video, the background video and the target voice. It should be understood that the key single sentence and the to-be-processed video are associated with the same time axis. The face fusion video, the background video and the target voice are synthesized according to the corresponding time nodes of the face fusion video, the background video and the target voice on the time axis, and the relative positions of the face fusion video, the background video and the target voice in the preprocessed video picture, so as to obtain the virtual image video in which the lip action is consistent with the target voice.

[0132] In summary, the virtual image video generation method provided by the application can be applied to a computer, and the virtual image video generation method specifically comprises: obtaining a to-be-processed video, and extracting a pre-processed video in the to-be-processed video based on facial features of a target object; extracting a facial video, a background video and voice information of the target object in the pre-processed video; determining lip feature information of the target object based on the facial video, and generating a lip action video and target voice corresponding to the lip action video according to the voice information and the lip feature information; generating a facial fusion video of the target object according to the lip action video and the facial video; and synthesizing a virtual image video according to the facial fusion video, the background video and the target voice, so that the lip action of the target object in the virtual image video is consistent with the virtual voice, thereby solving the problem of split between the lip action and the virtual voice in the prior virtual image video generation method, and the target lip action model is determined by combining static lip shape feature information, lip color feature information and dynamic lip movement feature information, so that the virtual image video can better restore the actual action of the lip of the target object.

[0133] Figure 5 A module structure schematic diagram of a virtual image video generation device provided by an embodiment of the application is shown in FIG. 6, which comprises: Figure 5

[0134] A pre-processing module 601 is configured to obtain a to-be-processed video, and extract a pre-processed video in the to-be-processed video based on facial features of a target object;

[0135] A target extraction module 602 is configured to extract a facial video, a background video and voice information of the target object in the pre-processed video;

[0136] A lip processing module 603 is configured to determine lip feature information of the target object based on the facial video, and generate a lip action video and target voice corresponding to the lip action video according to the voice information and the lip feature information;

[0137] A facial processing module 604 is configured to generate a facial fusion video of the target object according to the lip action video and the facial video;

[0138] A video synthesis module 605 is configured to synthesize a virtual image video according to the facial fusion video, the background video and the target voice.

[0139] In some embodiments, the target extraction module 602 extracts the facial video and the background video of the target object in the pre-processed video, specifically comprising:

[0140] extracting a plurality of single-frame images in the pre-processed video;

[0141] performing semantic segmentation processing on the single-frame images to obtain at least one facial connected domain corresponding to the face of the target object;​

[0142] generating a dynamic face mask model of the corresponding target object according to the face connected domain in the plurality of single-frame images;

[0143] extracting a face video and a background video in the preprocessed video based on the dynamic face mask model.

[0144] In some embodiments, the target extraction module 602 differentiates the plurality of face image frames to obtain a lip connected domain of the lips of the corresponding target object in the face image frames, specifically including:

[0145] differentiating the time-adjacent face image frames to obtain a difference image frame;

[0146] dividing the face image frames into a plurality of sub-regions;

[0147] determining an average difference value of the plurality of sub-regions in the face image frames according to the difference image frame;

[0148] merging the position-adjacent sub-regions according to the average difference value to obtain at least one pending connected domain, and determining the lip connected domain in the pending connected domain.

[0149] In some embodiments, the lip processing module 603 determines the lip feature information of the target object based on the face video, specifically including:

[0150] extracting a plurality of face image frames in the face video;

[0151] differentiating the plurality of face image frames to obtain a lip connected domain of the lips of the corresponding target object in the face image frames;

[0152] extracting a lip image in the face image frames according to the lip connected domain;

[0153] inputting the lip image into a preset feature extraction model to extract the lip feature information of the target object, wherein the lip feature information includes at least one of lip shape feature information, lip color feature information, and lip motion feature information.

[0154] In some embodiments, the lip processing module 603 differentiates the plurality of face image frames to obtain a lip connected domain of the lips of the corresponding target object in the face image frames, specifically including:

[0155] differentiating the time-adjacent face image frames to obtain a difference image frame;

[0156] dividing the face image frames into a plurality of sub-regions;

[0157] determining an average difference value of the plurality of sub-regions in the face image frames according to the difference image frame;

[0158] Merge the sub-regions adjacent to the position according to the average difference value to obtain at least one pending connected domain, and determine a lip connected domain in the pending connected domain.

[0159] In some embodiments, the lip processing module 603 generates a lip action video and target speech corresponding to the lip action video according to the speech information and the lip feature information, specifically including:

[0160] Converting the speech information into corresponding speech text;

[0161] Determining a target lip action model from a plurality of preset candidate lip action models according to the lip feature information;

[0162] Generating a target lip action video according to the target lip action model, the speech information and the speech text;

[0163] Generating target speech corresponding to the target lip action video according to the speech information and the speech text.

[0164] In some embodiments, the lip processing module 603 generates a lip action video and target speech corresponding to the lip action video according to the speech information and the lip feature information, specifically including:

[0165] Generating speech pause information corresponding to the speech text according to the speech information;

[0166] Generating a speech text sequence according to the speech pause information and the speech text;

[0167] Inputting the speech text sequence into a preset speech synthesis model to synthesize and obtain target speech.

[0168] In some embodiments, the face processing module 604 generates a face fusion video of a target object according to the lip action video and the face video, specifically including:

[0169] Generating a virtual face video according to the face video;

[0170] Identifying a pixel partition corresponding to the lip in the virtual face video, and performing elimination processing on the pixel partition to obtain a preliminary face video;

[0171] Obtaining position information of the pixel partition in the virtual face video, and fusing the lip action video and the preliminary face video according to the position information to obtain the face fusion video.

[0172] Please refer to Figure 6 , Figure 6 A structural schematic block diagram of a computer device provided by an embodiment of the present application.

[0173] As Figure 6As shown, the computer device 700 includes a processor 701 and a memory 702, which are connected through a bus 703, such as an I2C (Inter-integrated Circuit) bus.

[0174] Specifically, the processor 701 is configured to provide computing and control capabilities to support the operation of the entire computer device. The processor 701 can be a central processing unit (CPU), and the processor 701 can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gates or transistor logic components, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.

[0175] Specifically, the memory 702 can be a Flash chip, a read-only memory (ROM) disk, an optical disk, a U disk or a mobile hard disk, etc.

[0176] Those skilled in the art can understand that, Figure 6 The structure shown in the figure is only a block diagram of part of the structure related to the embodiment of the present application, and does not constitute a limitation on the computer device to which the embodiment of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.

[0177] The processor is configured to run a computer program stored in the memory, and implement the virtual image video generation method provided by any one of the embodiments of the present application when the computer program is executed.

[0178] In some embodiments, the processor 701 is configured to run a computer program stored in the memory 702, and implement the following steps when the computer program is executed:

[0179] Obtaining a to-be-processed video, and extracting a pre-processed video in the to-be-processed video based on a facial feature of a target object;

[0180] Extracting a facial video, a background video and voice information of the target object in the pre-processed video;

[0181] The lip feature information of the target object is determined based on the facial video, and a lip movement video and target speech corresponding to the lip movement video are generated according to the speech information and the lip feature information;

[0182] The facial fusion video of the target object is generated according to the lip movement video and the facial video;

[0183] The virtual image video is synthesized according to the facial fusion video, the background video and the target speech.

[0184] In some embodiments, the processor 701 includes the following when extracting the facial video and the background video of the target object in the pre-processed video:

[0185] extracting a plurality of single-frame images in the pre-processed video;

[0186] performing semantic segmentation processing on the single-frame images to obtain at least one facial connected domain corresponding to the face of the target object;

[0187] generating a dynamic facial mask model of the target object according to the facial connected domain in the plurality of single-frame images;

[0188] extracting the facial video and the background video in the pre-processed video based on the dynamic facial mask model.

[0189] In some embodiments, the processor 701 includes the following when determining the lip feature information of the target object based on the facial video:

[0190] extracting a plurality of facial image frames in the facial video;

[0191] performing difference processing on the plurality of facial image frames to obtain a lip connected domain corresponding to the lips of the target object in the facial image frames;

[0192] extracting a lip image in the facial image frames according to the lip connected domain;

[0193] inputting the lip image into a pre-set feature extraction model to extract the lip feature information of the target object, wherein the lip feature information includes at least one of lip shape feature information, lip color feature information and lip movement feature information.

[0194] In some embodiments, the processor 701 includes the following when performing difference processing on the plurality of facial image frames to obtain the lip connected domain corresponding to the lips of the target object in the facial image frames:

[0195] performing difference processing on the facial image frames adjacent in time to obtain a difference image frame;

[0196] dividing the facial image frame into a plurality of sub-regions;

[0197] determining average difference values of the plurality of sub-regions in the facial image frame according to the difference image frame;

[0198] merge the sub-regions adjacent to the position according to the average difference value to obtain at least one pending connected domain, and determine a lip connected domain in the pending connected domain.

[0199] In some embodiments, the processor 701 comprises, when generating the lip action video and the target speech corresponding to the lip action video according to the speech information and the lip feature information:

[0200] converting the speech information into corresponding speech text;

[0201] determining a target lip action model from a plurality of preset candidate lip action models according to the lip feature information;

[0202] generating a target lip action video according to the target lip action model, the speech information and the speech text;

[0203] generating a target speech corresponding to the target lip action video according to the speech information and the speech text.

[0204] In some embodiments, the processor 701 comprises, when generating a target speech corresponding to a target lip action video according to speech information and speech text:

[0205] generating speech pause information corresponding to the speech text according to the speech information;

[0206] generating a speech text sequence according to the speech pause information and the speech text;

[0207] inputting the speech text sequence into a preset speech synthesis model to synthesize and obtain the target speech.

[0208] In some embodiments, the processor 701 comprises, when generating a face fusion video of a target object according to the lip action video and the face video:

[0209] generating a virtual face video according to the face video;

[0210] identifying a pixel partition corresponding to the lip in the virtual face video, and performing elimination processing on the pixel partition to obtain a preliminary face video;

[0211] obtaining position information of the pixel partition in the virtual face video, and fusing the lip action video and the preliminary face video according to the position information to obtain the face fusion video.

[0212] It should be noted that those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the computer device described above can refer to the corresponding process in the foregoing virtual image video generation method embodiments, which will not be described here.

[0213] The embodiment of the present application further provides a storage medium, which stores a computer program, and the computer program can be executed by one or more processors to implement the steps of any virtual image video generation method provided in the specification of the embodiment of the present application.

[0214] The storage medium can be an internal storage unit of the computer device, for example, a hard disk or a memory of the computer device. The storage medium can also be an external storage device of the computer device, for example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card and the like.

[0215] Those skilled in the art can understand that all or some of the steps in the method disclosed above, the functions of the modules / units in the system and the device can be implemented as software, firmware, hardware or a suitable combination thereof. In the hardware embodiment, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, one physical component can have multiple functions, or one function or step can be performed by several physical components in cooperation. Some or all of the physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor or a microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer readable medium, which can include computer storage media (or non-transitory media) and communication media (or transitory media). As known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. In addition, it is known to those skilled in the art that communication media typically includes computer readable instructions, data structures, program modules or other data in modulated data signals such as carrier waves or other transport mechanisms, and can include any information delivery medium.

[0216] In the description of the application, unless otherwise clearly specified and limited, the terms "mounting", "connection", "connecting" should be understood in a broad sense, for example, can be fixedly connected, can be detachably connected, or integrally connected; can be mechanically connected, can be electrically connected; can be directly connected, can be indirectly connected through an intermediate medium, and can be internal communication of two elements. For those skilled in the art, the specific meanings of the above terms in the application can be understood according to the specific circumstances.

[0217] It should be understood that the terms used herein in the specification and the appended claims should not be construed as limiting the application to specific embodiments. As used in this specification and the appended claims, the singular forms "a", "an" and "the" include plural referents unless the context clearly dictates otherwise.

[0218] It should also be understood that the term "and / or" as used herein in the specification and in the claims, means any one of the items, or combinations of any of the items, listed together with all possible combinations thereof. It should be noted that, as used in this text, the terms "comprises", "comprising", or any other variation thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements, but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising a" does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.

[0219] The above-mentioned application embodiment serial number is only for description, not representing the advantages and disadvantages of the embodiments. The above description is only for specific embodiments of the application, but the protection scope of the application is not limited thereto. Any skilled person in the art can easily think of various equivalent modifications or replacements within the technical scope disclosed by the application, and these modifications or replacements should be covered in the protection scope of the application. Therefore, the protection scope of the application should be subject to the protection scope of the claims.

Claims

1. A method for generating a virtual image video, characterized in that: The method comprises: Acquire a video to be processed, and extract a pre-processed video from the video to be processed based on facial features of a target object; Extracting facial video, background video and voice information of the target object from the preprocessed video; Determining lip feature information of the target object based on the facial video, and generating a lip movement video and a target voice corresponding to the lip movement video according to the voice information and the lip feature information; generating a facial fusion video of the target object according to the lip movement video and the facial video; synthesizing a virtual image video according to the facial fusion video, the background video and the target voice; The determining of the lip feature information of the target object based on the facial video includes: extracting a plurality of facial image frames from the facial video; performing differential processing on the plurality of facial image frames to obtain a lip connected region corresponding to the lips of the target object in the facial image frames; extracting a lip image in the facial image frame according to the lip connected domain; Inputting the lip image into a preset feature extraction model to extract the lip feature information of the target object, wherein the lip feature information includes at least one of lip shape feature information, lip color feature information, and lip movement feature information; Generating a facial fusion video of the target object according to the lip movement video and the facial video includes: generating a virtual facial video based on the facial video; Identifying pixel partitions corresponding to lips in the virtual facial video, and performing elimination processing on the pixel partitions to obtain a preliminary facial video; The position information of the pixel partition in the virtual facial video is obtained, and the lip movement video and the preliminary facial video are fused according to the position information to obtain the facial fusion video.

2. The method according to claim 1, characterized in that Extracting the facial video and background video of the target object from the pre-processed video includes: Extracting multiple single-frame images from the preprocessed video; Performing semantic segmentation processing on the single frame image to obtain at least one facial connected domain corresponding to the face of the target object; generating a dynamic facial mask model corresponding to the target object according to the facial connected domains in the plurality of single-frame images; The facial video and the background video are extracted from the preprocessed video based on the dynamic facial mask model.

3. The method according to claim 1, characterized in that The step of performing differential processing on the plurality of facial image frames to obtain a lip connected region corresponding to the lips of the target object in the facial image frames includes: performing differential processing on the temporally adjacent facial image frames to obtain differential image frames; dividing the facial image frame into a plurality of sub-regions; Determining mapped pixels corresponding to pixels in the subregion of the facial image frame on the differential image frame, and calculating an average of pixel values ​​of a plurality of the mapped pixels corresponding to the subregion to obtain an average differential value of the plurality of subregions in the facial image frame; The adjacent sub-regions having similar average difference values ​​are merged to obtain at least one undetermined connected domain, and the lip connected domain is determined in the undetermined connected domain.

4. The method according to any one of claims 1 to 3, characterized in that The generating of the lip movement video and the target speech corresponding to the lip movement video according to the speech information and the lip feature information includes: Converting the voice information into corresponding voice text; Determining a target lip motion model from a plurality of preset candidate lip motion models according to the lip feature information; generating a target lip movement video according to the target lip movement model, the voice information, and the voice text; The target voice corresponding to the target lip movement video is generated according to the voice information and the voice text.

5. The method according to claim 4, characterized in that Generating the target voice corresponding to the target lip movement video according to the voice information and the voice text includes: Generating voice pause information corresponding to the voice text according to the voice information; generating a speech-to-text sequence according to the speech pause information and the speech text; The speech text sequence is input into a preset speech synthesis model to synthesize and obtain the target speech.

6. A virtual image video generating device, characterized in that: include: A preprocessing module is used to obtain a video to be processed and extract a preprocessed video from the video to be processed based on facial features of a target object; A target extraction module, configured to extract the facial video, background video, and voice information of the target object from the pre-processed video; a lip processing module, configured to determine lip feature information of the target object based on the facial video, and generate a lip movement video and a target voice corresponding to the lip movement video according to the voice information and the lip feature information; A facial processing module, configured to generate a facial fusion video of the target object based on the lip movement video and the facial video; A video synthesis module, configured to synthesize a virtual image video based on the facial fusion video, the background video, and the target voice; The determining of the lip feature information of the target object based on the facial video includes: extracting a plurality of facial image frames from the facial video; performing differential processing on the plurality of facial image frames to obtain a lip connected region corresponding to the lips of the target object in the facial image frames; extracting a lip image in the facial image frame according to the lip connected domain; Inputting the lip image into a preset feature extraction model to extract the lip feature information of the target object, wherein the lip feature information includes at least one of lip shape feature information, lip color feature information, and lip movement feature information; Generating a facial fusion video of the target object according to the lip movement video and the facial video includes: generating a virtual facial video based on the facial video; Identifying pixel partitions corresponding to lips in the virtual facial video, and performing elimination processing on the pixel partitions to obtain a preliminary facial video; The position information of the pixel partition in the virtual facial video is obtained, and the lip movement video and the preliminary facial video are fused according to the position information to obtain the facial fusion video.

7. A computer device, characterized in that: The computer device includes a memory and a processor; The memory is used to store computer programs; The processor is configured to execute the computer program and implement the virtual image video generation method according to any one of claims 1 to 5 when executing the computer program.

8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, enables the processor to implement the virtual image video generation method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Video synthesis method and device, equipment and storage medium

    CN112866586A

  • Lip shape synchronization face forgery generation method and system based on image completion

    CN114663962A