Virtual character video generation method, device, computer equipment and storage medium
By separating audio and video streams from template videos, using face recognition and audio database matching, generating target lip video streams and integrating, the problem that virtual human videos cannot be designed in personalized manner is solved, and the reality and fluency of virtual human videos are improved.
Patent Information
- Application Number
- CN202210582573.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-26
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2042-05-26
AI Technical Summary
In the prior art, virtual human video generation cannot be personalized according to the image or needs of different users, resulting in the inability to promote and use on a large scale.
By segmenting the video stream and audio stream from the template video, using the face recognition model to calibrate the face area, combining the audio database to match the audio type, performing speech synthesis and lip parameter generation, and finally integrating the target lip video stream with the template video stream, adding the audio stream to generate virtual character video.
The generated virtual character video sound is close to the template video, and the lip state is close to the target character, improving the user experience and the realism and fluency of the video.
Smart Images

Figure CN114998489B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image recognition, and in particular to a method, device, computer equipment and storage medium for generating a virtual character video. Background Art
[0002] In recent years, with the rapid development of video technology, virtual character video generation technology has emerged.
[0003] In the existing technology, virtual humans are generally generated based on three-dimensional animation. However, the generation cycle of virtual humans with three-dimensional animation is long and the cost is high, and they cannot be personalized according to the image or needs of different users. As a result, virtual human videos cannot be widely promoted and used, which has limitations. Summary of the Invention
[0004] Based on this, it is necessary to provide a virtual character video generation method, device, computer equipment and storage medium to address the above technical problems, so as to solve the problem in the existing technology that virtual people cannot be personalized according to the image or needs of different users, resulting in the inability to promote and use virtual human videos on a large scale.
[0005] A method for generating a virtual character video, comprising:
[0006] Splitting a template video stream and a template audio stream from a template video;
[0007] Performing facial region calibration on the template video stream using a face recognition model to generate a target face video stream; searching an audio database for an audio type that matches the template audio stream;
[0008] Performing speech synthesis on a preset text according to the audio type to generate a target audio stream;
[0009] Generate lip shape parameters from the audio type and the text, and process the lip shape parameters and the target face video stream through a lip shape generation model to obtain a target lip shape video stream;
[0010] The target lip shape video stream and the template video stream are merged, and the target audio stream is added to obtain a virtual character video.
[0011] A virtual character video generation device, comprising:
[0012] A template video segmentation module is used to segment the template video into a template video stream and a template audio stream;
[0013] The video and audio processing module is used to calibrate the face area of the template video stream through the face recognition model to generate a target face video stream; and search the audio database for an audio type that matches the template audio stream;
[0014] A speech synthesis module, configured to perform speech synthesis on a preset text according to the audio type to generate a target audio stream;
[0015] A lip shape video stream module is configured to generate lip shape parameters from the audio type and the text, and process the lip shape parameters and the target face video stream through a lip shape generation model to obtain a target lip shape video stream;
[0016] The first virtual character video module is used to merge the target lip shape video stream and the template video stream, and add the target audio stream to obtain the virtual character video.
[0017] A computer device includes a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor. When the processor executes the computer-readable instructions, the above-mentioned virtual character video generation method is implemented.
[0018] One or more readable storage media storing computer-readable instructions, wherein when the computer-readable instructions are executed by one or more processors, the one or more processors execute the above-mentioned virtual character video generation method.
[0019] The above-mentioned virtual character video generation method, device, computer equipment and storage medium are as follows: a template video stream and a template audio stream are segmented from a template video; a face region is calibrated on the template video stream through a face recognition model to generate a target face video stream; an audio type matching the template audio stream is searched in an audio database; a preset text is speech synthesized according to the audio type to generate a target audio stream; lip shape parameters are generated from the audio type and the text, and the lip shape parameters and the target face video stream are processed by a lip shape generation model to obtain a target lip shape video stream; the target lip shape video stream is merged with the template video stream, and the target audio stream is added to obtain a virtual character video. The present invention generates a target audio stream by searching an audio database for an audio type matching the template audio stream, and speech synthesizing the preset text according to the audio type to make the sound of the final generated virtual character video closer to the sound in the template video, thereby improving the user experience. Furthermore, the target face video stream is separated from the template video, and a target lip shape video stream is generated based on the target face video stream and the target audio stream, so that the lip shape state in the generated target lip shape video stream is closer to the state of the target character when speaking, so that the obtained virtual character video is more in line with the user's design requirements, and the realism and smoothness of the final generated virtual character video picture are improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0021] Figure 1 1 is a schematic diagram of an application environment of a method for generating a virtual character video according to an embodiment of the present invention;
[0022] Figure 2 This is a flow chart of a method for generating a virtual character video according to an embodiment of the present invention;
[0023] Figure 3 1 is a schematic structural diagram of a virtual character video generating device according to an embodiment of the present invention;
[0024] Figure 4 FIG. 1 is a schematic diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0025] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0026] The virtual character video generation method provided in this embodiment can be applied in Figure 1 In an application environment, a client communicates with a server. Clients include, but are not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented as a standalone server or a server cluster consisting of multiple servers.
[0027] In one embodiment, if Figure 2 As shown, a method for generating a virtual character video is provided, and the method is applied in Figure 1 The server in the example is used as an example, and the steps are as follows:
[0028] S10. Segment the template video into a template video stream and a template audio stream.
[0029] As can be understood, the template video contains information such as the target person, the target person's speaking voice, and movements. Movement information includes head movements, facial movements, lip movements, eye movements, and hand movements. Generally, the target person's face in the template video faces forward, with minimal eye movement and a natural facial expression. The template video stream refers to a video containing the target person and their speaking movement information. The template audio stream refers to an audio stream containing the target person's speaking voice information. Using video segmentation technology, the template video stream and template audio stream are segmented from the template video.
[0030] S20. Demarcate the face area of the template video stream using a face recognition model to generate a target face video stream; and search an audio database for an audio type that matches the template audio stream.
[0031] It can be understood that the face recognition model is used to identify the face of the target person in the template video stream, calibrate the face area, separate the face area from the template video stream, and generate a target face video stream. Among them, the target face video stream refers to the video stream containing the face of the target person. Generally, the face of the target person refers to the face area including the ears, forehead, eyes, nose and mouth. Calibrating the face area means positioning and marking the area where the face is located. For example, the coordinate position of the face outline is calibrated, and according to the calibrated coordinate position, the face area can be separated from the template video stream to generate a target face video stream containing the face area. The audio database refers to a database that pre-stores audio data of different audio types to be matched. By parsing the audio data in the template audio stream, the audio type of the audio data in the template audio stream can be obtained, and the audio type is recorded as the template audio type; through the audio similarity model, the similarity between the template audio type and several audio types to be matched in the audio database is calculated to obtain several audio similarities; the audio type to be matched corresponding to the maximum similarity among the several audio similarities is determined as the audio type that matches the template audio stream.
[0032] S30: Perform speech synthesis on the preset text according to the audio type to generate a target audio stream.
[0033] It is understandable that different audio types correspond to different audio data. Audio data refers to data containing information such as timbre, volume, and intonation. Generally, different people differ in the timbre, volume, and intonation of their speech. By classifying audio data according to information such as timbre, volume, and intonation, the audio type can be obtained. Preset text refers to pre-set text, which can be set according to actual conditions. Speech synthesis refers to the process of converting preset text into audio data through speech synthesis technology based on the audio type that matches the template audio stream. The target audio stream refers to the audio data obtained through speech synthesis technology, which includes a timestamp. Generating a target audio stream based on the audio type that matches the template audio stream can make the timbre information of the generated target audio stream closer to the timbre information of the target person, thereby improving the user experience.
[0034] S40: Generate lip shape parameters from the audio type and the text, and process the lip shape parameters and the target face video stream through a lip shape generation model to obtain a target lip shape video stream.
[0035] It can be understood that the lip shape parameters are obtained by parsing the audio type information and text information contained in the target audio stream through the lip shape generation model. For the same audio type, different texts in the target audio stream correspond to different pronunciations, and different lip shape parameters are generated. Different audio types, corresponding to the same text, also generate different lip shape parameters. Among them, the lip shape parameters refer to the parameters of the lip shape under different pronunciation states. Among them, the lip shape generation model can change the lip shape state of the lips in the target face video stream according to the lip shape parameters, generate lip movements corresponding to the lip shape parameters and a timestamp corresponding to the target audio stream, and then obtain the target lip shape video stream. The target lip shape video stream refers to a video stream containing lip movements generated by the lip shape generation model.
[0036] S50: Fusing the target lip shape video stream and the template video stream, and adding the target audio stream to obtain a virtual character video.
[0037] It can be understood that a virtual character video refers to a video generated based on a target lip-sync video stream and a template video stream. The target lip-sync video stream includes lip movements based on the target audio stream, while the template video stream includes the target character's head, eye, and hand movements while speaking naturally. By fusing the target lip-sync video stream and the template video stream, a virtual character video containing lip movements based on the target audio stream and the target character's head, eye, and hand movements while speaking naturally can be generated.
[0038] In steps S10-S50, a template video stream and a template audio stream are segmented from a template video; the face region of the template video stream is calibrated using a face recognition model to generate a target face video stream; an audio type matching the template audio stream is searched in an audio database; a preset text is speech synthesized based on the audio type to generate a target audio stream; lip shape parameters are generated from the audio type and the text, and the lip shape parameters and the target face video stream are processed using a lip shape generation model to obtain a target lip shape video stream; the target lip shape video stream is fused with the template video stream, and the target audio stream is added to obtain a virtual character video. The present invention searches for an audio type matching the template audio stream in an audio database, and speech synthesizes the preset text based on the audio type to generate a target audio stream, so that the sound of the final generated virtual character video is closer to the sound in the template video, thereby improving the user experience. Furthermore, the target face video stream is separated from the template video, and a target lip shape video stream is generated based on the target face video stream and the target audio stream, so that the lip shape state in the generated target lip shape video stream is closer to the state of the target character when speaking, so that the obtained virtual character video is more in line with the user's design requirements, and the realism and smoothness of the final generated virtual character video picture are improved.
[0039] Optionally, before step S10, that is, before segmenting the template video stream and the template audio stream from the template video, the following steps are included:
[0040] S101, obtaining character feature information from a target video using a character recognition model;
[0041] S102: Cut the target video according to the character feature information to obtain the template video.
[0042] It is understandable that the target video refers to a video recorded of the target person. In particular, the target video is generally no less than one minute long. The character feature information is obtained through a character recognition model, which refers to the feature information of the target person when speaking. The character feature information includes information such as eye features, lip features, facial orientation, and hand features. Among them, the character recognition model is used to identify the character features of the target person in the target video to obtain character feature information. After obtaining the character feature information, the eye features, lip features, facial orientation, hand features and other information in the character feature information are screened, and the character feature information that does not meet the preset requirements is eliminated, and the video clips corresponding to the eliminated character feature information are cut from the target video to obtain a template video. Among them, the preset requirements can be set according to different application scenarios. Generally, video clips in the target video with large eye movements and unnatural facial expressions can be deleted.
[0043] In steps S101-S102, character feature information is obtained from the target video using a character recognition model. The target video is then cropped based on the character feature information to obtain the template video. This cropping of the target video based on the character feature information ensures that the resulting template video better meets the user's design requirements, and the resulting virtual character video further meets the user's needs.
[0044] Optionally, in step S20, the face region calibration is performed on the template video stream using a face recognition model to generate a target face video stream, including:
[0045] S201, performing facial key point detection on video frames in the template video stream using a face recognition model to obtain face regions in a plurality of video frames;
[0046] S202: Segment the plurality of face regions from the template video stream to generate the target face video stream.
[0047] It can be understood that the face recognition model is used to identify the face of the target person in the template video stream, calibrate the face area, separate the face area from the template video stream, and generate a target face video stream. The template video stream is composed of several video frames. By using the face recognition model to detect facial key points on the video frames in the template video stream, several face areas in the video frames containing faces can be obtained. Among them, facial key point detection refers to detecting whether a video frame contains facial key points. If the detected video frame contains facial key points, the face area in the video frame is calibrated according to the detected facial key points, and the calibrated face area is separated from the template video stream. By performing facial key point detection on several video frames, a target face video stream consisting of several face areas can be obtained.
[0048] In steps S201 and S202, facial key point detection is performed on the video frames in the template video stream using a facial recognition model to obtain facial regions in the video frames. These facial regions are then segmented from the template video stream to generate the target face video stream. Segmenting the target person's facial region from the template video stream ensures that the generated target face video stream contains only the facial region, enhancing the realism of the resulting virtual person video.
[0049] Optionally, in step S20, searching the audio database for an audio type that matches the template audio stream includes:
[0050] S203, parsing the template audio stream to obtain a template audio type of the template audio stream;
[0051] S204. Calculate similarities between the template audio type and a plurality of to-be-matched audio types in the audio database using an audio similarity model to obtain a plurality of audio similarities.
[0052] S205: Determine the to-be-matched audio type corresponding to the maximum similarity among the plurality of audio similarities as the audio type that matches the template audio stream.
[0053] It is understandable that the template audio stream refers to audio containing the sound information of the target person speaking. In the process of parsing the template audio stream, audio information such as the timbre, pitch, and tone of the target person's speech can be obtained. Furthermore, the audio data in the template audio stream is classified according to the audio information such as the timbre, pitch, and tone of the target person's speech to obtain the template audio type. Preferably, the audio type can be classified according to the timbre information, or according to the pitch information, or according to the tone information. The classification method of the audio type is not limited here and can be set according to actual needs. The audio similarity model is used to identify the audio type that is closest to the template audio type from several audio types to be matched in the audio database. Specifically, the audio similarity model can calculate the similarity between the template audio type and each audio type to be matched in the audio database, obtain several audio similarities, and use the audio type to be matched corresponding to the maximum similarity among the several audio similarities as the audio type that matches the template audio stream. Among them, the audio similarity refers to the similarity value between the template audio type and the audio type to be matched. The maximum similarity refers to an audio similarity having the largest similarity value among a plurality of audio similarities.
[0054] In steps S203-S205, the template audio stream is parsed to obtain a template audio type for the template audio stream; the similarity between the template audio type and several to-be-matched audio types in an audio database is calculated using an audio similarity model to obtain several audio similarities; and the to-be-matched audio type corresponding to the maximum similarity among the several audio similarities is determined as the audio type that matches the template audio stream. By parsing the template audio stream and determining the to-be-matched audio type corresponding to the maximum similarity among the several audio similarities as the audio type that matches the template audio stream, the timbre information of the target audio stream generated based on the audio type can be made closer to the timbre information of the target person, thereby improving the user experience.
[0055] Optionally, in step S50, the target lip shape video stream and the template video stream are merged, and the target audio stream is added to obtain the virtual character video, including:
[0056] S501, performing character segmentation on the template video stream using a video segmentation technology to obtain a character video stream;
[0057] S502, fusing the character video stream and the target lip shape video stream using a lip shape fusion algorithm to obtain a character lip shape video stream;
[0058] S503: Add a preset background image and the target audio stream to the character's lip shape video stream to obtain the virtual character video.
[0059] It can be understood that video segmentation technology refers to the technology of dividing an image or video into several specific parts or subsets with unique properties according to certain principles, and extracting the target of interest to facilitate higher-level analysis and understanding. Character segmentation refers to the process of segmenting the image of the target character from the template video stream through video segmentation technology. Generally, the image of the target character includes the target character's hair, clothing (clothes, pants, hats and other accessories), hands, and other parts of the body. Preferably, the image of the target character does not include the face area. The character video stream refers to the video stream that includes the target character's image. The lip fusion algorithm is used to fuse the target lip video stream and the character video stream to obtain a character lip video stream. The character lip video stream contains the target character's image and lip movements. The lip movements are generated by the lip generation model based on preset text. According to the timestamps of the target audio stream and the target lip video stream, the target audio stream is added to the target lip video stream to synchronize the picture and sound, thereby improving the user experience. The preset background image refers to a pre-selected background image, which can be set according to actual needs. In actual operation, according to user needs, a background image corresponding to the preset text can be added to the character's lip shape video stream to improve the visual sense of the virtual character video and meet the user's design needs.
[0060] Optionally, in step S503, that is, adding a preset background image and the target audio stream to the character lip shape video stream to obtain the virtual character video, the step includes:
[0061] S5031. Analyze the target audio stream using a head gesture algorithm to obtain a first head movement.
[0062] S5032: updating the second head movement in the lip-shaping video stream of the target person according to the first head movement to obtain a lip-shaping video stream of the target person;
[0063] S5033. Add the target audio stream and the background image to the target character's lip shape video to obtain the virtual character video.
[0064] It is understandable that when the generated target lip-syncing video stream is longer than the template video stream, the head movement of the target person in the person's lip-syncing video stream will repeat the head movement in the template video stream. When a person speaks, different speaking tones and intonations correspond to different head movements. For example, when expressing approval, the head movement can be a nod. By parsing the target audio stream through the head posture algorithm, a head movement corresponding to the speaking tone, intonation, words, etc. in the target audio stream can be generated, that is, the first head movement. The second head movement refers to the original head movement in the person's lip-syncing video stream. According to the first head movement, the second head movement in the person's lip-syncing video stream is updated to obtain the target person's lip-syncing video stream. Among them, the target person's lip-syncing video stream refers to the person's lip-syncing video stream including the first head movement.
[0065] In steps S5031-S5033, the target audio stream is parsed through the head posture algorithm to obtain the first head movement, and the first head movement is updated into the character's lip shape video stream, so that the head movement of the character in the final generated virtual character video matches the speech content, thereby improving the picture sense and user experience of the virtual character video.
[0066] Optionally, after step S40, that is, after the lip shape parameters and the target face video stream are processed by the lip shape generation model to obtain a target lip shape video stream, the following steps are included:
[0067] S401, merging the target lip shape video stream and the target face video stream to obtain a face lip shape video stream;
[0068] S402: Fusing the human face and lip shape video with the template video stream, and adding the target audio stream to obtain a virtual character video.
[0069] Understandably, when the lip shape generation model processes the lip shape parameters and the target face video stream to generate the target lip shape video stream, the lip shape generation model focuses on the lip shape area within the face region, making other areas of the face region susceptible to distortion or blur. By merging the target lip shape video stream with the target face video stream, the face region can be better restored, making the face in the final generated anthropomorphic video closer to the target person, the picture more natural, and improving the user experience.
[0070] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0071] In one embodiment, a virtual character video generation device is provided, which corresponds to the virtual character video generation method in the above embodiment. Figure 3 As shown, the virtual character video generation device includes a template video segmentation module 10, a video and audio processing module 20, a speech synthesis module 30, a lip shape video stream module 40 and a first virtual character video module 50. The functional modules are described in detail as follows:
[0072] The template video segmentation module 10 is used to segment the template video into a template video stream and a template audio stream;
[0073] The video and audio processing module 20 is used to calibrate the face area of the template video stream using a face recognition model to generate a target face video stream; and search the audio database for an audio type that matches the template audio stream;
[0074] The speech synthesis module 30 is used to perform speech synthesis on the preset text according to the audio type to generate a target audio stream;
[0075] A lip shape video stream module 40 is configured to generate lip shape parameters from the audio type and the text, and process the lip shape parameters and the target face video stream using a lip shape generation model to obtain a target lip shape video stream;
[0076] The first virtual character video module 50 is used to merge the target lip shape video stream and the template video stream, and add the target audio stream to obtain the virtual character video.
[0077] Optionally, before the template video segmentation module 10, the following steps are included:
[0078] The character feature module is used to obtain character feature information from the target video through the character recognition model;
[0079] The template video module is used to cut the target video according to the character feature information to obtain the template video.
[0080] Optionally, the video and audio processing module 20 includes:
[0081] A face region unit, configured to detect facial key points on the video frames in the template video stream using a face recognition model to obtain face regions in a plurality of the video frames;
[0082] The target face video stream unit is used to segment the plurality of face regions from the template video stream to generate the target face video stream.
[0083] Optionally, the video and audio processing module 20 further includes:
[0084] A template audio type unit, configured to parse the template audio stream to obtain a template audio type of the template audio stream;
[0085] an audio similarity unit, configured to calculate, by using an audio similarity model, similarities between the template audio type and a plurality of audio types to be matched in an audio database, to obtain a plurality of audio similarities;
[0086] The audio type matching unit is configured to determine the to-be-matched audio type corresponding to the maximum similarity among the plurality of audio similarities as the audio type matching the template audio stream.
[0087] Optionally, the first virtual character video module 50 includes:
[0088] A character video stream unit, configured to perform character segmentation on the template video stream using a video segmentation technique to obtain a character video stream;
[0089] A character lip shape video stream unit is used to fuse the character video stream and the target lip shape video stream using a lip shape fusion algorithm to obtain a character lip shape video stream;
[0090] The adding unit is used to add a preset background image and the target audio stream to the character lip shape video stream to obtain the virtual character video.
[0091] Optionally, the adding unit includes:
[0092] A first head action unit, configured to parse the target audio stream using a head gesture algorithm to obtain a first head action;
[0093] a head movement updating unit, configured to update a second head movement in the character's lip shape video stream according to the first head movement, to obtain a target character's lip shape video stream;
[0094] The virtual character video unit is used to add the target audio stream and the background image to the target character lip shape video to obtain the virtual character video.
[0095] Optionally, after the lip shape video stream module 40, the following steps are included:
[0096] A face and lip shape video stream module is used to merge the target lip shape video stream and the target face video stream to obtain a face and lip shape video stream;
[0097] The second character video module is used to fuse the face and lip shape video with the template video stream and add the target audio stream to obtain a virtual character video.
[0098] For the specific definition of the virtual character video generation device, please refer to the definition of the virtual character video generation method above, and will not be repeated here. The various modules in the above-mentioned virtual character video generation device can be implemented in whole or in part by software, hardware, or a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of the above-mentioned modules.
[0099] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 4 As shown. The computer device includes a processor, memory, network interface, display screen and input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a readable storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer-readable instructions. The internal memory provides an environment for the operation of the operating system and computer-readable instructions in the readable storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer-readable instructions are executed by the processor, a method for generating a virtual character video is implemented. The readable storage medium provided in this embodiment includes a non-volatile readable storage medium and a volatile readable storage medium.
[0100] In one embodiment, a computer device is provided, comprising a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor. When the processor executes the computer-readable instructions, the following steps are implemented:
[0101] Splitting a template video stream and a template audio stream from a template video;
[0102] Performing facial region calibration on the template video stream using a face recognition model to generate a target face video stream; searching an audio database for an audio type that matches the template audio stream;
[0103] Performing speech synthesis on a preset text according to the audio type to generate a target audio stream;
[0104] Generate lip shape parameters from the audio type and the text, and process the lip shape parameters and the target face video stream through a lip shape generation model to obtain a target lip shape video stream;
[0105] The target lip shape video stream and the template video stream are merged, and the target audio stream is added to obtain a virtual character video.
[0106] In one embodiment, one or more computer-readable storage media storing computer-readable instructions are provided. The computer-readable storage media provided in this embodiment include non-volatile computer-readable storage media and volatile computer-readable storage media. The computer-readable storage media store computer-readable instructions that, when executed by one or more processors, implement the following steps:
[0107] Splitting a template video stream and a template audio stream from a template video;
[0108] Performing facial region calibration on the template video stream using a face recognition model to generate a target face video stream; searching an audio database for an audio type that matches the template audio stream;
[0109] Performing speech synthesis on a preset text according to the audio type to generate a target audio stream;
[0110] Generate lip shape parameters from the audio type and the text, and process the lip shape parameters and the target face video stream through a lip shape generation model to obtain a target lip shape video stream;
[0111] The target lip shape video stream and the template video stream are merged, and the target audio stream is added to obtain a virtual character video.
[0112] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through computer-readable instructions. The computer-readable instructions can be stored in a non-volatile readable storage medium or a volatile readable storage medium. When the computer-readable instructions are executed, they can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided by the present invention may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0113] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0114] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.
Claims
1. A method for generating a virtual character video, characterized in that: include: Splitting a template video stream and a template audio stream from a template video; Performing facial region calibration on the template video stream using a face recognition model to generate a target face video stream; Searching an audio database for an audio type that matches the template audio stream; Performing speech synthesis on a preset text according to the audio type to generate a target audio stream; generating lip shape parameters from the audio type and the text, changing the lip shape state in the target face video stream according to the lip shape parameters using a lip shape generation model, generating lip movements corresponding to the lip shape parameters and a timestamp corresponding to the target audio stream, and determining a target lip shape video stream based on the lip movement and the timestamp; Merging the target lip shape video stream and the template video stream, and adding the target audio stream to obtain a virtual character video; The step of fusing the target lip shape video stream and the template video stream and adding the target audio stream to obtain a virtual character video includes: Performing character segmentation on the template video stream using a video segmentation technology to obtain a character video stream; The character video stream and the target lip shape video stream are merged by a lip shape fusion algorithm to obtain a character lip shape video stream; Adding a preset background image and the target audio stream to the character lip shape video stream to obtain the virtual character video; The step of adding a preset background image and the target audio stream to the character lip shape video stream to obtain the virtual character video comprises: parsing the target audio stream using a head gesture algorithm to obtain a first head movement; According to the first head movement, updating the second head movement in the character's lip shape video stream to obtain a target character's lip shape video stream; The target audio stream and the background image are added to the target character lip shape video to obtain the virtual character video.
2. The method for generating a virtual character video according to claim 1, wherein: Before the template video stream and the template audio stream are segmented from the template video, the method includes: Obtain character feature information from the target video through a character recognition model; The target video is trimmed according to the character feature information to obtain the template video.
3. The method for generating a virtual character video according to claim 1, wherein: The face region calibration of the template video stream using a face recognition model to generate a target face video stream includes: Performing facial key point detection on the video frames in the template video stream using a face recognition model to obtain face regions in a number of the video frames; The plurality of face regions are segmented from the template video stream to generate the target face video stream.
4. The method for generating a virtual character video according to claim 1, wherein: The step of searching an audio database for an audio type that matches the template audio stream comprises: Parsing the template audio stream to obtain a template audio type of the template audio stream; Calculating the similarity between the template audio type and a plurality of audio types to be matched in the audio database using an audio similarity model to obtain a plurality of audio similarities; The to-be-matched audio type corresponding to the maximum similarity among the plurality of audio similarities is determined as the audio type matching the template audio stream.
5. The method for generating a virtual character video according to claim 1, wherein: Merging the target lip shape video stream and the target face video stream to obtain a face lip shape video stream; The human face lip shape video and the template video stream are fused, and the target audio stream is added to obtain a virtual character video.
6. A virtual character video generation device, characterized in that: include: A template video segmentation module is used to segment the template video into a template video stream and a template audio stream; The video and audio processing module is used to calibrate the face area of the template video stream through the face recognition model to generate a target face video stream; and search the audio database for an audio type that matches the template audio stream; A speech synthesis module, configured to perform speech synthesis on a preset text according to the audio type to generate a target audio stream; a lip shape video stream module, configured to generate lip shape parameters from the audio type and the text, modify the lip shape state in the target face video stream according to the lip shape parameters using a lip shape generation model, generate lip movements corresponding to the lip shape parameters and a timestamp corresponding to the target audio stream, and determine a target lip shape video stream based on the lip movements and the timestamp; a first virtual character video module, configured to fuse the target lip-sync video stream with the template video stream and add the target audio stream to obtain a virtual character video; The first virtual character video module includes: A character video stream unit, configured to perform character segmentation on the template video stream using a video segmentation technique to obtain a character video stream; A character lip shape video stream unit is used to fuse the character video stream and the target lip shape video stream using a lip shape fusion algorithm to obtain a character lip shape video stream; An adding unit, configured to add a preset background image and the target audio stream to the character lip shape video stream to obtain the virtual character video; The adding unit comprises: A first head action unit, configured to parse the target audio stream using a head gesture algorithm to obtain a first head action; a head movement updating unit, configured to update a second head movement in the character's lip shape video stream according to the first head movement, to obtain a target character's lip shape video stream; The virtual character video unit is used to add the target audio stream and the background image to the target character lip shape video to obtain the virtual character video.
7. A computer device comprising a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor, wherein: When the processor executes the computer-readable instructions, the method for generating a virtual character video according to any one of claims 1 to 5 is implemented.
8. One or more readable storage media storing computer-readable instructions, wherein when the computer-readable instructions are executed by one or more processors, the one or more processors execute the virtual character video generation method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Audio duplicate checking method and device
CN112241467A
Lip shape synchronization video generation method and device, equipment and storage medium
CN112562720A
Video synthesis method and device, equipment and storage medium
CN112866586A