Method for synthesizing video, electronic device and computer program product

By removing and splicing the limbs of virtual objects in the synthetic video, the problems of cumbersome and strict requirements in the synthesis process of virtual human interaction in the prior art are solved, and the effect of simplifying the production process, shortening the processing cycle and improving memory usage efficiency is achieved.

CN119323625BActive Publication Date: 2025-05-09IFLYTEK CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411866722.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-18
Publication Date
2025-05-09
Estimated Expiration
2044-12-18

AI Technical Summary

Technical Problem

The prior art when recording virtual human interaction actions in synthetic videos, the process is cumbersome and strict, resulting in a long processing cycle of synthetic videos. Especially when adding specified actions, it is time-consuming and labor-intensive to re-create the action video.

Method used

By obtaining the first video containing the virtual object and the second video containing the limb movement of the target virtual object, the target limb part in the first video is removed, the video to be synthesized is generated, and the target limb part of the second video is then spliced ​​to the missing place of the video to be synthesized, and the synthetic video is generated.

Benefits of technology

The production process of action videos is simplified, the production requirements and cycles are reduced, the processing cycle of synthetic videos is shortened, and the space usage of action videos is reduced, and the efficiency of hardware device memory is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119323625B_ABST
    Figure CN119323625B_ABST
Patent Text Reader

Abstract

The present application proposes a method for synthesizing videos, an electronic device, and a computer program product. The method for synthesizing videos includes: obtaining a first video containing a first virtual object and a second video containing a body movement of a target virtual object; for the first video, removing the target body part of the first virtual object in the target video segment to obtain a video to be synthesized, wherein the target video segment is a video segment corresponding to the action insertion period in the first video; based on the temporal correspondence between the second video and the target video segment, splicing the target body part of each video frame of the second video to the missing part of the target body part of each video frame of the video to be synthesized to generate a synthesized video. Since the second video only contains the target body part, the decoupling of the virtual object and the body movement can be achieved. When the virtual object is a virtual character, when making the second video / action video, there is no need to consider the dressing of the main body parts of the character, or even the identity of the character.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of video processing technology, and in particular to a method for synthesizing videos, an electronic device, and a computer program product. Background Art

[0002] With the development of science and technology, virtual reality technology has been widely used in various fields. Among them, the interactive action synthesis of virtual humans is an important part of virtual reality technology.

[0003] At present, most technical solutions for synthesizing the interactive actions of virtual people into videos. To generate a synthesized video of a virtual person doing a specified action, it is necessary to record the person's speech and silent video, and at the same time, it is necessary to record the action video of the same person doing the specified action, and then insert the action video as a video segment into the person video to generate the final synthesized video. Therefore, when playing the synthesized video, you can see the virtual person doing the specified action.

[0004] However, the process of recording action videos in the above technical solution is relatively cumbersome, and the production requirements are also relatively strict: the action video needs to be recorded with the same person, wearing the same clothes and the same dress. This makes the entire process of synthesizing videos very long. In particular, when adding additional specified actions to the synthesized video, the action video needs to be re-produced as required, which is time-consuming and laborious. Summary of the invention

[0005] Based on the above technical status, the present application proposes a method for synthesizing video, an electronic device and a computer program product to achieve the purpose of reducing the processing cycle of synthesized video in the process of synthesizing virtual character actions in the video.

[0006] According to a first aspect of the present application, a method for synthesizing videos is provided, the method comprising: obtaining a first video containing a first virtual object and a second video containing limb movements of a target virtual object, wherein the target virtual object includes the first virtual object or a second virtual object that belongs to the same biological type as the first virtual object; the second video only contains a target limb part of the target virtual object; for the first video, removing the target limb part of the first virtual object in a target video segment to obtain a video to be synthesized, wherein the target video segment is a video segment of the first video corresponding to the action insertion period; based on the temporal correspondence between the second video and the target video segment, splicing the target limb part of each video frame of the second video to the missing part of the target limb part of each video frame of the video to be synthesized to generate a synthesized video.

[0007] According to the second aspect of the present application, an electronic device is provided, comprising a memory and a processor; the memory is connected to the processor and is used to store programs; the processor is used to implement the method for synthesizing video as described in the first aspect by running the program in the memory.

[0008] According to a third aspect of the present application, a storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the method for synthesizing video as described in the first aspect is implemented.

[0009] According to a fourth aspect of the present application, a computer program product is provided, comprising: a computer program, wherein when the computer program is executed by a processor, the method for synthesizing video as described in the first aspect is implemented.

[0010] In the present application, after obtaining the first video of the first virtual object, the target limb part is removed therefrom, thereby generating a video to be synthesized. Then the video to be synthesized is spliced ​​with the second video containing the limb action, and the target limb part in the second video is spliced ​​to the missing part of the target limb part in the first video to generate a synthesized video. Since the second video only contains the target limb part, the decoupling of the virtual object and the limb action can be achieved. In the case where the virtual object is a virtual character, when making the second video / action video, there is no need to consider the dressing of the main body parts of the character (body parts other than the target limb part), and even the identity of the character. The production process of the action video is greatly simplified, the production requirements and the production cycle are reduced, and the processing cycle of the synthesized video is also greatly reduced. Further, compared with the action video containing the virtual character, the second video only contains the target limb part. Therefore, the space occupied by the second video will also be greatly reduced. When generating the synthesized video, more limb actions can be loaded at the same time without affecting the memory usage of the hardware device. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0012] Figure 1 One of the flowcharts of a method for synthesizing videos provided in an embodiment of the present application.

[0013] Figure 2 A schematic diagram of the process of constructing a body movement video library according to an embodiment of the present application.

[0014] Figure 3 The present invention provides a second flowchart of a method for synthesizing videos according to an embodiment of the present application.

[0015] Figure 4 A schematic diagram of the cutting edge of a target limb part provided in an embodiment of the present application.

[0016] Figure 5 The third flowchart of a method for synthesizing videos provided in an embodiment of the present application.

[0017] Figure 6 A schematic diagram of the structure of a device for synthesizing video provided in an embodiment of the present application.

[0018] Figure 7 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0019] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0020] Overview

[0021] As described in the background technology section, most technical solutions for synthesizing the interactive actions of virtual people into videos require pre-recording of action videos. For example, a popular virtual human interactive action synthesis solution only needs to provide the virtual person's speech and silent videos and action videos to produce a synthetic video of the virtual person reading aloud according to the specified content and making specified actions. In addition, the lip shape of the virtual character in the synthetic video is consistent with the content being read aloud. The implementation process can be, first, shooting a character video, including a certain length of speech and silent videos. Secondly, according to the established script, a specific action video of the character should be shot, and the video should be divided into independent action video segments. Then, the lip neural network is trained for the shot speech and silent videos to realize a neural network whose input is audio and output is a lip picture, and the trained neural network is used to generate a lip picture of the specified content. Finally, when synthesizing the video, if there is an action that needs to be synthesized, select the relevant action in the pre-made action video segment, splice it with the character video and the generated lip picture, and you can form a synthetic video of the virtual person.

[0022] However, the existing virtual human interactive action synthesis scheme mainly considers the synthesis effect of the lips. The action must be preset, and there are strict requirements on the identity and dress of the person shooting the action. This is because the action video, as a video segment of the synthesized video, needs to be inserted as a whole. Therefore, the person recording the action must be the same person as the person recording the person video, and the dress during shooting must also be the same. In particular, after synthesizing the video, if you want to add additional actions to the synthesized video, you need to shoot the action video again with the same person in the same dress. This is not only not flexible enough, but also brings great video production costs and cycles. In addition, in the process of using the action video, the action video that needs to be added / synthesized needs to be pre-loaded into the memory of the hardware device. Since the action video contains the entire person, its memory occupancy is very large, which will result in the number of action videos added at the same time being limited.

[0023] In view of the above technical status, the inventor of the present application found that the above problems are all caused by action videos, and are caused by inserting the action video as a video segment into the character video as a whole when synthesizing the video. In this regard, the inventor thought that the video can be processed by image splicing. When synthesizing the video, the limb parts directly related to the action are spliced ​​into the character video without limb parts, and a synthetic video with the same effect can be synthesized. Moreover, at this time, the action video only contains limb parts, which are decoupled from the character, and the requirements for the character will be greatly reduced, so that various problems surrounding the action video can be solved. Then the inventor proposed a method for synthesizing videos, which can remove the target limb parts from the first video of the first virtual object after obtaining the first video, thereby generating a video to be synthesized. Then the video to be synthesized is spliced ​​with the second video containing the limb action, and the target limb part in the second video is spliced ​​to the missing part of the target limb part in the first video to generate a synthetic video. Since the second video only contains the target limb part, the decoupling of the virtual object and the limb action can be achieved. When the virtual object is a virtual character, when making the second video / action video, there is no need to consider the dressing of the main body parts of the character (body parts other than the target limbs), and even the identity of the character. The production process of the action video is greatly simplified, the production requirements and the production cycle are reduced, and the processing cycle of the composite video is also greatly reduced. Furthermore, compared with the action video containing the virtual character, the second video only contains the target limbs. Therefore, the space occupied by the second video will also be greatly reduced. When generating the composite video, more body movements can be loaded at the same time without affecting the memory usage of the hardware device.

[0024] Exemplary Methods

[0025] See also Figure 1In an exemplary embodiment, a method for synthesizing videos is provided. The method provided in the present application can be applied to any electronic device with data processing capabilities, such as mobile phones, computers, servers, etc. It is worth noting that the virtual objects in the following embodiments of the present application include but are not limited to virtual characters, and may also include other virtual objects with limb parts, such as virtual animals. Therefore, in different application scenarios or business needs, first videos and second videos of different virtual objects can be prepared to generate synthetic videos of corresponding virtual objects. For example, when it is necessary to generate a synthetic video of a virtual person, the following virtual objects are all virtual characters, and a synthetic video of the virtual person can be generated in the manner provided in the following embodiments. When it is necessary to generate a synthetic video of a virtual animal, the following virtual objects are all virtual animals, and a synthetic video of the virtual animal can be generated in the manner provided in the following embodiments.

[0026] The method for synthesizing video may include:

[0027] S101: Acquire a first video including a first virtual object and a second video including a body movement of a target virtual object.

[0028] It should be noted that the first video can be a video generated by recording an object in the physical world, or it can be a video containing a first virtual object generated by other video generation means. In the case where the virtual object is a virtual character, the first virtual object can be a virtual character created corresponding to any character in the physical world. In some embodiments, a video of a certain character can be recorded, and the recorded video can be used as the first video. For example, a company needs to synthesize a promotional video for promotion, and requires the virtual character in the video to be the company's receptionist or the company's founder. Then, a video containing speaking and silence of the company's receptionist or the company's founder can be recorded by an image acquisition device, and the recorded video can be used as the first video. The first virtual object can also be a digital person created entirely by virtualization means and does not correspond to a character in the physical world. Accordingly, the first video can be generated by other video generation means.

[0029] In some embodiments, a video library may be pre-generated, and the video library may include multiple videos, and different videos may include different virtual objects. By receiving a selection operation of the user, the video selected by the user is used as the first video.

[0030] The second video only includes the target limb part of the target virtual object. Since the target limb part in the second video will be spliced ​​onto the first virtual object in the first video, in order to enhance the natural effect of the composite video, the target virtual object includes the first virtual object or the second virtual object of the same biological type as the first virtual object. The biological type includes: people, animals, wherein when the biological type is animals, animals of different species belong to different biological types. For example, if the first virtual object is a virtual person, then the second virtual object is also a virtual person. If the first virtual object is a virtual tiger, then the second virtual object is also a virtual tiger. In different business needs, the target limb part can be set to different limb parts. In some embodiments, the first virtual object and the second virtual object are both virtual people, and the target limb part includes: arms. Among them, the arms include hands, upper arms and lower arms. In some embodiments, the target limb part may include only hands. In some embodiments, the target limb part may also include legs. In some embodiments, the first virtual object and the second virtual object are both virtual animals, and the virtual animals have limbs, and the target limb part includes the limbs of the virtual animal. Regarding the limb movements in the second video, there is no limitation here, and it can be any limb movement.

[0031] It is understandable that, since the second video does not contain other parts of the target virtual object except the target limb part, if the second video is obtained by recording, it is not necessary to consider the dress of the main body parts of the person (body parts other than the target limb part), or even the identity of the person. For example, when the first video is a character video of person A, when making the second video, a video of the target limb part of person B can be recorded. It is sufficient that the target limb part of person B is similar to the target limb part of person A.

[0032] S102: For the first video, remove the target limb part of the first virtual object in the target video segment to obtain a video to be synthesized.

[0033] In this step, the target video segment is the video segment corresponding to the action insertion period in the first video. The action insertion period can be the period in the composite video where the body movement in the second video needs to be played. For example, if a welcome action needs to be played from the 10th second to the 15th second of the composite video, then the body movement in the second video is the welcome action, and the action insertion period is from the 10th second to the 15th second. In some embodiments, the user can freely select or input the action insertion period, and after receiving the user's operation, the electronic device determines the target video segment based on the user's selection or input of the action insertion period.

[0034] After the video to be synthesized is obtained, the first virtual object in the video segment corresponding to the action insertion period in the video to be synthesized will lack the target limb part, but the first virtual objects in the remaining periods all contain the target limb part. In some embodiments, each video frame in the target video segment can be processed to remove the target limb part.

[0035] S103: Based on the temporal correspondence between the second video and the target video segment, the target limb part of each video frame of the second video is spliced ​​to the missing part of the target limb part of each video frame of the video to be synthesized to generate a synthesized video.

[0036] It should be noted that, according to the time sequence correspondence between the second video and the target video segment, the video frame in the target video segment corresponding to each video frame in the second video can be determined. The correspondence between each frame in the target video segment and the video frame in the video to be synthesized can be determined by S102. Finally, the video frame in the video to be synthesized corresponding to each video frame in the second video can be determined. Each video frame in the second video is spliced ​​into its corresponding video frame in the video to be synthesized, and spliced ​​into the missing part of the target limb part. By playing the synthesized video, it can be seen that the limb movements in the second video will be played during the action insertion period. In some embodiments, relevant audio data can also be inserted into the synthesized video, so that the relevant audio data can be played while displaying the image data. For example, the image data in the synthesized video can be the image of a tour guide at a certain scenic spot. During the playback of the synthesized video, the tour guide image will make the specified limb movements and will also give a voice introduction to the scenic spot.

[0037] In some embodiments, the duration of the second video is greater than or equal to the duration of the target video segment. When the duration of the second video is equal to the duration of the target video segment, the video frames of the two correspond one to one. When the duration of the second video is greater than the duration of the target video segment, one frame can be extracted every specified number of video frames by extracting frames, and the extracted video frames correspond one to one with the video frames in the target video segment.

[0038] It is worth noting that the number of the second videos can be one or more. The number is equal to the number of action insertion periods, and the two correspond one to one. In the case where the number of the second videos is multiple, when generating the composite video, each second video is spliced ​​with the video frame of the corresponding action insertion period.

[0039] In an embodiment of the present application, after obtaining the first video of the first virtual object, the target limb part thereof is removed to generate a video to be synthesized. Then the video to be synthesized is spliced ​​with the second video containing the limb action, and the target limb part in the second video is spliced ​​to the missing part of the target limb part in the first video to generate a synthesized video. Since the second video only contains the target limb part, the decoupling of the virtual object and the limb action can be achieved. In the case where the virtual object is a virtual character, when making the second video / action video, there is no need to consider the dressing of the main body parts of the character (body parts other than the target limb part), and even the identity of the character. The production process of the action video is greatly simplified, the production requirements and the production cycle are reduced, and the processing cycle of the synthesized video is also greatly reduced. Further, compared with the action video containing the virtual character, the second video only contains the target limb part. Therefore, the space occupied by the second video will also be greatly reduced. When generating the synthesized video, more limb actions can be loaded at the same time without affecting the memory usage of the hardware device.

[0040] In order to effectively reduce the duplication of action videos and increase the number of actions available, in some embodiments of the present application, the step of obtaining a second video containing the limb movements of the target virtual object includes:

[0041] In response to the first input, a body movement video is selected from a body movement video library; wherein the body movement video library includes: body movement videos corresponding to different body movements; and the selected body movement video is determined as the second video.

[0042] It should be noted that the body movement video library includes multiple body movement videos, each of which includes one or more body movements. Each body movement video also only includes the target body part. In some embodiments, each body movement video only includes one body movement. In this way, the user can freely select each body movement that needs to appear in the composite video according to actual needs. For example, the user needs to have a hand-raising action in the first period of the composite video, and needs to have a welcoming action in the second period. Then, the user can select a body movement video including a hand-raising action and a body movement video including a welcoming action as the second video from the body movement video library through the first input. In some other embodiments, each body movement video can include at least two body movements. In this way, the user selects fewer body movement videos, so that more body movements can appear in the composite video. For example, body movement video A includes a welcoming action and a handshake action, and the user needs to have a welcoming action and a handshake action in the first period of the composite video. Then, the user can select body movement video A as the second video from the body movement video library through the first input.

[0043] It is understandable that the advantages of building a body movement video library are very obvious, and the usage is also very flexible. It can support a wide range of body movements, and the attributes of body movements can be continuously enriched. There is no limitation on the process of building a body movement video library. Assuming the target body part is the arm or upper limb, the process of building a body movement video library can be as follows: Figure 2 As shown, it includes: S201, you can first shoot a video, shoot some upper limb movement videos of some characters in advance, the upper limb movement needs to include the hands and arms, and the hand movement cannot cover the head. Taking into account the diversity of body movements, the upper limb movements of people of different ages and genders can be shot. Optionally, the clothes worn by the shooter can be limited, for example, only solid color clothes or sleeveless clothes can be worn. S202, process the upper limb movement video shot in the first step: first, cut out each frame of the video, and cut out all the arms and palms, where the left hand and the right hand are separated, which can reduce the resource occupation of each action. S203, perform video encoding operation on the video processed in the second step, which can effectively reduce the machine resource occupation of the video. At the same time, the encoded and compressed video resources are binary encrypted and packaged, and processed into data that can be quickly processed by the machine.

[0044] In an embodiment of the present application, by creating a library of body movement videos and selecting the required second video therefrom, the repeated production of action videos can be effectively reduced, the number of actions available for selection can be increased, and the first virtual object can perform more types of body movements, thereby improving the realism and richness of the interaction.

[0045] To facilitate the user to select the body movements that need to appear in the synthesized video, in some embodiments of the present application, each body movement video in the body movement video library is provided with a plurality of movement tags;

[0046] In response to the first input, a body movement video is selected from a body movement video library, including:

[0047] In response to the first input, based on the body movement insertion requirement input by the user, an action tag of the body movement to be inserted is determined; the action tags of the body movement videos in the body movement video library are matched based on the action tags of the body movement to be inserted; and the body movement videos in the body movement video library that are successfully matched are determined as the selected body movement videos.

[0048] It should be noted that the action tag may represent the relevant information of the body action contained in the body action video, so that different body action videos can be distinguished through the relevant information and the specific content of the body action can be determined. For example, the relevant information may include the action type, action duration, etc.

[0049] The action insertion requirement can indicate the action tag of the body movement video selected by the user, so that the user can input the body movement insertion requirement according to the body movement video actually needed by the user. In some embodiments, the body movement insertion requirement includes multiple action tags. For example, the body movement that needs to appear in the synthetic video is a hand-raising action, and the duration is 2 seconds. Then, the user can directly input "hand-raising action" and "2 seconds" so that the body movement video with the action tags of "hand-raising action" and "2 seconds" can be selected from the body movement video library based on "hand-raising action" and "2 seconds" in the future. For another example, the action tags of all body movement videos in the body movement video library can be directly displayed, and the user determines the required body movement video through the action tag, and then inputs the required body movement video, and the electronic device determines the body movement video targeted by the user as the selected body movement video.

[0050] In some embodiments, if the action tags of a certain body movement video in the body movement video library are all the same as the action tags of the body movement to be inserted, then the body movement video can be regarded as a successfully matched body movement video in the body movement video library.

[0051] In an embodiment of the present application, each limb movement video in the limb movement video library is differentiated in different dimensions through multiple action tags, so that the user can more finely select the limb movements that need to appear in the synthesized video.

[0052] In some embodiments of the present application, the plurality of action tags include: action type, action duration, and at least one character attribute, and the character attribute includes: at least one of age, gender, and clothing.

[0053] It should be noted that the more action tags there are, the more refined the division of body movements is, and the more natural or realistic the effect of the synthesized video is. In some embodiments, the character attribute may also include skin color. For the action tag, it may also include a speed tag indicating the speed of the action.

[0054] It is understandable that in the process of building a body movement video library, an action tag for each body movement video is generated at the same time. For example, in the second step of the above example of building a body movement video library, an action tag can be added to the body movement video. The form of the action tag can be as follows:

[0055] {

[0056] Age: 20,

[0057] Gender: Female,

[0058] Name: Raise your hand,

[0059] Duration:

[0060] {

[0061] Pre-silence duration: 5 frames,

[0062] Action duration: 50 frames,

[0063] Post-silence duration: 5 frames

[0064] }

[0065] Left hand: Yes,

[0066] other:

[0067] }

[0068] In some embodiments, when adding action tags to body movement videos, action tags can be added to each body movement video manually. In other embodiments, body movements in body movement videos can also be identified based on image recognition technology and action tags can be added.

[0069] In an embodiment of the present application, the body movement video contains at least one action label related to the character attributes, and the body movements are divided more finely based on the character attributes, so that the effect of the synthesized video is more natural and realistic.

[0070] In order to improve the success rate of generating a synthetic video, in some embodiments of the present application, matching the action tags of the body action videos in the body action video library based on the action tags of the body action to be inserted includes:

[0071] Based on the action type and action duration of the limb action to be inserted, the action type and action duration of the limb action videos in the limb action video library are matched to determine an intermediate limb action video with the same action type and the same action duration; based on the character attributes of the limb action to be inserted, the character attributes of the intermediate limb action video are matched, and if no intermediate limb action video with the same character attributes is matched, prompt information associated with the intermediate limb action video is displayed; in response to a second input made by the user based on the prompt information, the intermediate limb action video selected by the user is determined as the successfully matched limb action video.

[0072] It should be noted that the action type and action duration of the body movements that need to appear in the synthesized video are generally determined before executing S101, and the body movement video that is usually matched needs to meet these two requirements. Therefore, the action type and action duration are the basis for matching the body movement video. On this basis, the more the number of matched action tags is, the more natural and realistic the effect of the synthesized video is. If all the action tags of the body movements to be inserted are the same as the action tags of a certain body movement video, then the effect of the synthesized video generated based on this body movement video is the most natural and realistic.

[0073] However, there are usually multiple action tags related to character attributes, and in some cases, it is difficult to match character attributes and complete the same body movement video. At this time, if the matching failure is returned directly, the composite video cannot be generated. If any intermediate body movement video is directly used as the body movement video that matches successfully, the effect of the composite video cannot be guaranteed. To this end, the embodiment of the present application gives the user the right to choose, and by displaying prompt information associated with the intermediate body movement video, the user decides whether to continue to generate the composite video and decides which intermediate body movement video to use as the body movement video that matches successfully. In some embodiments, the prompt information associated with the intermediate body movement video can be a text description, thumbnail, action label, etc. of the video content of the intermediate body movement video. The second input includes: click, slide, long press, etc., but is not limited to this.

[0074] In an embodiment of the present application, when a limb movement video with all identical action tags cannot be matched, prompt information associated with the intermediate limb movement video can be displayed, leaving it to the user to decide whether to continue generating a composite video and which intermediate limb movement video to use as the successfully matched limb movement video, thereby improving the success rate of generating a composite video.

[0075] In the case where a synthesized video outputting a specified semantic content is required, in order to further improve the efficiency of generating a synthesized video and reduce manual operations, in some embodiments of the present application, in response to a first input, selecting a body movement video from a body movement video library includes:

[0076] In response to the first input, the body movement to be inserted is determined based on the semantic content of the target text or target voice input by the user; wherein the body movement to be inserted is a body movement that matches the semantic content; and a body movement video containing the body movement to be inserted is selected from a body movement video library.

[0077] It should be noted that in the scenario where the synthesized video needs to output the specified semantic content, the user is usually required to input a text or voice, and the semantic content of the text or voice is the specified semantic content. When the synthesized video is finally generated, the specified semantic content can be output in the form of voice at the same time when it is played. In addition, the body movements in the synthesized video are consistent with the specified semantic content. For example, when the synthesized video is used as a welcome video at the company's front desk, the user can input a text whose semantics include welcoming customers and introducing the company to customers. By performing semantic analysis on the text, the semantic content related to welcome and introduction can be determined, and the time period when these semantic contents appear can be determined. Then, a body movement video containing a welcome action and a body movement video containing an introduction action can be selected from the body movement video library as the second video. Among them, the time period when these semantic contents determined above appear can be used as an action insertion time period.

[0078] In some embodiments, when determining the body movement to be inserted, the semantics of the target text or target speech can be analyzed and recognized through a large language model or a speech recognition model, and the action label of the body movement to be inserted can be output. The output action label is then used to match the action label of the body movement video in the body movement video library; then the successfully matched body movement video in the body movement video library is determined as the selected body movement video, wherein the action label and the matching process using the action label are the same as the description related to the action label in the above embodiments, and will not be repeated here. For example, Figure 3As shown, the generation of a synthetic video of a virtual character is used as an example for explanation. A limb movement library can be pre-built, which is equivalent to the limb movement video library in the above embodiments, and will not be described in detail here. First, execute S301 to record a character video. According to the video recording specifications, a few minutes of speaking video can be recorded. For example, the video can be recorded according to actual needs, and can include the head, upper body, hands and arms, and the hand movements cannot cover the face. S302, the character video recorded in the previous step is trained using AI technology to train a neural network model whose input is audio and output is lip video / picture. S303, text input. S304, text processing, including converting the text into speech data and predicting the action of the text content. Among them, the process of action prediction is the same as the process of generating action labels based on the semantic content of the target text or target speech. S305, using the trained neural network model, synthesized speech data and predicted action labels, to generate a virtual human interaction video. Among them, the lip video can be obtained through the neural network model and the voice data, the body action video to be inserted can be obtained through the retrieval of the action tag, and the virtual human interaction video can be obtained by splicing the body action video and the lip video into the character video (the character video after removing the target body part), which is equivalent to the synthetic video in the above embodiments. S306, video output, output the virtual human interaction video as the processing result.

[0079] In an embodiment of the present application, when it is necessary to output a synthesized video with specified semantic content, the body movements to be inserted can be determined through the semantic content, which can improve the efficiency of generating the synthesized video and reduce manual operations.

[0080] In some embodiments of the present application, the step of obtaining a second video containing a body movement of the target virtual object includes:

[0081] Obtain an initial limb movement video of a target virtual object recorded in the same or similar clothing as the first virtual object; for the initial limb movement video, remove the target area in each video frame to obtain a second video; wherein the target area includes all areas except the target limb part of the target virtual object.

[0082] It should be noted that the process of making the second video can be similar to the process of making the body movement video when constructing the body movement video library in the above embodiment. In this embodiment, in order to improve the video quality of the body movement video, the body movements under the same or similar clothing can be recorded to generate an initial body movement video. Among them, the higher the video quality, the more natural and realistic the body movements in the synthesized video. For example, the first virtual object is character A, and the dress of character A in the first video is the target style. Then, when making the second video, the body movements of character A in the target style can be recorded. The body movements of character B in the target style can also be recorded. The recorded video is used as the initial body movement video, and needs to be subsequently processed to obtain the second video.

[0083] It is understandable that the body movements in the initial body movement video are related to specific business requirements. If the first body movement is required to appear in the synthesized video, then the initial body movement video needs to include the first body movement. When recording the initial body movement video, the relevant person needs to perform the first body movement.

[0084] In some embodiments, when the target limb part is a symmetrical body part of a person, the second video includes two parts, each part including one body part. For example, when the target body part is a hand, the second video includes a first part including only the left hand and a second part including only the right hand.

[0085] In an embodiment of the present application, the body movements of the target virtual object can be directly recorded, and then the second video can be obtained by removing the redundant parts.

[0086] In order to improve the display effect of the joints of the target limb parts in the composite video and reduce the difficulty of stitching, in some embodiments of the present application, the cutting edge of the target limb part is a horizontal or vertical straight line.

[0087] It should be noted that when splicing video frames, if the cutting edge of the target limb part is a horizontal or vertical straight line, the difficulty of splicing will be greatly reduced, and the display effect of the splicing will be improved. If the cutting edge of the target limb part is a diagonal line, the cutting edge will have a boundary burr problem. When displayed on a larger display screen, the boundary burr will be very obvious, affecting the viewing experience.

[0088] For example, Figure 4As shown, when the virtual object is a virtual character, a vertical cut can be made along the shoulder joint, and the cutting edge 401 is a vertical branch line. In the process of making the second video, when the target limb part is cut out from the video containing the virtual character, a vertical cut can be made along the shoulder joint to separate the target limb part from the body of the virtual character. When executing the above S102, a vertical cut can be made along the shoulder joint to separate the target limb part from the body of the virtual character, thereby removing the target limb part.

[0089] In an embodiment of the present application, by setting the cutting edge as a horizontal or vertical branch line, the display effect of the joints of the target limb parts in the composite video can be improved, while at the same time, the difficulty of splicing can be reduced.

[0090] In order to make the effect of the synthesized video more natural and realistic, in some embodiments of the present application, based on the temporal correspondence between the second video and the target video segment, the target limb part of each video frame of the second video is spliced ​​to the missing part of the target limb part of each video frame of the video to be synthesized, and after the synthesized video is generated, the method further includes:

[0091] Reduce the skin color difference and / or clothing color difference between the first stitching area and the second stitching area in each target video frame, wherein the target video frame is a video frame formed by stitching the target limb part in the synthesized video, the first stitching area is the area corresponding to the target limb part in the second video, and the second stitching area is the area corresponding to the first virtual object in the video to be synthesized.

[0092] It should be noted that since the target video frame in the composite video is spliced ​​from two video frames, the target limb parts and main body parts may come from different people, and considering the production process of the second video, the target video frame may involve problems such as clothing color difference and skin color difference.

[0093] To solve these possible problems, after generating the composite video, if these problems are detected, special effects processing means can be used to reduce the color difference of skin color and / or clothing color difference. Among them, the special effects processing means include but are not limited to whitening, sharpening, and feathering. In some embodiments, the tool for adjusting color difference in the open source tool library can be directly used to process the composite video.

[0094] In the embodiment of the present application, by adjusting the color difference of skin color and clothing color difference, the effect of the synthesized video can be made more natural and realistic.

[0095] When the virtual object is a virtual character and a synthesized video is required to output specified content, in order to make the mouth shape of the first virtual character in the synthesized video consistent with the specified output content, in some embodiments of the present application, obtaining a first video containing the first virtual object includes:

[0096] An initial character video of a first virtual character and a lip image during reading a target text are obtained; for the initial character video, the lip image is spliced ​​to the lip part of each video frame in the initial character video to obtain a first video.

[0097] It should be noted that the mouth shape of the first virtual character in the initial character video is usually inconsistent with the mouth shape during the process of reading the target text. In this embodiment, the mouth shape during the process of reading the target text can be separately produced based on the target text. For example, an AI lip generation model can be pre-trained and generated. Then the target text is converted into audio data, and the audio data is input into the AI ​​lip generation model to obtain the mouth shape during the process of reading the target text, that is, the lip image. Among them, the AI ​​lip generation model can be trained by a video corresponding to the relevant virtual person shot in advance, and the generation effect of the lips and teeth is usually better. In some embodiments, the audio data generated by any user reading the target text can also be directly used, and the audio data is input into the AI ​​lip generation model to obtain the mouth shape during the process of reading the target text. The splicing process of the lip image will not be described in detail here.

[0098] In some embodiments, while keeping the output content consistent with the lip shape in the composite video, suppose a composite video of 1 minute is to be generated, and an action needs to be inserted at the 20th second, and the action duration is 5 seconds. The general steps are as follows: First, use the complete character video and lip image to generate the first 20 seconds of video. Second, use the video with missing hands and lip image to generate the middle 5 seconds of video, which is the part between the 20th and 25th seconds of the entire video. Third, splice the video generated in the second step with the retrieved action video. Fourth, use the complete character video and lip image to generate the video from the 25th second to 1 minute, and this part of the complete character video uses the video after the 25th second in the original video.

[0099] In an embodiment of the present application, when the virtual object is a virtual character and a synthesized video output is required to specify content, the lip shape of the first virtual character in the synthesized video can be made consistent with the specified output content by splicing lip images.

[0100] In order to avoid the influence of the splicing of the target body parts on the lip image, resulting in mismatch between the mouth shape and the sound, in some embodiments of the present application, the first splicing area and the third splicing area of ​​each target video frame do not overlap, wherein the target video frame is a video frame formed by splicing the target body parts in the synthetic video, the first splicing area is the area corresponding to the target body part in the second video, and the third splicing area is the area corresponding to the lip image.

[0101] It should be noted that the target limbs and lips in the target video frame are spliced ​​together. If the two areas overlap, they will affect each other, causing the mouth shape and sound to be mismatched or out of sync. For this reason, when making the second video, there are certain requirements for the body movements of the target limbs. When making body movements, the target limbs cannot block the lips. For example, in the above embodiment, when recording the initial body movement video, the recorded object cannot make body movements such as covering the mouth or passing the arm through the lips.

[0102] In the embodiment of the present application, the areas corresponding to the target limb parts and the lip parts in the target video frame do not overlap, thereby avoiding the influence of the splicing of the target limb parts on the lip image, causing the mouth shape and sound to not match.

[0103] For ease of understanding, the method provided by the embodiment of the present application is described in detail below by taking the generation of a synthetic video of a virtual character as an example. Assume that a first body movement occurs in a first period of time of the synthetic video, and a second body movement occurs in a second period of time. And the synthetic video needs to output the content of the target text in the form of voice, and the mouth shape is consistent with the output content while the voice is output.

[0104] A limb movement library and an AI model may be prepared in advance. The limb movement library is equivalent to the limb movement video library in the above embodiment, and the AI ​​model is equivalent to the AI ​​lip generation model, which will not be described in detail here.

[0105] like Figure 5 As shown, first, S501 is executed to generate a lip video. The lip shape data of reading the target text is obtained by using the user audio and the AI ​​model. The user audio can be audio data converted from the target text.

[0106] S502, lip part stitching. The character video and the lip video are stitched together to generate a lip stitching video. The lip shape of the virtual character in the lip stitching video is the lip shape of reading the target text. The character video is a pre-recorded video of the virtual character, and the target body parts have been removed in the first and second time periods of the character video.

[0107] S503, retrieve the body movement video from the body movement library based on the action label. Among them, the action label can be input by the user himself; it can also be generated by inputting the target text into the large language model, which will not be described in detail here. It can be understood that the incoming action label will carry information about which action needs to be inserted in which frame. In this way, it is possible to determine in which frame to insert the action, and determine in which frame to end it based on the label of the action duration. Here, a first body movement video containing the first body movement and a second body movement video containing the second body movement will be retrieved. If the body movement video is in an encoded and encrypted state, it needs to be decoded and decrypted first.

[0108] S504, body movement splicing, splicing the retrieved first body movement video to the lip splicing video according to the first time period, and splicing the second body movement video to the lip splicing video according to the second time period, to generate a composite video.

[0109] S505, special effect processing, eliminating clothing color difference and skin color difference. This process can refer to the above embodiment of reducing clothing color difference and skin color difference, which will not be repeated here.

[0110] S506, obtaining a virtual human video. During the playing of the virtual human video, the virtual human's mouth shape is consistent with the output audio content, and the virtual human makes a first body movement in a first period and a second body movement in a second period.

[0111] Exemplary Devices

[0112] Accordingly, the embodiment of the present application also provides a device for synthesizing video, see Figure 6 As shown, the device comprises:

[0113] The video acquisition module 601 is used to acquire a first video including a first virtual object and a second video including a body movement of a target virtual object, wherein the target virtual object includes the first virtual object or a second virtual object of the same biological type as the first virtual object; and the second video only includes a target body part of the target virtual object;

[0114] The video processing module 602 is used to remove the target limb part of the first virtual object in the target video segment of the first video to obtain a video to be synthesized, wherein the target video segment is a video segment corresponding to the action insertion period in the first video;

[0115] The video generation module 603 is used to splice the target limb part of each video frame of the second video to the missing part of the target limb part of each video frame of the video to be synthesized based on the temporal correspondence between the second video and the target video segment, so as to generate a synthesized video.

[0116] In some embodiments of the present application, the video acquisition module 601, the response unit, is used to select a limb movement video from a limb movement video library in response to a first input; wherein the limb movement video library includes: limb movement videos corresponding to different limb movements; the determination unit is used to determine the selected limb movement video as the second video.

[0117] In some embodiments of the present application, each body movement video in the body movement video library is provided with a plurality of movement tags;

[0118] The response unit is specifically used to respond to the first input, determine the action tag of the body movement to be inserted based on the body movement insertion demand input by the user; match the action tag of the body movement video in the body movement video library based on the action tag of the body movement to be inserted; and determine the body movement video that has been successfully matched in the body movement video library as the selected body movement video.

[0119] In some embodiments of the present application, the plurality of action tags include: action type, action duration, and at least one character attribute, and the character attribute includes: at least one of age, gender, and clothing.

[0120] In some embodiments of the present application, the response unit is specifically used to match the action type and action duration of the limb movement videos in the limb movement video library based on the action type and action duration of the limb movement to be inserted, and determine the intermediate limb movement video with the same action type and the same action duration; match the character attributes of the intermediate limb movement video based on the character attributes of the limb movement to be inserted, and if no intermediate limb movement video with the same character attributes is matched, display prompt information associated with the intermediate limb movement video; in response to a second input made by the user based on the prompt information, determine the intermediate limb movement video selected by the user as the successfully matched limb movement video.

[0121] In some embodiments of the present application, the response unit is specifically used to determine the body movement to be inserted in response to the first input based on the semantic content of the target text or target voice input by the user; wherein the body movement to be inserted is a body movement that matches the semantic content; and a body movement video containing the body movement to be inserted is selected from the body movement video library.

[0122] In some embodiments of the present application, the step of the video acquisition module 601 acquiring the second video containing the body movement of the target virtual object includes:

[0123] Obtain an initial limb movement video of a target virtual object recorded in the same or similar clothing as the first virtual object; for the initial limb movement video, remove the target area in each video frame to obtain a second video; wherein the target area includes all areas except the target limb part of the target virtual object.

[0124] In some embodiments of the present application, the cutting edge of the target limb part is a horizontal or vertical straight line.

[0125] In some embodiments of the present application, the device also includes: a color difference processing module, used to reduce the skin color difference and / or clothing color difference between the first stitching area and the second stitching area in each target video frame, wherein the target video frame is a video frame formed by stitching the target limb parts in the synthesized video, the first stitching area is the area corresponding to the target limb part in the second video, and the second stitching area is the area corresponding to the first virtual object in the video to be synthesized.

[0126] In some embodiments of the present application, the first virtual object and the second virtual object are both virtual characters, and the target limb parts include: arms.

[0127] In some embodiments of the present application, the video acquisition module 601 acquires a first video containing a first virtual object, including: acquiring an initial character video of the first virtual character and a lip image during the process of reading a target text; for the initial character video, splicing the lip image to the lip part of each video frame in the initial character video to obtain a first video.

[0128] In some embodiments of the present application, the first stitching area and the third stitching area of ​​each target video frame do not overlap, wherein the target video frame is a video frame formed by stitching target limb parts in the synthetic video, the first stitching area is the area corresponding to the target limb part in the second video, and the third stitching area is the area corresponding to the lip image.

[0129] The apparatus for synthesizing video provided in this embodiment belongs to the same application concept as the method for synthesizing video provided in the above embodiments of this application, and can execute the method for synthesizing video provided in any of the above embodiments of this application, and has the corresponding functional modules and beneficial effects of the execution method. For the technical details not described in detail in this embodiment, please refer to the specific processing content of the method for synthesizing video provided in the above embodiments of this application, which will not be repeated here.

[0130] It should be understood that the modules in the above devices can be implemented in the form of a processor calling software. For example, the device includes a processor, the processor is connected to a memory, and instructions are stored in the memory. The processor calls the instructions stored in the memory to implement any of the above methods or realize the functions of each unit of the device, wherein the processor can be a general-purpose processor, such as a CPU or a microprocessor, etc., and the memory can be a memory in the device or a memory outside the device. Alternatively, the unit in the device can be implemented in the form of a hardware circuit, and the functions of some or all units can be realized by designing the hardware circuit. The hardware circuit can be understood as one or more processors; for example, in one implementation, the hardware circuit is an ASIC, and the functions of some or all of the above units are realized by designing the logical relationship of the components in the circuit; for another example, in another implementation, the hardware circuit can be implemented by PLD, taking FPGA as an example, which can include a large number of logic gate circuits, and the connection relationship between the logic gate circuits is configured by a configuration file, so as to realize the functions of some or all of the above units. All units of the above devices can be implemented in the form of a processor calling software, or in the form of a hardware circuit, or in part by a processor calling software, and the remaining part is implemented in the form of a hardware circuit.

[0131] In an embodiment of the present application, a processor is a circuit with the ability to process signals. In one implementation, the processor may be a circuit with the ability to read and run instructions, such as a CPU, a microprocessor, a GPU, or a DSP; in another implementation, the processor may implement certain functions through the logical relationship of a hardware circuit, and the logical relationship of the hardware circuit is fixed or reconfigurable, such as a hardware circuit implemented by an ASIC or PLD, such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration document to implement the hardware circuit configuration can be understood as the process of the processor loading instructions to implement the functions of some or all of the above units. In addition, it can also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as an NPU, TPU, DPU, etc.

[0132] It can be seen that each unit in the above device can be one or more processors (or processing circuits) configured to implement the above method, such as: CPU, GPU, NPU, TPU, DPU, microprocessor, DSP, ASIC, FPGA, or a combination of at least two of these processor forms.

[0133] In addition, all or part of the units in the above device can be integrated together, or can be implemented independently. In one implementation, these units are integrated together and implemented in the form of a SOC. The SOC may include at least one processor for implementing any of the above methods or implementing the functions of each unit of the device. The type of the at least one processor may be different, for example, including a CPU and an FPGA, a CPU and an artificial intelligence processor, a CPU and a GPU, etc.

[0134] Exemplary Electronic Devices

[0135] The present application embodiment provides an electronic device, see Figure 7 As shown, the device includes:

[0136] A memory 700 and a processor 710; wherein the memory 700 is connected to the processor 710 and is used to store programs; the processor 710 is used to implement the method for synthesizing video disclosed in any of the above embodiments by running the program stored in the memory 700.

[0137] Specifically, the electronic device may further include: a bus, a communication interface 720 , an input device 730 and an output device 740 .

[0138] The processor 710, the memory 700, the communication interface 720, the input device 730 and the output device 740 are connected to each other via a bus.

[0139] A bus may include a pathway that transfers information between components of a computer system.

[0140] The processor 710 may be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the scheme of the present invention. It may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0141] The processor 710 may include a main processor, and may also include a baseband chip, a modem, and the like.

[0142] The memory 700 stores a program for executing the technical solution of the present invention, and may also store an operating system and other key services. Specifically, the program may include a program code, and the program code includes a computer operation instruction. More specifically, the memory 700 may include a read-only memory (ROM), other types of static storage devices that can store static information and instructions, a random access memory (RAM), other types of dynamic storage devices that can store information and instructions, a disk storage, a flash, and the like.

[0143] The input device 730 may include a device for receiving data and information input by a user, such as a keyboard, a mouse, a camera, a scanner, a light pen, a voice input device, a touch screen, a pedometer, or a gravity sensor.

[0144] Output device 740 may include devices that allow information to be output to a user, such as a display screen, printer, speaker, etc.

[0145] The communication interface 720 may include any transceiver or the like to communicate with other devices or communication networks, such as Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc.

[0146] The processor 710 executes the program stored in the memory 700 and calls other devices, which can be used to implement each step of any method for synthesizing video provided in the above embodiments of the present application.

[0147] An embodiment of the present application also proposes a chip, which includes a processor and a data interface. The processor reads and runs a program stored in a memory through the data interface to execute the method for synthesizing video introduced in any of the above embodiments. The specific processing process and its beneficial effects can be found in the embodiment introduction of the above-mentioned method for synthesizing video.

[0148] Exemplary computer program products and storage media

[0149] In addition to the above-mentioned methods and devices, an embodiment of the present application may also be a computer program product, which includes computer program instructions, which, when executed by a processor, enable the processor to execute the steps in the method of synthesizing video according to various embodiments of the present application described in any of the above embodiments of this specification.

[0150] The computer program product may be written in any combination of one or more programming languages ​​to write program codes for performing the operations of the embodiments of the present application, including object-oriented programming languages, such as Java, C++, etc., and conventional procedural programming languages, such as "C" language or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as an independent software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0151] In addition, an embodiment of the present application may also be a storage medium on which a computer program is stored, and the computer program is executed by a processor to execute the steps of the method for synthesizing video according to various embodiments of the present application described in any of the above embodiments of this specification.

[0152] For the aforementioned method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the order of the actions described, because according to the present application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.

[0153] It should be noted that each embodiment in this specification is described in a progressive manner, and each embodiment focuses on the differences from other embodiments, and the same or similar parts between the embodiments can be referred to each other. For the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0154] The steps in the methods of each embodiment of the present application can be adjusted in order, combined and deleted according to actual needs, and the technical features recorded in each embodiment can be replaced or combined.

[0155] The modules and sub-modules in the devices and terminals of the various embodiments of the present application can be combined, divided and deleted according to actual needs.

[0156] In the several embodiments provided in the present application, it should be understood that the disclosed terminals, devices and methods can be implemented in other ways. For example, the terminal embodiments described above are only schematic, for example, the division of modules or submodules is only a logical function division, and there may be other division methods in actual implementation, for example, multiple submodules or modules can be combined or integrated into another module, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or modules, which can be electrical, mechanical or other forms.

[0157] The modules or submodules described as separate components may or may not be physically separated, and the components of the modules or submodules may or may not be physical modules or submodules, that is, they may be located in one place, or they may be distributed on multiple network modules or submodules. Some or all of the modules or submodules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0158] In addition, each functional module or submodule in each embodiment of the present application may be integrated into one processing module, or each module or submodule may exist physically separately, or two or more modules or submodules may be integrated into one module. The above-mentioned integrated modules or submodules may be implemented in the form of hardware or in the form of software functional modules or submodules.

[0159] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0160] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly by hardware, software units executed by a processor, or a combination of the two. The software units may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0161] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.

[0162] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for synthesizing video, characterized in that: The method comprises: Acquire a first video containing a first virtual object and a second video containing a limb movement of a target virtual object, wherein the target virtual object includes the first virtual object or a second virtual object of the same biological type as the first virtual object; and the second video contains only a target limb part of the target virtual object; For the first video, remove the target limb part of the first virtual object in the target video segment to obtain a video to be synthesized, wherein the target video segment is a video segment corresponding to the action insertion period in the first video; the target video segment is determined based on the action insertion period selected or input by a user; Based on the temporal correspondence between the second video and the target video segment, the target limb part of each video frame of the second video is spliced ​​to the missing part of the target limb part of each video frame of the video to be synthesized to generate a synthesized video.

2. The method according to claim 1, characterized in that: The step of obtaining a second video containing the body movement of the target virtual object includes: In response to the first input, a body movement video is selected from a body movement video library; wherein the body movement video library includes: body movement videos corresponding to different body movements; The selected body movement video is determined as the second video.

3. The method according to claim 2, characterized in that Each body movement video in the body movement video library is provided with a plurality of movement tags; In response to the first input, a body movement video is selected from a body movement video library, including: In response to the first input, based on the physical action insertion requirement input by the user, determining an action label of the physical action to be inserted; Matching action tags of body movement videos in the body movement video library based on the action tags of the body movement to be inserted; The successfully matched body movement videos in the body movement video library are determined as the selected body movement videos.

4. The method according to claim 3, characterized in that: The multiple action tags include: action type, action duration and at least one character attribute, and the character attribute includes: at least one of age, gender, and clothing.

5. The method according to claim 4, characterized in that Matching the action tags of the body movement videos in the body movement video library based on the action tags of the body movement to be inserted includes: Based on the action type and action duration of the to-be-inserted limb action, the action type and action duration of the limb action video in the limb action video library are matched, and an intermediate limb action video with the same action type and the same action duration is determined; Matching the character attributes of the intermediate limb action video based on the character attributes of the limb action to be inserted, and displaying prompt information associated with the intermediate limb action video if no intermediate limb action video with the same character attributes is matched; In response to a second input made by the user based on the prompt information, the intermediate limb movement video selected by the user is determined as a successfully matched limb movement video.

6. The method according to claim 2, characterized in that In response to the first input, a body movement video is selected from a body movement video library, including: In response to the first input, based on the semantic content of the target text or target speech input by the user, determining the body movement to be inserted; wherein the body movement to be inserted is a body movement matching the semantic content; A body movement video containing the body movement to be inserted is selected from the body movement video library.

7. The method according to claim 1, characterized in that The step of obtaining a second video containing the body movement of the target virtual object includes: Obtaining an initial body movement video of the target virtual object recorded in the same or similar clothing as the first virtual object; For the initial limb action video, the target area in each video frame is removed to obtain the second video; wherein the target area includes all areas except the target limb part of the target virtual object.

8. The method according to any one of claims 1 to 7, characterized in that: The cutting edge of the target limb part is a horizontal or vertical straight line.

9. The method according to any one of claims 1 to 7, characterized in that: Based on the time sequence correspondence between the second video and the target video segment, the target limb part of each video frame of the second video is spliced ​​to the missing part of the target limb part of each video frame of the video to be synthesized, and after generating the synthesized video, the method further includes: Reduce the skin color difference and / or clothing color difference between the first stitching area and the second stitching area in each target video frame, wherein the target video frame is a video frame formed by stitching the target limb part in the synthesized video, the first stitching area is an area corresponding to the target limb part in the second video, and the second stitching area is an area corresponding to the first virtual object in the video to be synthesized.

10. The method according to any one of claims 1 to 7, characterized in that: The first virtual object and the second virtual object are both virtual characters, and the target limb parts include: arm.

11. The method according to any one of claims 1 to 7, characterized in that: Obtaining a first video containing a first virtual object, including: Acquire an initial character video of the first virtual character and a lip image during reading a target text; For the initial person video, the lip image is spliced ​​to the lip part of each video frame in the initial person video to obtain the first video.

12. The method according to claim 11, characterized in that The first stitching area and the third stitching area of ​​each target video frame do not overlap, wherein the target video frame is a video frame formed by stitching the target limb parts in the synthetic video, the first stitching area is the area corresponding to the target limb part in the second video, and the third stitching area is the area corresponding to the lip image.

13. An electronic device, characterized in that: It comprises a memory and a processor; the memory is connected to the processor and is used to store programs; the processor is used to implement the method for synthesizing video as described in any one of claims 1 to 12 by running the program in the memory.

14. A computer program product, characterized in that include: A computer program, which, when executed by a processor, implements the method for synthesizing video as described in any one of claims 1 to 12.

Citation Information

Patent Citations

  • Virtual character generation method and device, equipment and storage medium

    CN114219879A

  • Virtual image action generation method and device and action library construction method and device

    CN116958342A

  • Composite image output device and composite image output processing program

    JP2009086785A