Video processing method and apparatus, electronic device, and storage medium
By generating multiple short action videos for input materials and sorting them by quality, users can select and combine them to generate high-quality long videos. This solves the problems of low quality and high computational resource consumption in long video generation, and achieves efficient generation of high-quality long videos.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANGHAI BILIBILI TECH CO LTD
- Filing Date
- 2024-08-07
- Publication Date
- 2026-08-04
AI Technical Summary
The quality of long videos is difficult to improve and they consume a lot of computing resources, making it difficult to improve existing technologies.
By acquiring input materials and determining multiple action templates, action videos with a duration of less than a predetermined threshold are generated. The videos are then sorted and output based on their quality. Users can choose to combine high-quality action videos to generate longer videos.
It improves the quality of long video generation, reduces computing resource consumption, simplifies the quality requirements for user-input materials, and enhances video generation efficiency.
Smart Images

Figure CN119071537B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of data processing technology, specifically to a video processing method, a video processing apparatus, an electronic device, a computer-readable storage medium, and a computer program product. Background Technology
[0002] Videos featuring relatively natural-looking virtual avatars can be generated based on footage actually shot by the user.
[0003] The methods described in this section are not necessarily methods that had been previously conceived or adopted. Unless otherwise specified, no method described in this section should be assumed to be prior art simply because it is included in this section. Similarly, unless otherwise specified, the issues mentioned in this section should not be considered to be accepted in any prior art. Summary of the Invention
[0004] This disclosure provides a data processing method, a video processing apparatus, an electronic device, a computer-readable storage medium, and a computer program product.
[0005] According to one aspect of this disclosure, a video processing method is provided, comprising: acquiring input material for generating a video; determining a plurality of action templates for processing the input material, wherein each action template corresponds to a different action performed by a virtual character; generating a plurality of action videos based on the input material and the plurality of action templates, wherein a virtual character in each action video performs an action in a corresponding action template and the duration of the action video is less than a predetermined duration threshold; sorting the plurality of action videos based on video quality; and outputting at least a portion of the plurality of action videos according to the sorting.
[0006] According to another aspect of this disclosure, a video processing apparatus is also provided, comprising: an input unit configured to acquire input material for generating a video; a template determination unit configured to determine a plurality of action templates for processing the input material, wherein each action template corresponds to a different action performed by a virtual avatar; a generation unit configured to generate a plurality of action videos based on the input material and the plurality of action templates, wherein a virtual avatar in each action video performs an action in a corresponding action template and the duration of the action video is less than a predetermined duration threshold; a sorting unit configured to sort the plurality of action videos based on video quality; and an output unit configured to output at least a portion of the plurality of action videos according to the sorting.
[0007] According to another aspect of this disclosure, an electronic device is also provided, comprising: at least one processor; and a memory communicatively connected to said at least one processor; wherein said memory stores a computer program that, when executed by said at least one processor, implements the method described above.
[0008] According to another aspect of this disclosure, a non-transitory computer-readable storage medium storing a computer program is also provided, wherein the computer program implements the method described above when executed by a processor.
[0009] According to another aspect of this disclosure, a computer program product is also provided, comprising a computer program, wherein the computer program, when executed by a processor, implements the method described above.
[0010] Using the embodiments provided in this disclosure, different motion videos can be generated from the same input material, and the output can be sorted based on video quality. In this way, users can select one or more of the output motion videos to combine according to their actual needs. When the video quality of each motion video is high, the quality of the combined output video will also be improved.
[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0012] The accompanying drawings exemplify embodiments and form part of the specification, serving together with the textual description to explain exemplary implementations of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. Throughout the drawings, the same reference numerals refer to similar but not necessarily identical elements.
[0013] Figure 1 An exemplary flowchart of a video processing method according to an embodiment of the present disclosure is shown;
[0014] Figure 2 An exemplary user interface for a video processing method according to an embodiment of the present disclosure is shown;
[0015] Figure 3 An exemplary block diagram of a video processing apparatus according to an embodiment of the present disclosure is shown;
[0016] Figure 4 A structural block diagram of an exemplary electronic device that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation
[0017] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0018] In this disclosure, unless otherwise stated, the use of terms such as "first," "second," etc., to describe various elements is not intended to limit the positional, temporal, or importance relationships of these elements; such terms are merely used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of that element, while in other cases, based on the context, they may refer to different instances.
[0019] The terminology used in the description of the various examples described in this disclosure is for the purpose of describing particular examples only and is not intended to be limiting. Unless the context explicitly indicates otherwise, an element may be one or more unless the number of elements is specifically limited. Furthermore, the term "and / or" as used in this disclosure covers any one of the listed items and all possible combinations thereof.
[0020] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.
[0021] Using user-shot video footage as input, relatively natural-looking virtual avatar videos can generally be generated. However, as video length increases, generating continuous long videos remains challenging. In related technologies, improving video generation algorithms to enhance the quality of long videos is difficult and computationally expensive.
[0022] To improve the generation of long videos, this disclosure provides a new video processing method.
[0023] Figure 1 An exemplary flowchart of a video processing method according to an embodiment of the present disclosure is shown.
[0024] In step S102, input materials for generating the video are obtained.
[0025] In step S104, multiple action templates for processing the input material are determined, wherein each action template corresponds to a different action performed by the virtual avatar.
[0026] In step S106, multiple action videos are generated based on the input materials and multiple action templates, wherein the virtual character in each action video performs the action in the corresponding action template and the duration of the action video is less than a predetermined duration threshold.
[0027] In step S108, the multiple action videos are sorted based on video quality.
[0028] In step S110, at least a portion of the multiple motion videos are output according to the sorting.
[0029] The video processing method provided by the embodiments of this disclosure can generate different motion videos for the same input material and output them based on video quality ranking. In this way, users can select one or more of the output motion videos to combine according to actual needs. When the video quality of each motion video is high, the quality of the combined output video will also be improved.
[0030] The principles of this disclosure will now be described in detail.
[0031] In step S102, input materials for generating the video are obtained. Users can upload the materials to the server via the network. The input materials may include audio, text, or a combination of audio and text. A trained model can be used to extract features from the input materials, such as textual and motion information. The model can be any suitable machine-executable model used for image or video processing.
[0032] In step S104, multiple action templates for processing the input material are determined, wherein each action template corresponds to a different action performed by the virtual avatar.
[0033] The action template can be predetermined and used to define the actions that the virtual avatar should perform. These actions can include gestures, facial expressions, etc.
[0034] In some embodiments, different action templates can be defined for different virtual avatars, enabling the customization of personalized action habits for each virtual avatar. Step S104 may include: determining a user-specified virtual avatar in response to user input, and determining multiple action templates corresponding to the user-specified virtual avatar. The action templates corresponding to the virtual avatar can be determined by receiving the user-specified virtual avatar input on the user interface. In the example, the user can select which action templates to use.
[0035] In step S106, multiple action videos are generated based on the input materials and multiple action templates, wherein the virtual character in each action video performs the action in the corresponding action template and the duration of the action video is less than a predetermined duration threshold.
[0036] The features extracted in step S102 can be combined with multiple predefined action templates to generate corresponding action videos. A trained machine learning model can be used to generate action videos based on input materials and action templates. The action video will include a virtual avatar performing actions from the corresponding action template and incorporating features extracted from the input materials. For example, the virtual avatar in the action video will perform gestures defined in the action template and will have lip movements that match the text information in the input materials.
[0037] Action videos corresponding to different action templates can be generated using a single input clip. Alternatively, different input clips can be selected for each action template. The predetermined duration threshold for the action video can be pre-determined based on actual conditions. By limiting the duration of the action video to no more than the predetermined threshold, the computational cost of video generation can be reduced, thereby improving the model's generation quality.
[0038] In step S108, the multiple action videos are sorted based on video quality.
[0039] In some embodiments, at least one qualified motion video can be determined from multiple motion videos based on predetermined image quality standards and duration standards, and the at least one qualified motion video can be sorted based on predetermined sorting rules.
[0040] In the example, the duration standard can include a video length of less than 30 seconds and greater than 10 seconds. By limiting the upper and lower limits of the video length, it is possible to ensure that sufficient output information is retained in the video and avoid the quality degradation and limitations on the computational requirements for generating the video caused by generating a long video.
[0041] Predetermined image quality standards may include one or more of the following:
[0042] - Ensure the proportion of the face in the frame is appropriate; avoid being too close to or too far from the camera.
[0043] - Pay attention to the changes in light on the face within the clip, and be careful to avoid obvious swallowing of saliva.
[0044] - Avoid long-distance displacement when adjusting standing or sitting posture, as well as large-amplitude body and head shaking.
[0045] - The character's eyes must look directly at the camera. Blinking is allowed, but glancing up, down, left, or right is not permitted.
[0046] - Avoid using closed eyes, body swaying, or hand gestures as the beginning and end points of the template.
[0047] - The character does not make frequent hand gestures, ensuring that the user's gesture frequency remains consistent.
[0048] - High-frequency editing or pauses within the video during this action sequence to avoid camera cuts.
[0049] - Avoid videos with excessive background noise higher than human voice, or videos with poor audio recording throughout.
[0050] The image quality standards mentioned above can be prioritized as follows: head pose > facial expression > frequency of hand movements.
[0051] When sorting the generated motion videos, they can be ranked according to the number of videos that meet the image quality standards. The more videos that meet these standards, the higher they are ranked. For videos with the same number of videos meeting the same image quality standards, they are ranked according to their duration. The longer the video, the higher it is ranked.
[0052] In step S110, at least a portion of the multiple motion videos are output according to their order. The output motion videos can be displayed sequentially on the user interface according to their order, allowing the user to intuitively understand the image quality of the output motion videos. In the example, a portion (e.g., the first three) or all of the multiple motion videos can be output. The number of output motion videos can be determined according to the actual situation.
[0053] In some embodiments, method 100 may further include displaying at least a portion of a plurality of motion videos in a sorted manner on a user interface, and generating an output video using the selected motion video in response to the user selecting at least one of the displayed motion videos.
[0054] As mentioned earlier, using the video processing method described above, users can obtain multiple high-quality action videos with limited durations based on input materials. By combining and editing these short but high-quality videos, high-quality long videos can be quickly generated.
[0055] In some embodiments, the output video may include multiple action videos selected by the user and transition segments between the different action videos. It is understood that the multiple action videos selected by the user are generated separately based on user-uploaded materials, thus lacking continuity between the different action videos. To make the synthesized output video smooth, transition segments can be inserted between the different action videos as transitions, thereby improving the user's viewing experience. In some instances, the transition segments may be generated using text-to-video conversion, and the transition segments may not include virtual avatars from the action videos. Using this method, multiple different videos can be combined in a natural way to obtain a longer output video.
[0056] Using the above method, the task of generating a single long video can be decomposed into multiple shorter video generation tasks. This reduces the difficulty of model training, decreases computational resource consumption, and simultaneously improves the quality of generated videos. By providing users with multiple different action videos as candidate segments for the output video, users can easily combine multiple shorter videos to obtain a high-quality long video output. Furthermore, by limiting the duration of each action video, the quality requirements for user-uploaded input materials are also reduced. Users only need to ensure that each input material is of high quality within a short duration, without needing to prepare longer input materials.
[0057] Figure 2 An exemplary user interface for a video processing method according to an embodiment of the present disclosure is shown.
[0058] like Figure 2 As shown, the video output results according to an embodiment of the present disclosure are displayed in user interface 200. A video preview is shown in area 210. Three motion videos are shown in areas 220-1 to 220-3. The three motion videos shown can be generated using a combination of... Figure 1 The process described generates videos corresponding to different action templates. Users can select from three action videos shown in areas 220-1 to 220-3, and a preview of the selected video will be displayed in area 210. In the example, users can modify the action names of the shown action videos. A prompt can be displayed on the user interface: "Multiple action clips have been intelligently generated for you. You can select the action with the higher matching degree according to the scene you appear in for digital avatar synthesis."
[0059] Figure 3 An exemplary block diagram of a video processing apparatus according to an embodiment of the present disclosure is shown.
[0060] like Figure 3 As shown, the video processing device 300 may include an input unit 310, a template determination unit 320, a generation unit 330, a sorting unit 340, and an output unit 350.
[0061] Input unit 310 can be configured to acquire input material for generating videos. Template determination unit 320 can be configured to determine multiple action templates for processing the input material, wherein each action template corresponds to a different action performed by a virtual avatar. Generation unit 330 can be configured to generate multiple action videos based on the input material and the multiple action templates, wherein the virtual avatar in each action video performs an action in a corresponding action template. Sorting unit 340 can be configured to sort the multiple action videos based on video quality. Output unit 350 can be configured to output at least a portion of the multiple action videos according to the sorting.
[0062] In some embodiments, the input material includes at least one of audio and text.
[0063] In some embodiments, determining multiple action templates for processing the input material includes: determining a user-specified virtual avatar in response to user input; and determining multiple action templates corresponding to the user-specified virtual avatar.
[0064] In some embodiments, sorting the plurality of motion videos based on video quality includes: determining at least one qualified motion video from the plurality of motion videos based on predetermined image quality standards and duration standards; and sorting the at least one qualified motion video based on predetermined sorting rules.
[0065] In some embodiments, the duration standard includes a video length of less than 30 seconds and greater than 10 seconds.
[0066] In some embodiments, the output unit is further configured to: display at least a portion of the plurality of motion videos on a user interface according to the sorting; and generate an output video using the selected motion videos in response to a user selecting at least two of the displayed motion videos.
[0067] In some embodiments, the output video includes at least two selected motion videos and a transition segment between the different motion videos.
[0068] In some embodiments, the connecting segments are generated using a text-to-video method and do not include the virtual avatar.
[0069] It should be understood that Figure 3 Each unit of the device 300 shown can be connected to a reference. Figure 1 The steps in method 100 described correspond to each other. Therefore, the operations, features, and advantages described above for method 100 also apply to apparatus 300 and its constituent units. For the sake of brevity, some operations, features, and advantages will not be repeated here.
[0070] It should also be understood that this article can describe various technologies in the general context of software and hardware components or program modules. The above regarding... Figure 3 The described units can be implemented in hardware or in hardware in combination with software and / or firmware. For example, these units can be implemented as computer program code / instructions configured to execute in one or more processors and stored in a computer-readable storage medium. Alternatively, these units can be implemented as hardware logic / circuit. For example, in some embodiments, one or more of the input unit 310, template determination unit 320, generation unit 330, sorting unit 340, and output unit 350 can be implemented together in a system on a chip (SoC). The SoC may include an integrated circuit chip (which includes a processor (e.g., a central processing unit (CPU), microcontroller, microprocessor, digital signal processor (DSP), etc.), memory, one or more communication interfaces, and / or one or more components of other circuitry) and may optionally execute received program code and / or include embedded firmware to perform functions.
[0071] According to another aspect of this disclosure, an electronic device is also provided, comprising: at least one processor; and a memory communicatively connected to said at least one processor; wherein said memory stores a computer program that, when executed by said at least one processor, implements the method described above.
[0072] According to another aspect of this disclosure, a non-transitory computer-readable storage medium storing a computer program is also provided, wherein the computer program implements the method described above when executed by a processor.
[0073] According to another aspect of this disclosure, a computer program product is also provided, comprising a computer program, wherein the computer program, when executed by a processor, implements the method described above.
[0074] See Figure 4The present invention describes a structural block diagram of an electronic device 400 that can serve as a server or client of the present disclosure, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device can be different types of computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0075] like Figure 4 As shown, the electronic device 400 may include at least one processor 401, working memory 402, input unit 404, display unit 405, speaker 406, storage unit 407, communication unit 408 and other output units 409 that are capable of communicating with each other via system bus 403.
[0076] Processor 401 may be a single processing unit or multiple processing units, and all processing units may include single or multiple computing units or multiple cores. Processor 401 may be implemented as one or more microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units, state machines, logic circuits, and / or any device that manipulates signals based on operating instructions. Processor 401 may be configured to acquire and execute computer-readable instructions stored in working memory 402, storage unit 407, or other computer-readable media, such as program code of operating system 402a, program code of application program 402b, etc.
[0077] Working memory 402 and storage unit 407 are examples of computer-readable storage media for storing instructions that are executed by processor 401 to perform the various functions described above. Working memory 402 may include both volatile and non-volatile memory (e.g., RAM, ROM, etc.). Furthermore, storage unit 407 may include hard disk drives, solid-state drives, removable media including external and removable drives, memory cards, flash memory, floppy disks, optical disks (e.g., CDs, DVDs), storage arrays, network-attached storage, storage area networks, etc. Working memory 402 and storage unit 407 may be collectively referred to herein as memory or computer-readable storage media, and may be non-transitory media capable of storing computer-readable, processor-executable program instructions as computer program code that can be executed by processor 401 as a specific machine configured to perform the operations and functions described in the examples herein.
[0078] Input unit 406 can be any type of device capable of inputting information to electronic device 400. Input unit 406 can receive input digital or character information and generate key signal input related to user settings and / or function control of electronic device, and can include, but is not limited to, a mouse, keyboard, touch screen, trackpad, trackball, joystick, microphone and / or remote control. Output unit can be any type of device capable of presenting information, and can include, but is not limited to, display unit 405, speaker 406 and other output units 409. Other output units 409 can include, but are not limited to, video / audio output terminals, vibrators and / or printers. Communication unit 408 allows electronic device 400 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and can include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers and / or chipsets, such as Bluetooth devices, 802.8 devices, WiFi devices, WiMax devices, cellular communication devices and / or the like.
[0079] The application program 402b in working register 402 can be loaded to execute the various methods and processes described above, for example... Figure 1 Steps S102-S108 in the above description. For example, in some embodiments, the method 100 described above may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 407. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 400 via storage unit 407 and / or communication unit 408. When the computer program is loaded and executed by processor 401, one or more steps of the method 100 described above may be performed. Alternatively, in other embodiments, processor 401 may be configured to perform method 100 by any other suitable means (e.g., by means of firmware).
[0080] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0081] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0082] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0083] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0084] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0085] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other.
[0086] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0087] While embodiments or examples of this disclosure have been described with reference to the accompanying drawings, it should be understood that the methods, systems, and devices described above are merely exemplary embodiments or examples, and the scope of the invention is not limited by these embodiments or examples, but only by the granted claims and their equivalents. Various elements in the embodiments or examples may be omitted or replaced by their equivalents. Furthermore, the steps may be performed in a different order than that described in this disclosure. Further, various elements in the embodiments or examples may be combined in various ways. Importantly, as the technology evolves, many elements described herein can be replaced by equivalents that appear after this disclosure.
Claims
1. A video processing method, comprising: Obtain the input materials used to generate the video; Multiple action templates are determined for processing the input material, wherein each action template corresponds to a different action performed by the virtual avatar; Multiple action videos are generated based on the input materials and the multiple action templates, wherein the virtual character in each action video performs the action in the corresponding action template and the duration of the action video is less than a predetermined duration threshold. At least one qualified action video is determined from the plurality of action videos based on predetermined image quality standards and duration standards, wherein the predetermined image quality standards include head posture, facial expression issues, and frequency of hand movements in a priority order. The at least one qualified action video is sorted according to a predetermined sorting rule; as well as At least a portion of the plurality of motion videos are output according to the sorting, wherein the output at least a portion of motion videos is used to generate an output video, the output video including the at least a portion of motion videos and connecting segments located between different motion videos.
2. The video processing method as described in claim 1, wherein, The input materials include at least one of audio and text.
3. The video processing method as described in claim 1, wherein, The multiple action templates used to process the input material include: In response to user input, determine the virtual avatar specified by the user; and Determine multiple action templates corresponding to the virtual avatar specified by the user.
4. The video processing method as described in claim 1, wherein, The duration standard includes videos with a length of less than 30 seconds and more than 10 seconds.
5. The video processing method as described in claim 1, further comprising: At least a portion of the plurality of motion videos are displayed on the user interface according to the sorting; In response to the user selecting at least two of the displayed motion videos, an output video is generated using the selected motion videos.
6. The video processing method as described in claim 5, wherein, The output video includes the selected at least two motion videos and the connecting segments between the different motion videos.
7. The video processing method as described in claim 6, wherein the connecting segments are generated using a text-to-video method and do not include the virtual avatar.
8. A video processing apparatus, comprising: The input unit is configured to acquire input materials for generating the video; The template determination unit is configured to determine a plurality of action templates for processing the input material, wherein each action template corresponds to a different action performed by the virtual avatar; The generation unit is configured to generate multiple action videos based on the input material and the multiple action templates, wherein the virtual image in each action video performs the action in the corresponding action template and the duration of the action video is less than a predetermined duration threshold. The sorting unit is configured to sort the plurality of motion videos based on video quality; An output unit is configured to output at least a portion of the plurality of motion videos according to the ordered sequence, wherein the output at least a portion of the motion videos is used to generate an output video, the output video including the at least a portion of the motion videos and transition segments located between the different motion videos. The sorting of the multiple action videos based on video quality includes: At least one qualified action video is determined from the plurality of action videos based on predetermined image quality standards and duration standards, wherein the predetermined image quality standards include head posture, facial expression issues, and frequency of hand movements in a priority order. The at least one qualified action video is sorted according to a predetermined sorting rule.
9. The video processing apparatus as claimed in claim 8, wherein, The input materials include at least one of audio and text.
10. The video processing apparatus of claim 9, wherein, The multiple action templates used to process the input material include: In response to user input, determine the virtual avatar specified by the user; and Determine multiple action templates corresponding to the virtual avatar specified by the user.
11. The video processing apparatus of claim 8, wherein, The duration standard includes videos with a length of less than 30 seconds and more than 10 seconds.
12. The video processing apparatus of claim 8, wherein the output unit is further configured to: At least a portion of the plurality of motion videos are displayed on the user interface according to the sorting; In response to the user selecting at least two of the displayed motion videos, an output video is generated using the selected motion videos.
13. The video processing apparatus of claim 12, wherein, The output video includes the selected at least two motion videos and the connecting segments between the different motion videos.
14. The video processing apparatus of claim 13, wherein the connecting segments are generated by text-to-video conversion and do not include the virtual avatar.
15. An electronic device comprising: At least one processor; as well as A memory that is communicatively connected to the at least one processor; in The memory stores a computer program that, when executed by the at least one processor, implements the method according to any one of claims 1-7.
16. A non-transitory computer-readable storage medium storing a computer program, wherein, The computer program, when executed by a processor, implements the method according to any one of claims 1-7.
17. A computer program product comprising a computer program, wherein, The computer program, when executed by a processor, implements the method according to any one of claims 1-7.