Video processing method and apparatus, electronic device, and storage medium
By generating digital human models and producing target videos based on source material and preview videos, the reliance on real-life appearance and environment in traditional video processing methods has been resolved, enabling the generation of diverse digital human images and improving video quality.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SHANGHAI BILIBILI TECH CO LTD
- Filing Date
- 2025-09-23
- Publication Date
- 2026-06-04
AI Technical Summary
Traditional video processing methods have strict requirements on the appearance, body movements, and facial expressions of real people, and are related to the shooting environment, making it difficult to generate diverse digital human images.
By acquiring input video and preview audio, a digital human model is generated, and a target video is generated based on the source material and the selected preview video. Video clips are screened using few-shot learning and image/text machine review technology to ensure the quality of the digital human's appearance and audio.
This allows users to select a suitable digital avatar to generate videos based on their actual situation, improving video quality and security and preventing the use of malicious materials.
Smart Images

Figure CN2025123331_04062026_PF_FP_ABST
Abstract
Description
Video processing methods and apparatus, electronic devices and storage media
[0001] This application claims priority to Chinese Patent Application No. 202411727914.8, filed on November 27, 2024, entitled "Video Processing Method and Apparatus, Electronic Device and Storage Medium", the entire contents of which are incorporated herein by reference. Technical Field
[0002] This disclosure relates to the field of data processing technology, specifically to a video processing method, a video processing apparatus, an electronic device, a computer-readable storage medium, and a computer program product. Background Technology
[0003] Traditional voice-over videos require real people to appear on camera, which places certain demands on the real person's appearance, body language, facial expressions, wording, and intonation of the reading script. It is also directly related to the shooting and recording equipment and the external environment.
[0004] The methods described in this section are not necessarily methods that had been previously conceived or adopted. Unless otherwise specified, no method described in this section should be assumed to be prior art simply because it is included in this section. Similarly, unless otherwise specified, the issues mentioned in this section should not be considered to be accepted in any prior art. Summary of the Invention
[0005] This disclosure provides a data processing method, a video processing apparatus, an electronic device, a computer-readable storage medium, and a computer program product.
[0006] According to one aspect of this disclosure, a video processing method is provided, comprising: acquiring an input video and preview audio for training a digital human model; generating a digital human model based on the input video; applying the preview audio to the digital human model to generate multiple preview videos, wherein the digital human images in the multiple preview videos correspond to the user images in the input video, and the audio of the multiple preview videos is the same as the preview audio; acquiring materials for generating a target video and selecting at least one preview video among the multiple preview videos; and generating the target video based on the materials and the selected preview video, wherein the target video has audio corresponding to the materials.
[0007] Optionally, generating a digital human model based on the input video includes:
[0008] Multiple video segments are determined from the input video, wherein each video segment has a corresponding user avatar; and
[0009] The digital human model is generated based on the multiple video clips.
[0010] Optionally, the digital human figure in each of the plurality of preview videos corresponds to the user figure in each of the plurality of video segments.
[0011] Optionally, determining multiple video segments from the input video includes:
[0012] The plurality of video segments are determined from the input video based on video review standards, wherein the video review standards include video quality standards and video content standards, and the determined plurality of video segments simultaneously meet the video quality standards and the video content standards.
[0013] Optionally, the material includes at least one of audio and text.
[0014] Optionally, generating the target video based on the source material and the selected preview video includes:
[0015] Determine whether the material meets the material review standards; and
[0016] In response to the material meeting the material review criteria, the target video is generated based on the material and the selected preview video.
[0017] Optionally, it also includes:
[0018] In response to the fact that the material does not meet the material review criteria, the generation of the target video based on the material and the selected preview video is abandoned.
[0019] Optionally, the material review standards include voiceprint review standards for audio and text content review standards for text.
[0020] Optionally, the user avatar and the digital human avatar include lip movements, facial expressions, and body movements.
[0021] Optionally, generating a digital human model based on the input video includes:
[0022] The digital human model is generated based on the input video using few-shot learning.
[0023] According to another aspect of this disclosure, a video processing apparatus is also provided, comprising: a first input unit configured to acquire an input video and a preview audio for training a digital human model; a model generation unit configured to generate a digital human model based on the input video; a preview unit configured to apply the preview audio to the digital human model to generate multiple preview videos, wherein the digital human images in the multiple preview videos correspond to the user images in the input video, and the audio of the multiple preview videos is the same as the preview audio; a second input unit configured to acquire materials for generating a target video and a selection of at least one preview video among the multiple preview videos; and an output unit configured to generate the target video based on the materials and the selected preview video, wherein the target video has audio corresponding to the materials.
[0024] According to another aspect of this disclosure, an electronic device is also provided, comprising: at least one processor; and a memory communicatively connected to said at least one processor; wherein said memory stores computer-readable instructions that, when executed by said at least one processor, implement the method described above.
[0025] According to another aspect of this disclosure, a non-transitory computer-readable storage medium storing computer-readable instructions is also provided, wherein the computer-readable instructions, when executed by a processor, implement the method described above.
[0026] According to another aspect of this disclosure, a computer program product is also provided, including computer-readable instructions, wherein the computer-readable instructions, when executed by a processor, implement the method described above.
[0027] Using the embodiments provided in this disclosure, multiple preview videos generated after the digital human model has been trained can be previewed, and a target video can be generated based on at least one preview video selected from the multiple preview videos. In this way, users can select at least one preview video from the multiple preview videos as the basis for inference according to actual conditions or personal preferences, thereby generating the desired digital human video.
[0028] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0029] The accompanying drawings exemplify embodiments and form part of the specification, serving together with the textual description to explain exemplary implementations of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. Throughout the drawings, the same reference numerals refer to similar but not necessarily identical elements.
[0030] Figure 1 shows an exemplary flowchart of a video processing method according to an embodiment of the present disclosure;
[0031] Figure 2 illustrates a schematic diagram of selecting multiple video segments in a video processing method according to an embodiment of the present disclosure;
[0032] Figure 3 illustrates an exemplary detailed flowchart of a video processing method according to an embodiment of the present disclosure;
[0033] Figure 4 illustrates an exemplary block diagram of a video processing apparatus according to an embodiment of the present disclosure; and
[0034] Figure 5 shows a structural block diagram of an exemplary electronic device that can be used to implement embodiments of the present disclosure. Embodiments of the present invention
[0035] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0036] In this disclosure, unless otherwise stated, the use of terms such as "first," "second," etc., to describe various elements is not intended to limit the positional, temporal, or importance relationships of these elements; such terms are merely used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of that element, while in other cases, based on the context, they may refer to different instances.
[0037] The terminology used in the description of the various examples described in this disclosure is for the purpose of describing particular examples only and is not intended to be limiting. Unless the context explicitly indicates otherwise, an element may be one or more unless the number of elements is specifically limited. Furthermore, the term "and / or" as used in this disclosure covers any one of the listed items and all possible combinations thereof.
[0038] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.
[0039] A digital human (or meta-human) is a virtual character or avatar generated through computer technology that can simulate human appearance, behavior, and language. Digital humans typically employ technologies such as 3D modeling, animation, artificial intelligence (AI), and natural language processing (NLP) to achieve interaction with users. The emergence of digital humans has broken the strong dependence of video works on users and their surrounding environment. Users can shoot video footage of a certain length to train digital human models. After the digital human model is trained, it can infer near-realistic video content based on the textual material the user wants to understand.
[0040] In related technologies, user-uploaded videos used for training digital human models often only produce digital human models that generate a single digital human image.
[0041] To enable users to select a suitable digital human image from multiple digital human images for digital human video synthesis, this disclosure provides a new video processing method.
[0042] Figure 1 shows an exemplary flowchart of a video processing method according to an embodiment of the present disclosure.
[0043] In step S102, the input video and preview audio for training the digital human model are obtained.
[0044] In step S104, a digital human model is generated based on the input video.
[0045] In step S106, the preview audio is applied to the digital human model to generate multiple preview videos, wherein the digital human images in the multiple preview videos correspond to the user images in the input video, and the audio of the multiple preview videos is the same as the preview audio.
[0046] In step S108, materials for generating the target video are obtained and at least one preview video is selected from a plurality of preview videos.
[0047] In step S110, a target video is generated based on the source material and the selected preview video, wherein the target video has audio corresponding to the source material.
[0048] The video processing method provided by the embodiments of this disclosure allows for the previewing of multiple preview videos generated after the digital human model has been trained, and the generation of a target video based on at least one preview video selected from the multiple preview videos. In this way, users can select at least one preview video from the multiple preview videos as the basis for inference according to actual conditions or personal preferences, thereby generating the desired digital human video.
[0049] The principles of this disclosure will now be described in detail.
[0050] In step S102, the input video and preview audio for training the digital human model are acquired. Users can upload their personal spoken video and preview audio for digital human model training to the server via the network.
[0051] In step S104, a digital human model is generated based on the input video.
[0052] In some embodiments, step S104 may include: determining multiple video segments from the input video, wherein each video segment has a corresponding user image; and generating a digital human model based on the multiple video segments.
[0053] In some examples, user-uploaded videos may include multiple user avatars, for example, at different times. For instance, if the user avatar is a doctor for the first 0-10 minutes of the input video and a teacher for the second 10-20 minutes, then the user-uploaded input video is determined to be two video segments, with the first segment featuring the user avatar of a doctor and the second segment featuring the user avatar of a teacher, and a digital human model is generated based on these two video segments.
[0054] In some embodiments, step S104 may include: determining multiple video segments from the input video based on video review standards, wherein the video review standards include video quality standards and video content standards, and the determined multiple video segments simultaneously meet both video quality standards and video content standards.
[0055] In some examples, video quality standards may include one or more of the following:
[0056] - Ensure the proportion of the face in the frame is appropriate; avoid being too close to or too far from the camera.
[0057] - Pay attention to the changes in light on the face within the clip, and be careful to avoid obvious swallowing of saliva.
[0058] - Avoid long-distance displacement when adjusting standing or sitting posture, as well as large-amplitude body and head swaying.
[0059] - The character's eyes must look directly at the camera. Blinking is allowed, but glancing up, down, left, or right is not permitted.
[0060] - Avoid using closed eyes, body swaying, or hand gestures as the beginning and end points of the template.
[0061] - The character does not make frequent hand gestures, ensuring that the user's gesture frequency remains consistent.
[0062] - High-frequency editing or pauses within the video during this action sequence to avoid camera cuts.
[0063] - Avoid videos with excessive background noise higher than human voice, or videos with poor audio recording throughout.
[0064] In some examples, video content standards may include confirming that the video content does not involve illegal, irregular, or other content that is unsuitable for dissemination in the video.
[0065] In some examples, automated image and text review can be used to filter input videos and identify multiple video clips that simultaneously meet video quality and content standards for use in generating digital human models. Automated image and text review (also known as automatic image and text moderation) is a technology that uses artificial intelligence and machine learning to automatically review image and text content. It is widely used in content moderation, social media, and advertising moderation to ensure that published content complies with relevant laws, regulations, and platform rules. Compared to manual review, it can reduce review costs and improve review efficiency.
[0066] By using video review standards to filter user-uploaded videos, we can not only improve the quality of the videos used for training digital humans, thereby improving the quality of the generated digital human models, but also eliminate malicious videos uploaded by users.
[0067] In some embodiments, step S104 may include: generating a digital human model based on the input video using few-shot learning. Few-shot learning is a machine learning method that allows a model to learn new concepts or tasks with very few training samples. Unlike traditional machine learning, which requires large amounts of labeled data, few-shot learning can achieve effective learning and prediction even with scarce data. In some examples, users only need to upload a few minutes to half an hour of personal spoken video to train an effective digital human model.
[0068] In step S106, the preview audio is applied to the digital human model to generate multiple preview videos, wherein the digital human images in the multiple preview videos correspond to the user images in the input video, and the audio of the multiple preview videos is the same as the preview audio.
[0069] In some examples, preview audio is applied to a trained digital human model, which can then use this preview audio to perform inference and generate multiple preview videos with the same audio as the preview audio.
[0070] In some embodiments, the digital human figure in each of the multiple preview videos corresponds to the user figure in each of the multiple video clips.
[0071] In some examples, the user personas in multiple (e.g., two) video clips are doctors and teachers, respectively. Then, the digital personas in each of the multiple (e.g., two) preview videos are doctors and teachers corresponding to the user personas in the multiple video clips.
[0072] In some embodiments, user avatars and digital human avatars include lip movements, facial expressions, and body movements.
[0073] In step S108, materials for generating the target video are obtained and at least one preview video is selected from a plurality of preview videos.
[0074] In some examples, users can upload materials for inference. Users can preview at least one of multiple preview videos and then (e.g., based on their actual needs or personal preferences) select at least one preview video as a template video for inference.
[0075] In some embodiments, the material includes at least one of audio and text. In some examples, users can upload audio, text, or a combination of audio and text as material for reasoning purposes.
[0076] In step S110, a target video is generated based on the source material and the selected preview video, wherein the target video has audio corresponding to the source material.
[0077] In some examples, the user-uploaded materials include audio. The uploaded audio is then used to perform digital human audio-to-video inference on the selected preview video (this process is also known as image-driven, which combines a digital human model trained on video samples with timbre information to simulate lip movements, facial expressions, voices, body movements, etc., to infer the complete video content), generating a digital human video. The audio of this digital human video is the uploaded audio.
[0078] In some examples, user-uploaded materials include text, which is then converted into audio (e.g., via TTS (Text-to-Speech) technology). The generated audio is then used to perform digital human audio-to-video inference on a selected preview video, generating a digital human video whose audio is the audio generated from the text conversion. TTS (Text-to-Speech) technology is a technique that converts written text into audible speech by analyzing the text content to generate natural and fluent speech output.
[0079] In some embodiments, step S110 includes: determining whether the material meets the material review criteria; and in response to the material meeting the material review criteria, generating a target video based on the material and the selected preview video.
[0080] In some examples, image and text machine review can be used to determine whether user-uploaded materials meet the material review standards.
[0081] In some embodiments, in response to the material not meeting the material review criteria, the generation of the target video based on the material and the selected preview video is abandoned.
[0082] In some examples, in response to (e.g., using image and text machine review) determining that the material does not meet the material review criteria, the generation of the target video based on the material and the selected preview video is abandoned, and the user can also be notified that the uploaded material does not meet the material review criteria.
[0083] In some embodiments, the material review criteria include voiceprint review criteria for audio and text content review criteria for text.
[0084] In some examples, text content review criteria for text include determining (e.g., through automated review of images and text) that the text uploaded by users involves illegal, irregular, or other content unsuitable for public dissemination.
[0085] Voiceprint verification is a biometric technology that verifies identity by analyzing and recognizing an individual's vocal characteristics. Similar to fingerprints, each person's voice has unique frequencies, pitches, and pronunciation patterns, which can be used for identity verification. In some examples, voiceprint verification criteria for audio include determining (e.g., through voiceprint verification) that the voiceprint of an audio file uploaded by a user matches the voiceprint of an input video previously uploaded by the user for training a digital human model, thereby preventing users from misusing others' voices and causing infringement or other risks.
[0086] As mentioned earlier, using the video processing method described above, multiple preview videos generated after the digital human model has been trained can be previewed, and a target video can be generated based on at least one preview video selected from these multiple preview videos. In this way, users can choose at least one preview video from the multiple preview videos as the basis for inference, according to actual needs or personal preferences, thereby generating the desired digital human video.
[0087] Using the above method, users only need to upload a short personal video for digital human model training, and can preview multiple preview videos generated by the digital human model. After the digital human model is trained, users can upload materials and select preview videos as the basis for material inference according to the actual situation or personal preferences, thereby generating the digital human video that the user wants. It can also prevent users from stealing other people's voices and uploading malicious videos for digital human models or malicious materials for inference.
[0088] Figure 2 illustrates a schematic diagram of selecting multiple video segments in a video processing method according to an embodiment of the present disclosure.
[0089] Figure 2 illustrates a schematic diagram of selecting multiple video segments from a video sequence 200 according to an embodiment of the present disclosure. In the example shown in Figure 2, the entire video sequence comprises N frames, and relevant features (lip movements, body movements, and facial expressions, etc.) of the person in each frame are extracted. Frames 2 and 3 are selected as segment 1 from the N frames, and frames 5 and 6 are selected as segment 2.
[0090] Figure 3 shows an exemplary detailed flowchart of a video processing method 300 according to an embodiment of the present disclosure.
[0091] An exemplary specific process of the video processing method 300 may include the following steps:
[0092] Step 301: The user submits video materials and audio materials for generating the preview video to the creative-interface.
[0093] Step 302: The creation gateway stores the materials in the database and sends them to the creation service (bcut-interface).
[0094] Step 303: The creative team transcodes the source video into different video resolutions, so that the reviewers or the machine can select different resolutions for high-quality review.
[0095] Step 304: Conduct review training for creative work;
[0096] Step 305: The Studio backend reviews the footage frame by frame, checks for any substandard footage, selects the segments that can be used for training, and marks the excellent segments.
[0097] Step 306: The Studio backend will send the approved materials and annotation information to the algorithm engineering team for digital human model training;
[0098] Step 307: The algorithm engineering team sends the trained digital human model assets back to the Studio backend.
[0099] Step 308: The Studio backend performs manual or machine review of the trained digital human model;
[0100] Step 309: After the model is approved, the Studio backend will send the assets trained by the digital human model (which may include preview videos inferred from audio materials) to the user, for example, via SMS or web pop-up notification.
[0101] Step 310: The user uploads the reasoning text to the authoring gateway. In some examples, the user can also upload the reasoning audio or a combination of text and audio.
[0102] Step 311: The creation gateway can save users' drafts and works (text, audio recordings, background images and background music) to the creation business at any time, so that users can see the previous draft content when they enter the website or application again after an abnormal shutdown.
[0103] Step 312: The creation process stores the user's drafts and works in a database, such as MySQL;
[0104] Step 313: Create an inference task for the creation business and send it to the AI gateway (aikit);
[0105] Step 314: The AI gateway sends the inference task to the algorithm engineering for inference execution. The algorithm engineering depends on the following components: text and image machine review, TTS, voiceprint verification, and image-driven processing. When the inference material uploaded by the user is text, the TTS model needs to be called to convert the text into an audio file. Before conversion, the text and image machine review is used to check for sensitive content. If any is found, the synthesis request is rejected and the user is informed. When the inference material uploaded by the user is audio, voiceprint verification is required to prevent the user from stealing other people's voiceprints. Then, the audio file or text uploaded by the user is converted into an audio file by TTS, and combined with the digital human inference model to adapt lip movements, facial expressions, body movements, etc. The user can also select different segments from the preview video for synthesis, and finally infer the digital human video.
[0106] Step 315: The algorithm engineer will train and infer the digital human videos and send them back to the AI gateway. The user data is stored in the cloud, and users can submit their work directly through the website.
[0107] In some examples, the Studio backend can also report the approval content of each node to the HIVE data warehouse.
[0108] Figure 4 shows an exemplary block diagram of a video processing apparatus according to an embodiment of the present disclosure.
[0109] As shown in Figure 4, the video processing device 400 may include a first input unit 410, a model generation unit 420, a preview unit 430, a second input unit 440, and an output unit 450.
[0110] The first input unit 410 can be configured to acquire input video and preview audio for training the digital human model. The model generation unit 420 can be configured to generate a digital human model based on the input video. The preview unit 430 can be configured to apply the preview audio to the digital human model to generate multiple preview videos, wherein the digital human image in the preview videos corresponds to the user image in the input video, and the audio of the multiple preview videos is the same as the preview audio. The second input unit 440 can be configured to acquire material for generating a target video and select at least one preview video from the multiple preview videos. The output unit 450 can be configured to generate a target video based on the material and the selected preview video, wherein the target video has audio corresponding to the material.
[0111] In some embodiments, generating a digital human model based on an input video includes: determining multiple video segments from the input video, wherein each video segment has a corresponding user image; and generating a digital human model based on the multiple video segments.
[0112] In some embodiments, the digital human figure in each of the multiple preview videos corresponds to the user figure in each of the multiple video clips.
[0113] In some embodiments, the material includes at least one of audio and text.
[0114] In some embodiments, the output unit is further configured to: determine whether the material meets the material review criteria; and in response to the material meeting the material review criteria, generate a target video based on the material and the selected preview video.
[0115] In some embodiments, the output unit is further configured to: abandon generating a target video based on the material and the selected preview video in response to the material not meeting the material review criteria.
[0116] In some embodiments, the material review criteria include voiceprint review criteria for audio and text content review criteria for text.
[0117] In some embodiments, determining multiple video segments from an input video includes: determining multiple video segments from an input video based on video review standards, wherein the video review standards include video quality standards and video content standards, and the determined multiple video segments simultaneously meet both video quality standards and video content standards.
[0118] In some embodiments, user avatars and digital human avatars include lip movements, facial expressions, and body movements.
[0119] In some embodiments, generating a digital human model based on an input video includes: using few-shot learning to generate a digital human model based on an input video.
[0120] It should be understood that the various units of the apparatus 400 shown in FIG. 4 can correspond to the various steps in the method 100 described with reference to FIG. 1. Therefore, the operations, features, and advantages described above for method 100 also apply to apparatus 400 and its constituent units. For the sake of brevity, some operations, features, and advantages will not be repeated here.
[0121] It should also be understood that various techniques can be described in the general context of software hardware elements or program modules. The various units described above with respect to Figure 4 can be implemented in hardware or in hardware combined with software and / or firmware. For example, these units can be implemented as computer-readable instruction code / instructions configured to execute in one or more processors and stored in a computer-readable storage medium. Alternatively, these units can be implemented as hardware logic / circuit. For example, in some embodiments, one or more of the first input unit 410, model generation unit 420, preview unit 430, second input unit 440, and output unit 450 can be implemented together in a System on Chip (SoC). The SoC may include an integrated circuit chip (which includes a processor (e.g., a Central Processing Unit (CPU), microcontroller, microprocessor, digital signal processor (DSP), etc.), memory, one or more communication interfaces, and / or one or more other components of circuitry) and may optionally execute received program code and / or include embedded firmware to perform functions.
[0122] According to another aspect of this disclosure, an electronic device is also provided, comprising: at least one processor; and a memory communicatively connected to said at least one processor; wherein said memory stores computer-readable instructions that, when executed by said at least one processor, implement the method described above.
[0123] According to another aspect of this disclosure, a non-transitory computer-readable storage medium storing computer-readable instructions is also provided, wherein the computer-readable instructions, when executed by a processor, implement the method described above.
[0124] According to another aspect of this disclosure, a computer program product is also provided, including computer-readable instructions, wherein the computer-readable instructions, when executed by a processor, implement the method described above.
[0125] Referring to Figure 5, a structural block diagram of an electronic device 500 that can serve as a server or client of this disclosure is described below, which is an example of hardware devices that can be applied to various aspects of this disclosure. The electronic device can be different types of computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the disclosure described and / or claimed herein.
[0126] As shown in Figure 5, the electronic device 500 may include at least one processor 501, working memory 502, input unit 504, display unit 505, speaker 506, storage unit 507, communication unit 508 and other output units 509 that are capable of communicating with each other via system bus 503.
[0127] Processor 501 may be a single processing unit or multiple processing units, and all processing units may include single or multiple computing units or multiple cores. Processor 501 may be implemented as one or more microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units, state machines, logic circuits, and / or any device that manipulates signals based on operating instructions. Processor 501 may be configured to acquire and execute computer-readable instructions stored in working memory 502, storage unit 507, or other computer-readable media, such as program code of operating system 502a, program code of application program 502b, etc.
[0128] Working memory 502 and storage unit 507 are examples of computer-readable storage media for storing instructions that are executed by processor 501 to perform the various functions described above. Working memory 502 may include both volatile and non-volatile memory (e.g., RAM, ROM, etc.). Furthermore, storage unit 507 may include hard disk drives, solid-state drives, removable media including external and removable drives, memory cards, flash memory, floppy disks, optical disks (e.g., CDs, DVDs), storage arrays, network-attached storage, storage area networks, etc. Working memory 502 and storage unit 507 may be collectively referred to herein as memory or computer-readable storage media, and may be non-transitory media capable of storing computer-readable, processor-executable program instructions as computer program (computer-readable instruction) code, which may be executed by processor 501 as a specific machine configured to perform the operations and functions described in the examples herein.
[0129] Input unit 506 can be any type of device capable of inputting information to electronic device 500. Input unit 506 can receive input digital or character information and generate key signal input related to user settings and / or function control of electronic device, and can include, but is not limited to, a mouse, keyboard, touch screen, trackpad, trackball, joystick, microphone and / or remote control. Output unit can be any type of device capable of presenting information, and can include, but is not limited to, display unit 505, speaker 506 and other output units 509. Other output units 509 can include, but are not limited to, video / audio output terminals, vibrators and / or printers. Communication unit 508 allows electronic device 500 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and can include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers and / or chipsets, such as Bluetooth devices, 802.8 devices, WiFi devices, WiMax devices, cellular communication devices and / or the like.
[0130] The application program 502b in working register 502 can be loaded to execute the various methods and processes described above, such as steps S102-S110 in FIG. 1. For example, in some embodiments, the method 100 described above can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 507. In some embodiments, part or all of the computer-readable instructions can be loaded and / or installed on the electronic device 500 via storage unit 507 and / or communication unit 508. When the computer-readable instructions are loaded and executed by processor 501, one or more steps of the method 100 described above can be performed. Alternatively, in other embodiments, processor 501 can be configured to execute method 100 by any other suitable means (e.g., by means of firmware).
[0131] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: implementations of one or more computer-readable instructions that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transferring data and instructions to the storage system, the at least one input device, and the at least one output device.
[0132] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0133] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0134] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0135] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0136] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via a communication network. The client-server relationship is created by computer-readable instructions running on the respective computers and having a client-server relationship with each other.
[0137] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0138] While embodiments or examples of this disclosure have been described with reference to the accompanying drawings, it should be understood that the methods, systems, and devices described above are merely exemplary embodiments or examples, and the scope of this disclosure is not limited by these embodiments or examples, but only by the granted claims and their equivalents. Various elements in the embodiments or examples may be omitted or replaced by their equivalents. Furthermore, the steps may be performed in a different order than that described in this disclosure. Further, various elements in the embodiments or examples may be combined in various ways. Importantly, as technology evolves, many elements described herein can be replaced by equivalents that appear after this disclosure.
Claims
1. A video processing method, comprising: Acquire input video and preview audio for training the digital human model; A digital human model is generated based on the input video; The preview audio is applied to the digital human model to generate multiple preview videos, wherein the digital human images in the multiple preview videos correspond to the user images in the input video, and the audio of the multiple preview videos is the same as the preview audio. Acquire materials for generating the target video and select at least one preview video from the plurality of preview videos; and A target video is generated based on the source material and the selected preview video, wherein the target video has audio corresponding to the source material.
2. The video processing method as described in claim 1, wherein, Generating a digital human model based on the input video includes: Multiple video segments are determined from the input video, wherein each video segment has a corresponding user avatar; and The digital human model is generated based on the multiple video clips.
3. The video processing method as described in claim 2, wherein, The digital human figure in each of the multiple preview videos corresponds to the user figure in each of the multiple video clips.
4. The video processing method as described in claim 2, wherein, Determining multiple video segments from the input video includes: The plurality of video segments are determined from the input video based on video review standards, wherein the video review standards include video quality standards and video content standards, and the determined plurality of video segments simultaneously meet the video quality standards and the video content standards.
5. The video processing method as described in claim 1, wherein, The materials include at least one of audio and text.
6. The video processing method as described in claim 5, wherein, Generating the target video based on the aforementioned materials and the selected preview video includes: Determine whether the material meets the material review standards; and In response to the material meeting the material review criteria, the target video is generated based on the material and the selected preview video.
7. The video processing method as described in claim 6, further comprising: In response to the fact that the material does not meet the material review criteria, the generation of the target video based on the material and the selected preview video is abandoned.
8. The video processing method as described in any one of claims 6-7, wherein, The material review standards include voiceprint review standards for audio and text content review standards for text.
9. The video processing method as described in claim 1, wherein, The user avatar and the digital human avatar include lip movements, facial expressions, and body movements.
10. The video processing method as described in claim 1, wherein, Generating a digital human model based on the input video includes: The digital human model is generated based on the input video using few-shot learning.
11. A video processing apparatus, comprising: The first input unit is configured to acquire input video and preview audio for training the digital human model; The model generation unit is configured to generate a digital human model based on the input video; A preview unit is configured to apply the preview audio to the digital human model to generate multiple preview videos, wherein the digital human images in the multiple preview videos correspond to the user images in the input video, and the audio of the multiple preview videos is the same as the preview audio. The second input unit is configured to acquire materials used to generate the target video and to select at least one preview video from the plurality of preview videos; and The output unit is configured to generate a target video based on the source material and a selected preview video, wherein the target video has audio corresponding to the source material.
12. An electronic device, comprising: At least one processor; as well as A memory that is communicatively connected to the at least one processor; in The memory stores computer-readable instructions that, when executed by the at least one processor, perform the following operations: Acquire input video and preview audio for training the digital human model; A digital human model is generated based on the input video; The preview audio is applied to the digital human model to generate multiple preview videos, wherein the digital human images in the multiple preview videos correspond to the user images in the input video, and the audio of the multiple preview videos is the same as the preview audio. Acquire materials for generating the target video and select at least one preview video from the plurality of preview videos; and A target video is generated based on the source material and the selected preview video, wherein the target video has audio corresponding to the source material.
13. The electronic device of claim 12, wherein, Generating a digital human model based on the input video includes: Multiple video segments are determined from the input video, wherein each video segment has a corresponding user avatar; and The digital human model is generated based on the multiple video clips.
14. The electronic device of claim 13, wherein, The digital human figure in each of the multiple preview videos corresponds to the user figure in each of the multiple video clips.
15. The electronic device of claim 13, wherein, Determining multiple video segments from the input video includes: The plurality of video segments are determined from the input video based on video review standards, wherein the video review standards include video quality standards and video content standards, and the determined plurality of video segments simultaneously meet the video quality standards and the video content standards.
16. The electronic device of claim 12, wherein, The materials include at least one of audio and text.
17. The electronic device of claim 16, wherein, Generating the target video based on the aforementioned materials and the selected preview video includes: Determine whether the material meets the material review standards; and In response to the material meeting the material review criteria, the target video is generated based on the material and the selected preview video.
18. The electronic device of claim 17, wherein the at least one processor further performs the following operations when executed: In response to the fact that the material does not meet the material review criteria, the generation of the target video based on the material and the selected preview video is abandoned.
19. A non-transitory computer-readable storage medium storing computer-readable instructions, wherein, The computer-readable instructions, when executed by a processor, implement the method according to any one of claims 1-10.
20. A computer program product comprising computer-readable instructions, wherein, The computer-readable instructions, when executed by a processor, implement the method according to any one of claims 1-10.