Video editing method and device, medium and electronic equipment
The method addresses the challenge of aligning lip movements with audio in video editing by separating and synchronizing audio and video frames, achieving high accuracy and reduced computational demands, thereby enhancing video editing quality and efficiency.
Patent Information
- Application Number
- CN202410059654.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-15
- Publication Date
- 2025-07-15
AI Technical Summary
The accuracy of lip editing and sound editing is limited, the computing resource requirements are too high, and the alignment of the dynamic effects of the lips with the sound cannot be guaranteed, resulting in poor video editing.
By separating the original video, obtaining audio and image, performing frame processing separately, and establishing the correspondence between audio and image, realizing the lip shape and sound synthesis processing, forming an end-to-end video editing system.
Improve the accuracy and efficiency of audio and image editing, simplify workflow, reduce computing resource requirements, meet the needs of real-time editing, and improve the quality of video editing.
Smart Images

Figure CN120321449A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure belongs to the field of computer technology, and particularly relates to a video editing method, a video editing device, a computer-readable storage medium, and an electronic device. Background Art
[0002] During video editing, usually certain video clips are edited, and corresponding editing processing is performed on the relevant audio. When editing scenes involving relevant people, often after editing the video, relevant dubbing processing is carried out. However, it is often difficult to have a high matching degree for the language description process of people and the similarity of voices. In this regard, by performing similarity matching on lip shapes and voices, it is possible to achieve the effect of matching voices and lip shapes for the edited video frames, increasing the realism after video editing, and ultimately improving the quality and final effect of video editing.
[0003] However, the technologies of lip shape editing and voice editing are still in the stage of continuous development and improvement. Although some achievements have been made, there are still many problems to be solved. For example, the accuracy of lip shape editing and voice editing is limited, the computational resource requirements are too high, and the method of performing similarity matching on lip shapes and voices cannot ensure the alignment of lip dynamic effects and voices. Summary of the Invention
[0004] To overcome the problems existing in the related art, the present disclosure provides a video editing method, a video editing device, a computer-readable storage medium, and an electronic device.
[0005] According to the first aspect of the embodiments of the present disclosure, a video editing method is provided. The method includes:
[0006] Obtain an original video, and separate the original video to obtain a first audio and a first image, where the first image includes a lip image;
[0007] Perform frame splitting on the first audio to obtain a second audio, and perform frame splitting on the first image to obtain a second image;
[0008] Establish a correspondence between the second audio and the second image, and perform synthesis processing on the second audio and the second image according to the correspondence to obtain a target video.
[0009] Optionally, the separating the original video to obtain a first audio and a first image includes:
[0010] Perform lip detection on the original video to determine a first image including a lip image in the original video;
[0011] Separate the original video to obtain a first audio and a first image according to the lip image.
[0012] Optionally, the processing of framing the first audio to obtain a second audio includes:
[0013] Extracting the fundamental frequency feature of the first audio and obtaining a target text;
[0014] Performing synthesis processing on the fundamental frequency feature and the target text to obtain a third audio, and performing framing processing on the third audio to obtain a second audio.
[0015] Optionally, the performing synthesis processing on the fundamental frequency feature and the target text to obtain a third audio includes:
[0016] Performing preprocessing on the target text to obtain a first text, and performing text conversion processing on the first text to obtain a phoneme sequence;
[0017] Using a speech synthesis model to perform speech synthesis processing on the phoneme sequence and the fundamental frequency feature to obtain a third audio.
[0018] Optionally, the performing framing processing on the third audio to obtain a second audio includes:
[0019] Obtaining the video duration of the original video and a preset window size;
[0020] Based on the video duration, performing framing processing on the third audio according to the window size to obtain a second audio.
[0021] Optionally, after the performing synthesis processing on the fundamental frequency feature and the target text to obtain a third audio, the method further includes:
[0022] Performing synthesis processing on the third audio and the first image to obtain a target video.
[0023] Optionally, the processing of framing the first image to obtain a second image includes:
[0024] Obtaining the video duration of the original video and a preset window size;
[0025] Based on the video duration, performing framing processing on the first image according to the window size to obtain a second image.
[0026] Optionally, the establishing the correspondence relationship between the second audio and the second image includes:
[0027] Extracting the Fourier feature of the second audio to obtain a spectrogram, and performing mapping processing on the spectrogram to obtain an audio space;
[0028] Perform mapping processing on the second image to obtain an image space, and perform feature alignment processing on the audio space and the image space to obtain the correspondence between the second audio and the second image.
[0029] Optionally, the synthesizing the second audio and the second image according to the correspondence to obtain a target video includes:
[0030] Perform mapping processing on the image space to obtain the lip feature in the second image, and use a decoder to restore the lip feature to the lip image;
[0031] Perform image stitching processing on the lip image to obtain a third image, and perform synthesizing processing on the second audio and the third image according to the correspondence to obtain a target video.
[0032] According to a second aspect of the embodiments of the present disclosure, there is provided a video editing device, including:
[0033] A video separation module, configured to obtain an original video, and separate the original video to obtain a first audio and a first image, where the first image includes a lip image;
[0034] A frame processing module, configured to perform frame processing on the first audio to obtain a second audio, and perform frame processing on the first image to obtain a second image;
[0035] A video synthesis module, configured to establish the correspondence between the second audio and the second image, and perform synthesizing processing on the second audio and the second image according to the correspondence to obtain a target video.
[0036] Optionally, the video separation module includes:
[0037] A lip detection sub-module, configured to perform lip detection on the original video to determine a first image including a lip image in the original video;
[0038] A separation processing sub-module, configured to separate the original video according to the lip image to obtain a first audio and a first image.
[0039] Optionally, the frame processing module includes:
[0040] A feature extraction sub-module, configured to extract the fundamental frequency feature of the first audio and obtain a target text;
[0041] An audio frame sub-module, configured to perform synthesizing processing on the fundamental frequency feature and the target text to obtain a third audio, and perform frame processing on the third audio to obtain a second audio.
[0042] Optionally, the audio framing sub-module includes:
[0043] A text processing unit, configured to preprocess the target text to obtain a first text, and perform text conversion processing on the first text to obtain a phoneme sequence;
[0044] A speech synthesis unit, configured to perform speech synthesis processing on the phoneme sequence and the fundamental frequency feature by using a speech synthesis model to obtain a third audio.
[0045] Optionally, the audio framing sub-module includes:
[0046] A first acquisition unit, configured to acquire the video duration of the original video and a preset window size;
[0047] A first framing unit, configured to perform framing processing on the third audio according to the window size based on the video duration to obtain a second audio.
[0048] Optionally, the video editing device further includes:
[0049] A synthesis processing module, configured to perform synthesis processing on the third audio and the first image to obtain a target video.
[0050] Optionally, the framing processing module includes:
[0051] A second acquisition sub-module, configured to acquire the video duration of the original video and a preset window size;
[0052] A second framing sub-module, configured to perform framing processing on the first image according to the window size based on the video duration to obtain a second image.
[0053] Optionally, the video synthesis module includes:
[0054] A mapping processing sub-module, configured to extract Fourier features of the second audio to obtain a spectrogram, and perform mapping processing on the spectrogram to obtain an audio space;
[0055] A feature alignment sub-module, configured to perform mapping processing on the second image to obtain an image space, and perform feature alignment processing on the audio space and the image space to obtain a correspondence between the second audio and the second image.
[0056] Optionally, the video synthesis module includes:
[0057] A feature restoration sub-module, configured to perform mapping processing on the image space to obtain lip features in the second image, and use a decoder to restore the lip features into a lip image;
[0058] The splicing processing sub-module is configured to perform image splicing processing on the lip-shaped image to obtain a third image, and perform synthesis processing on the second audio and the third image according to the corresponding relationship to obtain a target video.
[0059] According to the third aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium having computer program instructions stored thereon, and when the program instructions are executed by a processor, the steps of the video editing method provided in any one of the first aspects of the present disclosure are implemented.
[0060] According to the fourth aspect of the embodiments of the present disclosure, there is provided an electronic device, including:
[0061] A processor;
[0062] A memory for storing processor-executable instructions;
[0063] Wherein, the processor is configured to: execute the executable instructions to implement the steps of the video editing method provided in any one of the first aspects of the present disclosure.
[0064] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects:
[0065] In the method and device provided by the exemplary embodiments of the present disclosure, separating the original video to obtain the first audio and the first image provides a data basis and theoretical support for independent editing operations on the first audio and the first image simultaneously, ensuring the effect of not damaging the fidelity of the original video. Further, frame-by-frame processing is respectively performed on the first audio and the first image, meeting the real-time processing requirements of frame-by-frame processing and generation, realizing an editable process that can be processed in real time, and improving the operation efficiency and inference speed of image editing and audio editing. Furthermore, establishing the corresponding relationship between the second audio and the second image, and performing synthesis processing on the second audio and the second image according to the corresponding relationship, captures the driving relationship between the lip features and the audio features in the original video, improves the alignment effect of the image and the audio, and improves the accuracy of image editing and audio editing. In addition, compared with the method that requires multiple independent modules to implement different functions in the related art, the present application provides an end-to-end video editing method, realizing the functional integration of lip shape editing and sound editing, forming a complete video editing system, being able to better utilize the information of the original video data itself, being able to more accurately understand and generate semantic information, and also being able to better understand and generate the second audio and the second image, simplifying the work design process, reducing artificial design components, improving the parallel computing ability and operation speed of the hardware, solving the problem of excessive computing resources brought by lip shape editing and sound editing, thus completing the editing task faster and better meeting the requirements of real-time editing.
[0066] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and do not limit the present disclosure. Description of the Drawings
[0067] The drawings herein are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure.
[0068] Figure 1 Schematically shows a flowchart of a video editing method in an exemplary embodiment of the present disclosure;
[0069] Figure 2 Schematically shows a flowchart of a method for separating an original video in an exemplary embodiment of the present disclosure;
[0070] Figure 3 Schematically shows a flowchart of a method for frame - dividing a first audio in an exemplary embodiment of the present disclosure;
[0071] Figure 4 Schematically shows a flowchart of a method for synthesizing fundamental frequency features and target text in an exemplary embodiment of the present disclosure;
[0072] Figure 5 Schematically shows a flowchart of a method for frame - dividing a third audio in an exemplary embodiment of the present disclosure;
[0073] Figure 6 Schematically shows a flowchart of a method for frame - dividing a first image in an exemplary embodiment of the present disclosure;
[0074] Figure 7 Schematically shows a flowchart of a method for establishing a correspondence relationship in an exemplary embodiment of the present disclosure;
[0075] Figure 8 Schematically shows a flowchart of a method for synthesizing a second audio and a second image according to the correspondence relationship in an exemplary embodiment of the present disclosure;
[0076] Figure 9 Schematically shows a flowchart of a video editing method in an application scenario in an exemplary embodiment of the present disclosure;
[0077] Figure 10 Schematically shows an interface diagram of a video editing function in an application scenario in an exemplary embodiment of the present disclosure;
[0078] Figure 11 Schematically shows a structural diagram of a video editing device in an exemplary embodiment of the present disclosure;
[0079] Figure 12A schematic diagram schematically shows the structure of another video editing device in an exemplary embodiment of the present disclosure;
[0080] Figure 13 The structural diagram of another video editing device in an exemplary embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION
[0081] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0082] It should be noted that all actions of acquiring signals, information or data in the present disclosure are carried out in compliance with the relevant data protection laws and policies of the country where the device is located and with the authorization given by the owner of the corresponding device.
[0083] In the process of video editing, some screen segments are usually edited, and the relevant audio is edited accordingly. When editing scenes involving relevant characters, the relevant dubbing is often performed after the screen is edited.
[0084] However, it is often difficult to have a high degree of matching between the language description process and the similarity of the voice of the character. In this regard, by matching the lip shape and the voice similarly, the effect of matching the voice and the lip shape can be achieved for the edited video screen, which increases the realism of the video after editing and ultimately improves the quality and final effect of video editing.
[0085] Among them, lip editing and voice editing are two different fields, but they are both related to speech and language.
[0086] Lip editing, also known as lip reading editing, refers to the editing process of changing a person's lip movements to adapt to the sounds produced.
[0087] At present, there are relatively many technologies for lip shape editing. For example, it can be achieved through computer vision and machine learning algorithms, and is mainly used in fields such as speech recognition and speech synthesis. Many researchers are developing new algorithms and technologies to improve the accuracy and practicality of lip shape editing.
[0088] Voice editing refers to the editing and processing of audio to change its sound characteristics or improve its quality. This technology can be used in fields such as music production, film production, and speech recognition. Currently, voice editing technology has also become quite mature and can be achieved through methods such as digital signal processing and machine learning algorithms.
[0089] In terms of lip shape editing, deep learning technology, especially Convolutional Neural Networks (CNN) and Recurrent Neural Network (RNN), which are relatively popular and have good effects, have been widely used in the process of lip shape recognition and generation. Through these models, it is relatively easy to establish the connection between the sound content and the movement of the lips. In addition, some studies have also explored how to use technologies such as Generative Adversarial Network (GAN) to generate lip shape images, and how to apply lip shape editing to fields such as speech recognition and speech synthesis.
[0090] In terms of voice editing, the same deep learning technology has also been widely used in the editing and creation of voices. New voices can be created by changing the quality and characteristics of the voice. Among them, using deep learning technology to analyze human voices and how to apply these analysis results to edit and create new voices has become a research hotspot.
[0091] The technologies of lip shape editing and voice editing are still in the stage of continuous development and improvement. Although some achievements have been made, there are still many problems to be solved, such as improving the accuracy of recognition and generation, optimizing algorithms and models, etc.
[0092] First, the accuracy of lip shape editing and voice editing is limited.
[0093] The accuracy of lip shape editing is affected by many factors, such as light, lip shape, pronunciation method, etc., resulting in a poor match between the shape in the lip expression process and the sound, and there is a time delay.
[0094] The effect of voice editing is affected by many factors, such as recording quality, noise interference, audio format, etc., resulting in a large difference between the edited voice content and the real voice, and there is noise interference, etc.
[0095] Second, the computational resource requirements for lip shape editing and voice editing are too high.
[0096] Lip shape editing and voice editing require a large amount of computational resources to process audio and video data, so higher computing power and storage space are needed, which will increase the hardware cost and computing time.
[0097] Third, privacy and security issues.
[0098] Lip shape editing and voice editing need to process the user's voice data and image data, which involves the user's privacy and security issues, and corresponding measures need to be taken to ensure the security of the user's data.
[0099] Fourth, the video editing software can only ensure the update and replacement of the sound, and cannot ensure the alignment of the lip movement effect and the sound. The actual editing effect has the problem of "sound-lip misalignment", resulting in a very poor final visual effect.
[0100] In view of the problems existing in the related technologies, the present disclosure provides a video editing method. Figure 1 It is a flowchart of a video editing method shown according to an exemplary embodiment, as Figure 1 shown, the method may at least include the following steps:
[0101] Step S110. Obtain the original video, and separate the original video to obtain the first audio and the first image, and the first image includes a lip shape image.
[0102] Step S120. Perform frame division processing on the first audio to obtain the second audio, and perform frame division processing on the first image to obtain the second image.
[0103] Step S130. Establish a correspondence relationship between the second audio and the second image, and perform synthesis processing on the second audio and the second image according to the correspondence relationship to obtain the target video.
[0104] In an exemplary embodiment of the present disclosure, separating the original video to obtain the first audio and the first image provides a data basis and theoretical support for independent editing operations on the first audio and the first image simultaneously, ensuring the effect of not damaging the fidelity of the original video. Further, performing frame splitting on the first audio and the first image respectively meets the real-time processing requirements of frame-by-frame processing and generation, realizes an editable process that can be processed in real time, and improves the computing efficiency and inference speed of image editing and audio editing. Furthermore, establishing the correspondence between the second audio and the second image and performing synthesis processing on the second audio and the second image according to the correspondence captures the driving relationship between the lip features and the audio features in the original video, improves the alignment effect of the image and the audio, and improves the accuracy of image editing and audio editing. In addition, compared with the related art methods that require multiple independent modules to implement different functions, the present application provides an end-to-end video editing method, realizes the functional integration of lip shape editing and sound editing, forms a complete video editing system, can better utilize the information of the original video data itself, can more accurately understand and generate semantic information, can also better understand and generate the second audio and the second image, simplifies the work design process, reduces artificial design components, improves the parallel computing ability and computing speed of the hardware, solves the problem of excessive computing resources brought by lip shape editing and sound editing, thus completing the editing task faster and better meeting the requirements of real-time editing.
[0105] The following will elaborate on each step of the video editing method.
[0106] In step S110, obtain the original video and separate the original video to obtain the first audio and the first image, where the first image includes a lip image.
[0107] In an exemplary embodiment of the present disclosure, the original video may be a video to be edited uploaded by the user. After the user uploads the video, the editing area can be selected and the actual duration of the original video to be edited can be selected by dragging the corresponding control.
[0108] Generally, the maximum editing duration of the original video can be 10 seconds, or other duration settings. This exemplary embodiment does not make special limitations on this.
[0109] In an alternative embodiment, Figure 2 shows a schematic flow diagram of the method for separating the original video, as Figure 2 shown, the method may at least include the following steps: In step S210, perform lip detection on the original video to determine the first image including the lip image in the original video.
[0110] Lip detection, as part of face recognition technology, has been widely used. Therefore, to implement lip detection, image processing algorithms are required. For example, edge detection algorithms, image segmentation algorithms, and contour extraction algorithms, etc.
[0111] Among them, the edge detection algorithm is a technology that improves the image quality and resolution by detecting edges in the image. In lip detection, the edge detection algorithm is used to detect the contour of the lips. Through the edge detection algorithm, the edge contour of the lip region in the face image can be obtained, which is particularly important for dynamic lip shape recognition.
[0112] The image segmentation algorithm is a technology that classifies the pixels in the image according to certain criteria. In lip detection, the image segmentation algorithm is used to separate the lip region and other regions in the face image. The image segmentation algorithm can provide more accurate boundary information for lip detection.
[0113] The contour extraction algorithm is a technology that extracts contours from binary images. In lip detection, the contour extraction algorithm is used to extract the contour information of the lip region. Through the contour extraction algorithm, the shape information of the lip contour can be obtained.
[0114] Through the lip detection technology, it is possible to determine a frame including a lip image in the original video as the first image.
[0115] In step S220, the first audio and the first image are obtained by separating the original video according to the lip image.
[0116] After determining the first image containing the lip image, the video and audio in the first image can be separated, that is, the audio track is extracted to obtain the first audio, and the image track is extracted as the first image.
[0117] In this exemplary embodiment, separating the first audio and the first image in the original video provides a data basis and theoretical support for subsequent editing of the original video through two channels of image and audio.
[0118] In step S120, the first audio is frame-processed to obtain the second audio, and the first image is frame-processed to obtain the second image.
[0119] In the exemplary embodiment of the present disclosure, after separating the first audio and the first image, the first audio and the first image can be respectively frame-processed.
[0120] In an alternative embodiment, Figure 3 shows a schematic flowchart of a method for frame-processing the first audio, as Figure 3As shown, the method may at least include the following steps: In step S310, extract the fundamental frequency feature of the first audio and obtain the target text.
[0121] Fundamental frequency extraction has extensive applications in sound processing. Its most direct application is to identify the melody of music, and it can also be used in speech processing, such as assisting in the speech recognition of tonal languages (such as Chinese), and identifying emotions in speech.
[0122] Since the fundamental frequency of sound often changes over time, fundamental frequency extraction usually first divides the signal into frames (the frame length is usually dozens of milliseconds), and then extracts the fundamental frequency frame by frame.
[0123] The methods for extracting the fundamental frequency of a frame of sound can be roughly divided into time-domain methods and frequency-domain methods.
[0124] The time-domain method takes the waveform of the sound as the input, and its basic principle is to find the smallest positive period of the waveform. Of course, the periodicity of the actual signal can only be approximate.
[0125] The frequency-domain method will first perform a Fourier transform on the signal to obtain the spectrum (only take the magnitude spectrum and discard the phase spectrum). There will be spikes at integer multiples of the fundamental frequency on the spectrum, and the basic principle of the frequency-domain method is to find the greatest common divisor of these spike frequencies.
[0126] After extracting the fundamental frequency feature of the first audio, the target text that needs to be corrected input by the user through the UI (User Interface) can be obtained. The target text can be the content that is expected to be added to the first audio, or the content to modify the first audio, etc., and this exemplary embodiment does not make special limitations on this.
[0127] Among them, the text length of the target text can be set to 200 characters, or it can be set to other lengths, and this exemplary embodiment does not make special limitations on this.
[0128] In step S320, perform synthesis processing on the fundamental frequency feature and the target text to obtain the third audio, and perform frame division processing on the third audio to obtain the second audio.
[0129] In an alternative embodiment, Figure 4 shows a schematic flowchart of the method for performing synthesis processing on the fundamental frequency feature and the target text, as Figure 4 shown, the method may at least include the following steps: In step S410, preprocess the target text to obtain the first text, and perform text conversion processing on the first text to obtain the phoneme sequence.
[0130] After extracting the fundamental frequency feature of the first audio and obtaining the target text, sound synthesis can be performed based on the fundamental frequency feature and the input target text.
[0131] Specifically, the target text is preprocessed first, including word segmentation, word splitting, and speech intonation annotation, etc.
[0132] Among them, the word segmentation technology is the cornerstone of natural language processing. Word segmentation refers to splitting a sequence of characters into individual words one by one. Word segmentation is the process of recombining a continuous sequence of characters into a sequence of words according to certain specifications.
[0133] Word segmentation algorithms can include word list-based word segmentation algorithms, statistical model-based word segmentation algorithms, and end-to-end word segmentation methods based on deep learning.
[0134] Among them, the word list-based word segmentation algorithms include the forward maximum matching algorithm, the backward maximum matching algorithm, and the bidirectional maximum matching algorithm; the statistical model-based word segmentation algorithms include the word segmentation method based on the N-gram (Chinese language model) language model, the word segmentation method based on the HMM (Hidden Markov Model), and the word segmentation method based on the CRF (Conditional Random Fields).
[0135] Specifically, the forward maximum matching method is to segment the longest word at the current position from left to right in a greedy manner for a given text input. The forward maximum matching method is a word segmentation method based on a dictionary. Its word segmentation principle is that the larger the granularity of a word, the more precise the meaning it can represent.
[0136] The corresponding algorithm steps are: (1) Generally start from the beginning position of a string and select a fragment with the maximum word length. If the sequence is less than the maximum word length, select the entire sequence; (2) First, check if the fragment is in the dictionary. If it is, it is counted as a segmented word. If not, start from the right and reduce one character, then check if the shorter fragment is in the dictionary, and loop in turn until only one character remains; (3) After the sequence becomes the remaining partial sequence after the segmentation in step 2.
[0137] The backward maximum matching method is the same as the forward method, but for a given text input, it segments the longest word at the current position from right to left in a greedy manner.
[0138] The algorithm steps are: (1) Generally start from the beginning position of a string and select a fragment with the maximum word length. If the sequence is less than the maximum word length, select the entire sequence; (2) First, check if the fragment is in the dictionary. If it is, it is counted as a segmented word. If not, start from the left and reduce one character, then check if the shorter fragment is in the dictionary, and loop in turn until only one character remains; (3) After the sequence becomes the remaining partial sequence after the segmentation in step 2.
[0139] The bidirectional maximum matching method compares the word segmentation results obtained by the forward maximum matching method with those obtained by the reverse maximum matching method to determine the correct word segmentation method.
[0140] The main core of the word segmentation algorithm based on the statistical model is that words are stable combinations. The more frequently adjacent characters appear together in the context, the more likely they are to form a word. Therefore, the probability or frequency of adjacent characters appearing together can better reflect the credibility of word formation. Based on this, the frequency of combinations of adjacent characters in the text can be calculated and statistically analyzed to obtain their co-occurrence information. The co-occurrence information reflects the tightness of the combination relationship between Chinese characters. When the tightness is higher than a certain threshold, it can be considered that this character group may form a word. This method is also called dictionary-free word segmentation.
[0141] In addition, it can also be implemented through common word segmentation tools, such as jieba word segmentation, etc.
[0142] Jieba word segmentation is currently widely used in Chinese word segmentation. Jieba word segmentation supports three modes: exact mode, full mode, and search engine mode.
[0143] Among them, the exact mode attempts to cut the sentence most precisely and is suitable for text analysis; the full mode scans out all the words that can form words in the sentence, with very fast speed, but it cannot solve ambiguity; the search engine mode is based on the exact mode and further cuts long words to improve the recall rate, which is suitable for search engine word segmentation.
[0144] Speech annotation is a relatively common type of annotation in the data annotation industry. Speech annotation is to first "extract" the text information and various sounds contained in the speech, then transcribe or synthesize them, and add corresponding labels. The annotated data is mainly used in artificial intelligence machine learning and can be applied in fields such as speech recognition and dialogue robots.
[0145] The languages of speech annotation are generally divided into Chinese, dialects, English, etc. According to the speech duration, it can be divided into long speech and short speech. Among them, factors such as the length of the speech, the sound quality, whether there is a pre-annotation result, and whether cutting is required will have a greater impact on the speed of speech transcription.
[0146] Common annotation types in speech annotation include ASR (Automatic Speech Recognition) speech transcription, speech cutting, speech cleaning, emotion determination, voiceprint recognition, phoneme annotation, prosody annotation, pronunciation proofreading, etc.
[0147] Among them, speech transcription is a technology that converts human speech into text. Speech transcription is the process of transcribing speech data into text data, which is a relatively common annotation form in the field of data annotation.
[0148] Transcription is the process of converting characters in one alphabet into characters in another alphabet. Simply put, transcription is the corresponding conversion between characters. Speech transcription can only be correspondingly converted into characters in another alphabet to ensure a complete, unambiguous, and reversible conversion between the two alphabets. Therefore, transcription is for the conversion between alphabetic writing systems. ASR speech transcription is a high-tech that transforms speech signals into corresponding text or commands through recognition and understanding processes.
[0149] Speech segmentation is the process of identifying the boundaries between words, syllables, or phonemes in natural language. Speech segmentation is an important sub-problem in the field of speech recognition technology. Just like most natural language processing problems, speech segmentation needs to consider context, grammar, and semantics.
[0150] Speech cleaning is the process of re-reviewing and validating speech, aiming to delete duplicate information, correct existing errors, and provide speech consistency. Speech cleaning is the first step in speech data preprocessing and an important link to ensure the correctness of subsequent results.
[0151] Human speech contains a lot of information. The emotional information in speech is a very important behavioral signal reflecting human emotions. At the same time, identifying the emotional information contained in speech is an important part of realizing natural human-computer interaction. For the same speech content, when spoken with different emotions, the semantics it carries may be completely different. Only when the computer can simultaneously identify the content of the speech and the emotion carried by the speech can it accurately understand the semantics of the language. Therefore, understanding the emotion of speech can make human-computer interaction more meaningful.
[0152] Emotion judgment is to judge the emotional tendency of speech content and distinguish their emotional attitudes. By extracting specific words or phrases in speech to judge whether a piece of content has a positive, negative, or neutral attitude.
[0153] Voiceprint recognition is a type of biometric technology that aims to identify unknown voices through the feature analysis of one or more speech signals. Simply put, it is a technology to identify whether a certain sentence is spoken by a certain person.
[0154] The vocal organs used by different people when speaking are different in size and shape, so each person's voiceprint spectrogram has certain differences, mainly reflected in four aspects: resonance mode characteristics, voice purity characteristics, average pitch characteristics, and pitch range characteristics. Voiceprint recognition is to convert the sound signal into an electrical signal and then use a computer for recognition.
[0155] At present, the commonly used methods for voiceprint recognition include template matching method, nearest neighbor method, neural network method, etc.
[0156] A phoneme is the smallest speech unit divided according to the natural attributes of speech. Analyzing based on the pronunciation actions in a syllable, one action constitutes one phoneme. A phoneme is the smallest unit or the smallest speech segment that makes up a syllable, and it is the smallest linear speech unit divided from the perspective of speech quality.
[0157] The method of using the International Phonetic Alphabet to mark speech is called phonetic transcription, and there are two types: broad transcription and narrow transcription. The broad transcription method marks with phonemes that can distinguish meanings, while the narrow transcription method marks with strict phoneme distinctions to try to show the differences between each phoneme. The symbols used in the broad transcription method are limited, while the symbols used in the narrow transcription method are extremely numerous, but both have their uses.
[0158] Simply put, phoneme annotation is to annotate speech according to phonetic symbols, constituent phonemes, and pronunciations.
[0159] Prosody annotation generally adopts the method of predicting prosody based on text information. Taking Chinese annotation as an example, for prosody prediction based on text information, the prosody prediction results are usually determined according to information such as initials, finals, words, phrases, paragraphs, etc. Prosody annotation refers to determining prosody information from speech data and then annotating the annotated text with prosody symbols, which is common in speech synthesis technology.
[0160] Pronunciation correction is the process of collecting data during the entire oral training process and correcting unstandardized pronunciations.
[0161] Then, through the synthesis technology, the first text after preprocessing is converted into a corresponding phoneme sequence. Among them, the synthesis technology can refer to the WaveNet model (a deep generation model of the original audio waveform), or other technologies, and this exemplary embodiment does not make special limitations on this.
[0162] The main component of the WaveNet model is a convolutional network. Each convolutional layer performs convolution on the previous layer. The larger the convolutional kernel and the more layers, the stronger the perception ability in the time domain and the larger the perception range. During the generation process, for each generated point, place this point at the last point of the input layer and continue to iterate and generate. And a softmax (activation function) layer is used as the output layer.
[0163] In step S420, the phoneme sequence and the fundamental frequency feature are processed by the speech synthesis model for speech synthesis to obtain the third audio.
[0164] After obtaining the phoneme sequence, the phoneme sequence and the fundamental frequency feature can be input into a speech synthesis model to generate speech synthesis parameters. The speech synthesis parameters can include the pitch, volume, speech rate, prosody, etc. of the voice, so as to control the generated speech features to obtain the synthesized third audio.
[0165] It should be noted that when performing lip detection on the original video, it is also possible that no lip image is detected. In this case, only the voice editing function needs to be executed to obtain the third audio. That is, the picture of the input original video is not modified, and only the voice editing path under the guidance of the text is performed.
[0166] In an alternative embodiment, the third audio and the first image are synthesized to obtain a target video.
[0167] Specifically, the audio track of the third audio and the image track of the first image are synthesized to obtain a target video.
[0168] Among them, the target video does not include a lip image, or does not include a complete lip image, and this exemplary embodiment does not make special limitations on this.
[0169] In an alternative embodiment, Figure 5 shows a schematic flow diagram of a method for frame-dividing the third audio, as Figure 5 shown, the method may at least include the following steps: In step S510, the video duration of the original video and a preset window size are obtained.
[0170] The video duration of the original video can be 10 seconds, or other durations, and this exemplary embodiment does not make special limitations on this.
[0171] The preset window size can be 10 milliseconds, or 3 milliseconds, etc., or other sizes, and this exemplary embodiment does not make special limitations on this.
[0172] It should be noted that in order to ensure the dynamic consistency between the audio and the image, the audio sampling rate and the frame-dividing parameters set here should be consistent with the frame-dividing parameters of the image.
[0173] In step S520, based on the video duration, the third audio is frame-divided according to the window size to obtain the second audio.
[0174] After determining the video duration and the window size, the third audio can be frame-divided to obtain the second audio.
[0175] In an alternative embodiment, Figure 6 shows a schematic flow diagram of a method for frame-dividing the first image, as Figure 6As shown, the method may at least include the following steps: In step S610, obtain the video duration of the original video and a preset window size.
[0176] The video duration of the original video may be 10 seconds or other durations, and this exemplary embodiment does not make special limitations thereto.
[0177] The preset window size may be 10 milliseconds, or 3 milliseconds, etc., or other sizes, and this exemplary embodiment does not make special limitations thereto.
[0178] In step S620, based on the video duration, frame the first image according to the window size to obtain a second image.
[0179] For the image editing part, frame the first image without audio in the same window manner as the audio to obtain independent framed images as the second image.
[0180] To ensure the smoothness of the subsequent lip changes and the coherence of the video, sample the first image in time sequence. Under normal circumstances, perform frame-by-frame sampling at a certain time interval, such as 10 milliseconds. And, the first image can be sampled with a time delay, which can be called time sequence feedback.
[0181] First, start frame-by-frame sampling from 0 milliseconds and obtain the first batch of video frame image groups; then perform the second time-sequence frame-by-frame sampling at an interval of 3 milliseconds to obtain the second batch of video frame image groups, and so on to obtain multiple groups of images until the repeated sampling stops, so as to finally obtain the second image.
[0182] After obtaining the second image, lip image recognition can also be performed, or directly perform frame selection or segmentation recognition operations. When performing area editing, the frame selection area can be mainly used, and a confidence rate of 0.999 can be set to ensure that the edited second image is the lip, rather than other areas that may be the lip.
[0183] It should be noted that when performing segmentation operations on the lip image, an irregular lip area or a rectangular area containing the lip area can be segmented, and this exemplary embodiment does not make special limitations thereto.
[0184] In this exemplary embodiment, frame the first audio and the first image respectively, and realize a real-time editable process through frame-by-frame processing and short-time sliding window methods, which can process audio and image data in real time, improve the operation efficiency and inference speed of lip recognition and audio editing, and ensure work efficiency and quality.
[0185] In step S130, a correspondence between the second audio and the second image is established, and the second audio and the second image are synthesized according to the correspondence to obtain a target video.
[0186] In an exemplary embodiment of the present disclosure, after obtaining the second audio and the second image, a correspondence between the second audio and the second image can be further established.
[0187] In an alternative embodiment, Figure 7 A flowchart showing a method for establishing a correspondence is shown, as Figure 7 shown. The method may at least include the following steps: In step S710, Fourier features of the second audio are extracted to obtain a spectrogram, and the spectrogram is mapped to obtain an audio space.
[0188] For multiple sets of second images, in order to establish a correspondence between the second images and the second audio, iterative editing is performed after aligning the audio features with the lip images. The guiding condition here is to generate a spectrogram after extracting short-time Fourier features or fast Fourier features of the second audio.
[0189] Its essence is to convert the relationship between "sound-image" into establishing the relationship between "image-image". In this way, not only can the relationship between sound data and lip features be established, but also the dynamic consistency between audio and image can be ensured due to setting the same sampling rate and framing parameters for the framing process of the first audio and the framing process of the first image.
[0190] For the spectrogram features of the second audio, they can be mapped into high-dimensional features through multiple linear layers. For the time-frequency dimension and the channel dimension, both are mapped into a 256-dimensional space to obtain an audio space. Its essence should be a cube space, but the features at different corner positions of the cube can be ignored, and at the same time, the computational amount can be reduced. Therefore, this audio space can be a sphere.
[0191] In step S720, the second image is mapped to obtain an image space, and the audio space and the image space are feature-aligned to obtain a correspondence between the second audio and the second image.
[0192] Similarly, for different images of the second image, they can also be mapped into high-dimensional features through multiple linear layers. For the time-frequency dimension and the channel dimension, both are mapped into a 256-dimensional space to obtain an image space. However, due to the reason of image resolution, the image space here is no longer a sphere, but an ellipsoid.
[0193] In this way, by feature-aligning the two feature spaces of the audio space and the image space, a consistency relationship between audio and image is finally established, which can be expressed by the following formula:
[0194]
[0195] Among them, m i = f(I i ), k i = g(S i ), where i represents sample points in different dimensions, I represents image features, S represents audio features, f and g represent linear mappings, and δ controls the smoothness of cross-entropy.
[0196] In an alternative embodiment, Figure 8 shows a schematic flowchart of a method for synthesizing a second audio and a second image according to a correspondence relationship, as Figure 8 shown. The method may at least include the following steps: In step S810, mapping processing is performed on the image space to obtain lip features in the second image, and the lip features are restored to a lip image by using a decoder.
[0197] After establishing the correspondence relationship between the second audio and the second image, the lip features of different frames can be controlled by the audio features at the same position, and finally the high-dimensional features in the image space are restored to the lip features in the original second image by using a linear layer.
[0198] Furthermore, the original image resolution is restored through a Decoder (decoder) to obtain a lip image. It should be noted that the lip image here is a regular rectangular image.
[0199] The purpose of adopting the frame-by-frame generation method for multiple groups of data here is to ensure the integrity of information and prevent the loss of dynamic features.
[0200] In step S820, image splicing processing is performed on the lip image to obtain a third image, and the second audio and the third image are synthesized according to the correspondence relationship to obtain a target video.
[0201] After generating the lip image, the inverse process of video frame division can be performed again to perform image splicing processing on multiple groups of lip images to synthesize the edited video data as the third image.
[0202] Meanwhile, the image channels of the third image and the audio channels of the edited second audio are synthesized according to the correspondence relationship to obtain the final target video.
[0203] The video editing method in the embodiments of the present disclosure will be described in detail below with reference to an application scenario.
[0204] Figure 9 shows a schematic flowchart of the video editing method in the application scenario, as Figure 9 shown. In step S910, an editing area is selected.
[0205] Figure 10 The schematic diagram of the interface of the video editing function in the application scenario is shown, such as Figure 10 shown. After the user inputs a video, the editing area can be selected as the original video by dragging the "progress bar" control 1010.
[0206] After selecting the original video, person detection or lip detection can be performed.
[0207] If no relevant person or lip is detected, only the sound editing function will be executed, that is, the picture of the input video will not be modified, and only the sound editing path under the guidance of text will be performed, such as Figure 9 shown on the left; if the images of the relevant person's face or lips are detected, the functions of the lip editing path and the audio editing path will be executed.
[0208] For example, through the lip detection technology, the picture including the lip image can be determined as the first image in the original video.
[0209] After determining the first image including the lip image, the picture and audio in the first image can be separated, that is, the audio track is extracted to obtain the first audio, and the image track is extracted as the first image.
[0210] For the audio part, first extract the fundamental frequency features of the first audio and obtain the target text.
[0211] After extracting the fundamental frequency features of the first audio, the target text that needs to be corrected input by the user through the UI can be obtained. The target text can be the content that is expected to be added to the first audio, or the content that modifies the first audio, etc. This exemplary embodiment does not make special limitations on this.
[0212] After extracting the fundamental frequency features of the first audio and obtaining the target text, the synthesis of sound can be performed according to the fundamental frequency features and the input target text.
[0213] Specifically, first preprocess the target text, including word segmentation, word splitting, and speech intonation annotation, etc.
[0214] Then, convert the first text after preprocessing into the corresponding phoneme sequence through the synthesis technology. Among them, the synthesis technology can refer to the WaveNet model, or other technologies. This exemplary embodiment does not make special limitations on this.
[0215] In step S920, TTS speech synthesis.
[0216] After obtaining the phoneme sequence, the phoneme sequence and the fundamental frequency feature can be input into a speech synthesis model to generate speech synthesis parameters. The speech synthesis parameters can include the pitch, volume, speech rate, rhythm, etc. of the voice, so as to control the generated speech features to obtain the synthesized third audio.
[0217] When performing lip detection on the original video, it is also possible that no lip image is detected. In this case, only perform the voice editing function to obtain the third audio. That is, do not modify the picture of the input original video, and only perform the voice editing path under the guidance of the text.
[0218] Specifically, synthesize the audio track of the third audio and the image track of the first image to obtain the target video.
[0219] Among them, the target video does not include a lip image or does not include a complete lip image, and this exemplary embodiment does not make special limitations on this.
[0220] In step S930, image extraction.
[0221] For the image editing part, the first image without audio is framed in the same window as the audio to obtain independent framed images as the second image.
[0222] In order to ensure the smoothness of the later lip changes and the coherence of the video, sample the first image in time series. Under normal circumstances, sample frame by frame at a certain time interval, such as 10 milliseconds. And the first image can be sampled with a time delay, which can be called time series feedback.
[0223] First, start sampling frame by frame from 0 milliseconds and obtain the first batch of video frame image groups; then perform the second time series frame-by-frame sampling at an interval of 3 milliseconds to obtain the second batch of video frame image groups, and so on to obtain multiple groups of images until the repeated sampling stops, so as to finally obtain the second image.
[0224] In step S940, lip recognition.
[0225] After obtaining the second image, lip image recognition can also be performed, or directly perform the recognition operations of boxing or segmentation. When editing the area, the boxed area can be mainly used, and a confidence rate of 0.999 can be set to ensure that the edited second image is the lip, rather than other areas that may be the lip.
[0226] Similarly, the third audio is also framed.
[0227] Obtain the video duration of the original video and the preset window size.
[0228] The video duration of the original video can be 10 seconds or other durations, and this exemplary embodiment does not make special limitations on this.
[0229] The preset window size can be 10 milliseconds, 3 milliseconds, etc., or other sizes, and this exemplary embodiment does not make special limitations on this.
[0230] It should be noted that, in order to ensure the dynamic consistency between audio and images, the audio sampling rate and the frame division parameters set here should be consistent with the frame division parameters of the images.
[0231] Based on the video duration, the third audio is frame-divided according to the window size to obtain the second audio.
[0232] After determining the video duration and the window size, the third audio can be frame-divided to obtain the second audio.
[0233] In step S950, FFT.
[0234] For multiple groups of second images, in order to establish the correspondence between the second images and the second audio, iterative editing is performed after aligning the audio features with the lip images. The guiding condition here is the spectrogram generated after extracting the fast Fourier features of the second audio. This spectrogram can also be obtained by extracting the short-time Fourier features, and this exemplary embodiment does not make special limitations on this.
[0235] Its essence is to convert the relationship between "sound - image" to establish the relationship between "image - image". In this way, not only can the relationship between the sound data and the lip features be established, but also the dynamic consistency between the audio and the images can be ensured due to setting the same sampling rate and frame division parameters for the frame division processing of the first audio and the frame division processing of the first image.
[0236] For the spectrogram features of the second audio, they can be mapped to high-dimensional features through multiple linear layers. For the time-frequency dimension and the channel dimension, both are mapped to a 256-dimensional space to obtain the audio space. Its essence should be a cube space, but the features at different corner positions of the cube can be ignored, and at the same time, the calculation amount can be reduced. Therefore, this audio space can be a sphere.
[0237] Similarly, for different second images, they can also be mapped to high-dimensional features through multiple linear layers. For the time-frequency dimension and the channel dimension, both are mapped to a 256-dimensional space to obtain the image space. However, due to the image resolution, the image space here is no longer a sphere but an ellipsoid.
[0238] In this way, by performing feature alignment on the two feature spaces of the audio space and the image space, the consistency relationship between the audio and the images is finally established, which can be represented by formula (1).
[0239] After establishing the correspondence between the second audio and the second image, the lip features of different frames can be controlled by the audio features at the same position, and finally the high-dimensional features in the image space are restored to the lip features in the original second image by using a linear layer.
[0240] In step S960, the Decoder.
[0241] Furthermore, the original image resolution is restored through the decoder to obtain a lip-shaped image. It should be noted that the lip-shaped image here is a regular rectangular image.
[0242] The purpose of adopting the frame-by-frame generation method for multiple groups of data here is to ensure the integrity of information and prevent the loss of dynamic features.
[0243] In step S970, edit the video.
[0244] After generating the lip-shaped image, the inverse process of video frame division can be performed again to splice multiple groups of lip-shaped images for image splicing processing, so as to synthesize the edited video data as the third image.
[0245] In step S980, audio-visual synthesis.
[0246] The image channels of the third image and the audio channels of the edited second audio are synthesized according to the correspondence relationship to obtain the final target video, so as to output the edited target video.
[0247] Therefore, in the editing of videos containing human images, the video editing method of the present application can achieve the coordinated editing of human voices and lip shapes, ensure the matching of the mouth shapes and voices of the edited characters, and improve the quality of video editing.
[0248] The video editing method of the present application is applicable to only editing sounds, including scenes such as editing background sounds, music, and other background sounds in video pictures, and is also applicable to editing human lip language, including human lip shapes and voice content in video pictures.
[0249] Based on this, the video editing method of the present application can also provide an editing function for the photo album or video software of terminals such as mobile phones and tablets, and can also provide corresponding function interfaces for third-party APPs (applications), so that users can use the corresponding lip shape editing function to automatically recognize and convert the voices in the video, or use the corresponding sound editing tool to adjust the audio characteristics in the video to help users obtain better video effects.
[0250] In related technologies, due to the inherent nature of video editing technology, it is difficult for the editing of detailed areas such as lip shapes to the generative model. Coupled with the difficulty of audio editing, the difficulty increases again.
[0251] In an exemplary embodiment of the present disclosure, an end-to-end video editing method is provided, realizing the technologies of multimodal fusion and real-time processing.
[0252] Specifically, multimodal fusion: Lip shape editing and sound editing can simultaneously process video, video frame images, and audio data. The lip shape and lip position are recognized through video frame images, and a driving relationship is forcibly established between the lip dynamic features and audio features captured from the video to improve the alignment effect of lip reading and sound.
[0253] Real-time processing technology: Lip shape editing and sound editing need to process audio and video data in real time. Through frame-by-frame processing and short-time sliding window methods, a real-time editable process is realized, improving the speed and efficiency of lip reading recognition and sound editing.
[0254] The video editing method of this application satisfies the working mechanism of real-time processing. Since the frame-by-frame method with the same audio and video parameters is adopted, and the single-frame features of the current audio information are also used as a reference in the image editing part, it can meet the real-time processing method of frame-by-frame processing and frame-by-frame generation, improving the computing efficiency and inference speed, and ensuring the work efficiency and quality.
[0255] End-to-end improved algorithm: Current lip shape editing and sound editing technologies usually require multiple independent modules to achieve different functions, while the end-to-end video editing method can integrate these modules together to form a complete system, thus better understanding and generating lip reading and sound. The video editing method of this application establishes a complete system and establishes a complete end-to-end solution only by providing data interfaces and parameters.
[0256] The video editing method of this application is completely based on the end-to-end processing method, which can make full use of the information of the original video data itself, can more accurately understand and generate semantic information, simplifies the design work process, reduces manually designed components, improves the hardware parallel computing ability, increases the computing speed, and thus completes the video editing task faster, meeting the real-time video editing requirements.
[0257] In addition, to ensure the reliability of hardware computing and the effectiveness of editing results, the video editing method of this application sets the maximum editing duration to 10 seconds, and the text length to 200 characters, fully meeting the application requirements under normal real situations and greatly improving the application experience.
[0258] The video editing ability of this application is efficient and lightweight, expands the efficiency of video editing details, and at the same time realizes the ability of audio editing and lip movement dynamics, providing a high user experience.
[0259] The two technologies of image editing and audio editing in the video can be used simultaneously for independent editing operations, resulting in more diverse outputs without compromising the fidelity of the original video semantics.
[0260] The encoding and decoding process of video editing is fast, has low memory occupancy, is easy to deploy on mobile hardware, does not involve complex algorithms, and is easy to deploy.
[0261] In addition, the video editing method of this application has strong scalability and customizability, and can provide an end-to-end interface for third-party APPs.
[0262] In addition, in an exemplary embodiment of the present disclosure, a video editing device is also provided. Figure 11 A schematic structural diagram of the video editing device is shown, as Figure 11 shown, the video editing device 1100 may include:
[0263] A video separation module 1110, configured to obtain an original video and separate the original video to obtain a first audio and a first image, where the first image includes a lip image;
[0264] A frame processing module 1120, configured to perform frame processing on the first audio to obtain a second audio, and perform frame processing on the first image to obtain a second image;
[0265] A video synthesis module 1130, configured to establish a correspondence between the second audio and the second image, and perform synthesis processing on the second audio and the second image according to the correspondence to obtain a target video.
[0266] In some embodiments of the present disclosure, the video separation module 1110 includes:
[0267] A lip detection sub-module, configured to perform lip detection on the original video to determine a first image including a lip image in the original video;
[0268] A separation processing sub-module, configured to separate the original video according to the lip image to obtain a first audio and a first image.
[0269] In some embodiments of the present disclosure, the frame processing module 1120 includes:
[0270] A feature extraction sub-module, configured to extract the fundamental frequency feature of the first audio and obtain a target text;
[0271] The audio framing sub-module is configured to perform synthesis processing on the fundamental frequency feature and the target text to obtain a third audio, and perform framing processing on the third audio to obtain a second audio.
[0272] In some embodiments of the present disclosure, the audio framing sub-module includes:
[0273] The text processing unit is configured to preprocess the target text to obtain a first text, and perform text conversion processing on the first text to obtain a phoneme sequence;
[0274] The speech synthesis unit is configured to perform speech synthesis processing on the phoneme sequence and the fundamental frequency feature by using a speech synthesis model to obtain a third audio.
[0275] In some embodiments of the present disclosure, the audio framing sub-module includes:
[0276] The first acquisition unit is configured to acquire the video duration of the original video and a preset window size;
[0277] The first framing unit is configured to perform framing processing on the third audio according to the window size based on the video duration to obtain a second audio.
[0278] In some embodiments of the present disclosure, the video editing device 1100 further includes:
[0279] The synthesis processing module is configured to perform synthesis processing on the third audio and the first image to obtain a target video.
[0280] In some embodiments of the present disclosure, the framing processing module includes:
[0281] The second acquisition sub-module is configured to acquire the video duration of the original video and a preset window size;
[0282] The second framing sub-module is configured to perform framing processing on the first image according to the window size based on the video duration to obtain a second image.
[0283] In some embodiments of the present disclosure, the video synthesis module 1130 includes:
[0284] The mapping processing sub-module is configured to extract the Fourier features of the second audio to obtain a spectrogram, and perform mapping processing on the spectrogram to obtain an audio space;
[0285] The feature alignment sub-module is configured to perform mapping processing on the second image to obtain an image space, and perform feature alignment processing on the audio space and the image space to obtain the correspondence between the second audio and the second image.
[0286] In some embodiments of the present disclosure, the video synthesis module 1130 includes:
[0287] A feature restoration sub-module, configured to perform mapping processing on the image space to obtain the lip feature in the second image, and use a decoder to restore the lip feature into the lip image;
[0288] An image splicing processing sub-module, configured to perform image splicing processing on the lip image to obtain a third image, and perform synthesis processing on the second audio and the third image according to the corresponding relationship to obtain a target video.
[0289] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0290] The present disclosure also provides a computer-readable storage medium, on which computer program instructions are stored, and when the program instructions are executed by a processor, the steps of the video editing method provided by the present disclosure are implemented.
[0291] Figure 12 is a block diagram of another video editing device 1200 shown according to an exemplary embodiment. For example, the device 1200 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0292] Referring to Figure 12 , the device 1200 may include one or more of the following components: a processing component 1202, a memory 1204, a power supply component 1206, a multimedia component 1208, an audio component 1210, an input / output interface 1212, a sensor component 1214, and a communication component 1216.
[0293] The processing component 1202 generally controls the overall operation of the device 1200, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 1202 may include one or more processors 1220 to execute instructions to complete all or part of the steps of the above method. In addition, the processing component 1202 may include one or more modules to facilitate the interaction between the processing component 1202 and other components. For example, the processing component 1202 may include a multimedia module to facilitate the interaction between the multimedia component 1208 and the processing component 1202.
[0294] The memory 1204 is configured to store various types of data to support the operation of the device 1200. Examples of such data include instructions for any application or method operating on the device 1200, contact data, phone book data, messages, pictures, videos, and the like. The memory 1204 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0295] The power supply component 1206 provides power to various components of the device 1200. The power supply component 1206 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device 1200.
[0296] The multimedia component 1208 includes a screen that provides an output interface between the device 1200 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of the touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 1208 includes a front camera and / or a rear camera. When the device 1200 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.
[0297] The audio component 1210 is configured to output and / or input audio signals. For example, the audio component 1210 includes a microphone (MIC) that is configured to receive external audio signals when the device 1200 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 1204 or transmitted via the communication component 1216. In some embodiments, the audio component 1210 further includes a speaker for outputting audio signals.
[0298] The input / output interface 1212 provides an interface between the processing component 1202 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power-on button, and a lock button.
[0299] The sensor assembly 1214 includes one or more sensors for providing a status assessment of various aspects of the device 1200. For example, the sensor assembly 1214 can detect the on / off state of the device 1200, the relative positioning of components, such as the display and keypad of the device 1200. The sensor assembly 1214 can also detect a change in the position of the device 1200 or a component of the device 1200, the presence or absence of user contact with the device 1200, the orientation or acceleration / deceleration of the device 1200, and the temperature change of the device 1200. The sensor assembly 1214 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 1214 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 1214 can also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0300] The communication component 1216 is configured to facilitate communication between the device 1200 and other devices in a wired or wireless manner. The device 1200 can access a wireless network based on communication standards, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 1216 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 1216 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0301] In an exemplary embodiment, the device 1200 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.
[0302] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions, such as the memory 1204 including instructions, is also provided. The above instructions can be executed by the processor 1220 of the device 1200 to complete the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0303] In addition to being an independent electronic device, the above-mentioned device can also be a part of an independent electronic device. For example, in one embodiment, the device can be an integrated circuit (IC) or a chip. The integrated circuit can be a single IC or a collection of multiple ICs. The chip can include, but is not limited to, the following types: GPU (Graphics Processing Unit), CPU (Central Processing Unit), FPGA (Field Programmable Gate Array), DSP (Digital Signal Processor), ASIC (Application Specific Integrated Circuit), SOC (System on Chip), etc. The above-mentioned integrated circuit or chip can be used to execute executable instructions (or code) to implement the above-mentioned video editing method. The executable instructions can be stored in the integrated circuit or chip, or obtained from other devices or equipment. For example, the integrated circuit or chip includes a processor, a memory, and an interface for communicating with other devices. The executable instructions can be stored in the memory, and when the executable instructions are executed by the processor, the above-mentioned video editing method is implemented; or, the integrated circuit or chip can receive the executable instructions through the interface and transmit them to the processor for execution to implement the above-mentioned video editing method.
[0304] In another exemplary embodiment, a computer program product is also provided. The computer program product includes a computer program that can be executed by a programmable device. The computer program has a code portion for executing the above-mentioned video editing method when executed by the programmable device.
[0305] Figure 13 is a block diagram of yet another video editing device 1300 shown according to an exemplary embodiment. For example, the device 1300 can be provided as a server. Referring to Figure 13 , the device 1300 includes a processing component 1322, which further includes one or more processors, and memory resources represented by a memory 1332 for storing instructions executable by the processing component 1322, such as application programs. The application programs stored in the memory 1332 can include one or more modules each corresponding to a set of instructions. In addition, the processing component 1322 is configured to execute instructions to perform the above method...
[0306] Apparatus 1300 may further include a power supply component 1326 configured to perform power management of apparatus 1300, a wired or wireless network interface 1350 configured to connect apparatus 1300 to a network, and an input / output interface 1358. Apparatus 1300 may operate based on an operating system stored in memory 1332.
[0307] Those skilled in the art will readily conceive of other embodiments of the present disclosure upon considering the specification and practicing the present disclosure. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and examples are only illustrative, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0308] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A video editing method, characterized in that, Including: Obtain an original video, and separate the original video to obtain a first audio and a first image, where the first image includes a lip image; Perform frame splitting on the first audio to obtain a second audio, and perform frame splitting on the first image to obtain a second image; Establish a correspondence between the second audio and the second image, and perform synthesis processing on the second audio and the second image according to the correspondence to obtain a target video.
2. The video editing method according to claim 1, wherein The separating the original video to obtain a first audio and a first image includes: Perform lip detection on the original video to determine a first image including a lip image in the original video; Separate the original video according to the lip image to obtain a first audio and a first image.
3. The video editing method according to claim 1, characterized in that The performing frame splitting on the first audio to obtain a second audio includes: Extract the fundamental frequency feature of the first audio, and obtain a target text; Perform synthesis processing on the fundamental frequency feature and the target text to obtain a third audio, and perform frame splitting on the third audio to obtain a second audio.
4. The video editing method according to claim 3, wherein The performing synthesis processing on the fundamental frequency feature and the target text to obtain a third audio includes: Perform preprocessing on the target text to obtain a first text, and perform text conversion processing on the first text to obtain a phoneme sequence; Use a speech synthesis model to perform speech synthesis processing on the phoneme sequence and the fundamental frequency feature to obtain a third audio.
5. The video editing method according to claim 3, wherein The performing frame splitting on the third audio to obtain a second audio includes: Obtain the video duration of the original video and a preset window size; Based on the video duration, perform frame splitting on the third audio according to the window size to obtain a second audio.
6. The video editing method according to claim 3, wherein After the performing synthesis processing on the fundamental frequency feature and the target text to obtain a third audio, the method further includes: Perform synthesis processing on the third audio and the first image to obtain a target video.
7. The video editing method according to claim 1, wherein The performing frame splitting on the first image to obtain a second image includes: Obtain the video duration of the original video and a preset window size; Based on the video duration, perform frame splitting on the first image according to the window size to obtain a second image.
8. The video editing method according to claim 1, wherein The establishing the correspondence between the second audio and the second image includes: Extract the Fourier feature of the second audio to obtain a spectrogram, and perform mapping processing on the spectrogram to obtain an audio space; Perform mapping processing on the second image to obtain an image space, and perform feature alignment processing on the audio space and the image space to obtain the correspondence between the second audio and the second image.
9. The video editing method according to claim 8, wherein The performing synthesis processing on the second audio and the second image according to the correspondence to obtain a target video includes: Perform mapping processing on the image space to obtain the lip feature in the second image, and use a decoder to restore the lip feature to the lip image; Perform image stitching processing on the lip image to obtain a third image, and perform synthesis processing on the second audio and the third image according to the correspondence to obtain a target video.
10. A video editing device, characterized in that, Including: A video separation module, configured to obtain an original video and separate the original video to obtain a first audio and a first image, wherein the first image includes a lip image; A frame processing module, configured to perform frame processing on the first audio to obtain a second audio and perform frame processing on the first image to obtain a second image; A video synthesis module, configured to establish a correspondence between the second audio and the second image, and perform synthesis processing on the second audio and the second image according to the correspondence to obtain a target video.
11. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the program instructions are executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.
12. An electronic device, characterized in that, Comprising: A memory, on which a computer program is stored; A processor, configured to execute the computer program in the memory to implement the steps of the method according to any one of claims 1 to 9.