Video processing method, video processing device, electronic equipment and medium

By acquiring and generating voice-over timbre and rhythmic features, the problem of low efficiency in video dubbing in existing technologies is solved, and a flexible and efficient video dubbing method is realized.

CN119653173BActive Publication Date: 2026-08-04VIVO MOBILE COMM CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
VIVO MOBILE COMM CO LTD
Filing Date
2024-12-06
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

In existing technologies, when electronic devices add audio to videos, users need to continuously record the audio of the entire video, resulting in low efficiency and inflexible methods for adding audio to videos.

Method used

By receiving user input, acquiring voice-over timbre and rhythm features, generating voice-over audio, and generating new videos based on these features and the original video, the need for complete audio recording is reduced.

Benefits of technology

It improves the efficiency and flexibility of video dubbing, allowing users to generate dubbing audio instantly based on input without having to continuously record the audio of the entire video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119653173B_ABST
    Figure CN119653173B_ABST
Patent Text Reader

Abstract

The application discloses a video processing method, a video processing device, an electronic equipment and a medium, and belongs to the technical field of computers. The video processing method comprises the following steps: receiving a first input to a first interface, the first interface comprising a picture of a first video; in response to the first input, obtaining a dubbing timbre feature and a dubbing rhythm feature; based on the dubbing timbre feature and the dubbing rhythm feature, obtaining a dubbing audio; and based on the dubbing audio and the first video, generating a second video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of computer technology, and specifically relates to a video processing method, a video processing device, an electronic device, and a medium. Background Technology

[0002] As electronic devices, such as mobile phones, continue to improve their functionality, using electronic devices to add voiceovers to videos has become increasingly common.

[0003] Currently, when using electronic devices to add voiceovers to videos, the video must first be divided into multiple sub-videos based on the dialogue or conversations within it, and the corresponding voiceover text for each sub-video must be obtained. Each sub-video corresponds to one line of dialogue or conversation, and the voiceover text for each sub-video is the text of that line of dialogue or conversation. Then, the electronic device can use the audio recorded by the user based on the voiceover text for each sub-video, along with the corresponding sub-video, to generate a new video.

[0004] Therefore, when using electronic devices to dub videos, users need to continuously record the audio of the entire video, resulting in low efficiency and inflexible dubbing methods. Summary of the Invention

[0005] The purpose of this application is to provide a video processing method, video processing apparatus, electronic device, and medium that can improve the efficiency and flexibility of video dubbing.

[0006] In a first aspect, embodiments of this application provide a video processing method, the method comprising:

[0007] Receive the first input to the first interface, the first interface including the screen of the first video;

[0008] In response to the first input, acquire the voice timbre features and voice prosody features;

[0009] Based on the timbre and rhythmic features of the dubbing, the dubbing audio is obtained;

[0010] A second video is generated based on the dubbing audio and the first video.

[0011] Secondly, embodiments of this application provide a video processing apparatus, the apparatus comprising:

[0012] A receiving module is used to receive the first input to the first interface, which includes the image of the first video.

[0013] The processing module is used to obtain the dubbing timbre features and dubbing prosody features in response to the first input;

[0014] Based on the timbre and rhythmic features of the dubbing, the dubbing audio is obtained;

[0015] A second video is generated based on the dubbing audio and the first video.

[0016] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, wherein the memory stores a program or instructions executable on the processor, and the program or instructions, when executed by the processor, implement the steps of the video processing method as described in the first aspect.

[0017] Fourthly, embodiments of this application provide a readable storage medium storing a program or instructions that, when executed by a processor, implement the steps of the video processing method as described in the first aspect.

[0018] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the video processing method as described in the first aspect.

[0019] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the video processing method as described in the first aspect.

[0020] In this embodiment, a first input is received to a first interface, which includes a frame of a first video. In response to the first input, voice-over timbre features and voice-over rhythm features are acquired. Based on the voice-over timbre features and voice-over rhythm features, voice-over audio is acquired. Based on the voice-over audio and the first video, a second video is generated. Thus, the electronic device can acquire corresponding voice-over timbre features and voice-over rhythm features according to the user's input, and then generate voice-over audio based on these features, thereby improving the flexibility of the video dubbing method. Furthermore, it eliminates the need for the user to continuously record the audio of the entire video, thus improving the efficiency of video dubbing. Attached Figure Description

[0021] Figure 1 This is a flowchart illustrating a video processing method provided in an embodiment of this application;

[0022] Figure 2A This is a schematic diagram of a video processing interface provided in an embodiment of this application;

[0023] Figure 2B This is a schematic diagram of a video processing interface provided in an embodiment of this application;

[0024] Figure 2C This is a schematic diagram of a video processing interface provided in an embodiment of this application;

[0025] Figure 2D This is a schematic diagram of a video processing interface provided in an embodiment of this application;

[0026] Figure 2E This is a schematic diagram of a video processing interface provided in an embodiment of this application;

[0027] Figure 2F This is a schematic diagram of a video processing interface provided in an embodiment of this application;

[0028] Figure 2G This is a schematic diagram of a video processing interface provided in an embodiment of this application;

[0029] Figure 2H This is a schematic diagram of a video processing interface provided in an embodiment of this application;

[0030] Figure 2I This is a schematic diagram of a video processing interface provided in an embodiment of this application;

[0031] Figure 2J This is a schematic diagram of a video processing interface provided in an embodiment of this application;

[0032] Figure 2K This is a schematic diagram of a video processing interface provided in an embodiment of this application;

[0033] Figure 3 This is a flowchart illustrating a video processing method provided in an embodiment of this application;

[0034] Figure 4 This is a flowchart illustrating a video processing method provided in an embodiment of this application;

[0035] Figure 5 This is a flowchart illustrating a video processing method provided in an embodiment of this application;

[0036] Figure 6 This is a flowchart illustrating a video processing method provided in an embodiment of this application;

[0037] Figure 7 This is a flowchart illustrating a method for generating dubbing audio provided in an embodiment of this application;

[0038] Figure 8 This is a schematic diagram of the structure of a timbre acquisition module provided in an embodiment of this application;

[0039] Figure 9 This is a schematic diagram of the structure of a prosody acquisition module provided in an embodiment of this application;

[0040] Figure 10 This is a schematic diagram of a phoneme sequence provided in an embodiment of this application;

[0041] Figure 11 This is a flowchart illustrating a video processing method provided in an embodiment of this application.

[0042] Figure 12 This is a schematic diagram of the structure of a video processing device provided in an embodiment of this application;

[0043] Figure 13 This is a schematic diagram of the structure of a video processing device provided in an embodiment of this application;

[0044] Figure 14 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;

[0045] Figure 15 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0046] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0047] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0048] The terms "at least one," "at least one of," etc., used in the specification and claims of this application refer to any one, any two, or a combination of two or more of the included items. For example, at least one of a, b, and c can mean: "a," "b," "c," "a and b," "a and c," "b and c," and "a, b, and c," where a, b, and c can be single or multiple. Similarly, "at least two" refers to two or more items, and its meaning is similar to that of "at least one."

[0049] The video processing method, video processing device, electronic device, and medium provided in this application will be described in detail below with reference to the accompanying drawings and through specific embodiments and application scenarios.

[0050] The video processing method, video processing device, electronic device, and medium provided in this application embodiment can be applied to video dubbing scenarios, such as scenarios where a mobile phone is used to dub a video.

[0051] The video processing method provided in this application can be executed by a video processing device. Exemplarily, the video processing device can be an electronic device, or a functional component or entity within that electronic device. The following will use an electronic device as an example to illustrate the video processing method provided in this application.

[0052] Figure 1 This is a flowchart illustrating the video processing method provided in an embodiment of this application, as shown below. Figure 1 As shown, the video processing method provided in this application embodiment may include the following steps 101 to 104.

[0053] Step 101: The electronic device receives the first input to the first interface.

[0054] In some embodiments of this application, the first interface includes a frame of a first video, which is a video to be dubbed.

[0055] In some embodiments of this application, the first input mentioned above may include any of the following: user click input, swipe input, press input, voice input, gesture input, or other feasible inputs, which are not limited in this application embodiment.

[0056] In some embodiments of this application, the above-mentioned gesture input may include, but is not limited to, at least one of the following: click gesture, swipe gesture, drag gesture, pressure recognition gesture, long press gesture, area change gesture, double press gesture, double tap gesture, specific gesture input or other possible gesture inputs. The specific gesture input form can be determined according to actual needs, and is not limited in some embodiments.

[0057] In some embodiments of this application, the above-mentioned click input can be single-click input, double-click input, or any number of clicks, or it can be long-press input or short-press input. In some embodiments, this is not limited.

[0058] In some embodiments of this application, the above-mentioned sliding input can be a sliding input in any direction, such as sliding up, sliding down, sliding left, or sliding right, etc., and in some embodiments, this is not limited.

[0059] Step 102: The electronic device responds to the first input and acquires the voice-over timbre features and voice-over rhythm features.

[0060] In some embodiments of this application, the first input mentioned above includes a third sub-input and a fourth sub-input, and step 102 can be implemented by the following steps 102a and 102b:

[0061] Step 102a: The electronic device responds to the third sub-input and displays at least one dubbing mode on the first interface.

[0062] In some embodiments of this application, each dubbing mode corresponds to a method for obtaining dubbing timbre features and dubbing rhythm features.

[0063] In some embodiments of this application, the aforementioned third sub-input may include any of the following: user click input, swipe input, press input, voice input, gesture input, or other feasible inputs, which are not limited in this application embodiment.

[0064] For example, when the electronic device is a mobile phone, such as Figure 2A As shown, the phone screen 20a can display a first interface 20, which includes a first video frame 21 and a voice-over window prompt control 22. When the phone receives a user's tap input on the voice-over window prompt control 22, the phone can respond to the user's tap input on the voice-over window prompt control 22, combined with... Figure 2A ,like Figure 2B As shown, a dubbing window 23 is displayed on the first interface 20, and the dubbing window 23 includes a dubbing mode selection control 231. When the mobile phone receives a user's click input on the dubbing mode selection control 231, the mobile phone can respond to the user's click input on the dubbing mode selection control 231 and combine it with... Figure 2B ,like Figure 2C As shown, a dubbing mode selection window 24 is displayed on the first interface 21. The dubbing mode selection window 24 includes dubbing mode 1, dubbing mode 2 and dubbing mode 3. The user's click input on the dubbing mode selection control 231 is the aforementioned third sub-input.

[0065] It should be noted that the shapes of the aforementioned dubbing window prompt control, dubbing window, and dubbing mode selection window can be any possible shape, such as circle, rectangle, triangle, rhombus, annulus, or polygon. Figure 2A , Figure 2B , Figure 2C The shapes of the dubbing window prompt control, dubbing window, and dubbing mode selection window are only provided as examples. The specific shapes can be determined according to actual usage requirements. This embodiment of the invention does not impose any limitations.

[0066] Step 102b: In response to the fourth sub-input of the target dubbing mode, the electronic device acquires the dubbing timbre features and dubbing rhythm features according to the target dubbing timbre features and dubbing rhythm features corresponding to the target dubbing mode.

[0067] In some embodiments of this application, the target dubbing mode is a dubbing mode among at least one of the dubbing modes.

[0068] In some embodiments of this application, the aforementioned fourth sub-input may include any of the following: user click input, swipe input, press input, voice input, gesture input, or other feasible inputs, which are not limited in this application embodiment.

[0069] Combination Figure 2C As shown, dubbing mode 1 can acquire user timbre features and user prosodic features; dubbing mode 2 can acquire video timbre features and user prosodic features; and dubbing mode 3 can acquire user timbre features and video prosodic features. When the phone receives a fourth sub-input for dubbing mode 1, the phone can respond to the fourth sub-input for dubbing mode 1 by extracting dubbing timbre features and dubbing prosodic features from the user audio. When the phone receives a fourth sub-input for dubbing mode 2, the phone can respond to the fourth sub-input for dubbing mode 2 by extracting dubbing timbre features from the first audio and dubbing prosodic features from the user audio. When the phone receives a fourth sub-input for dubbing mode 3, the phone can respond to the fourth sub-input for dubbing mode 3 by extracting dubbing timbre features from the user audio and dubbing prosodic features from the first audio.

[0070] Step 103: The electronic device acquires the dubbing audio based on the dubbing timbre and dubbing rhythm features.

[0071] In some embodiments of this application, the aforementioned voice-over timbre features are user timbre features, and the aforementioned voice-over prosodic features are video prosodic features; combined with Figure 1 ,like Figure 3 As shown, step 103 above can be achieved through the following steps 103a1 and 103a2:

[0072] Step 103a1: Extract user timbre features from user audio and extract video prosodic features from the first audio.

[0073] In some embodiments of this application, the first audio is the audio in the first video, the user timbre features extracted from the user audio are used to characterize the timbre of the user's voice, and the video prosodic features extracted from the first audio are used to characterize the prosodic of the first video sound.

[0074] In some embodiments of this application, the electronic device can extract audio features from the user's audio to obtain the audio features of the user's audio, and then extract timbre features from the audio features to obtain the user's timbre features.

[0075] In some embodiments of this application, the electronic device can extract prosodic features and perform vector quantization on the first audio to obtain the aforementioned video prosodic features.

[0076] Step 103a2: The electronic device generates dubbing audio based on the user's timbre characteristics and video prosodic characteristics.

[0077] It should be noted that the implementation process of step 103a2 above can refer to step A1 below. To avoid repetition, this embodiment will not repeat it here.

[0078] In this way, the electronic device extracts user timbre features from the user's audio and video prosodic features from the first audio recording; based on the user timbre features and video prosodic features, it generates dubbing audio. This allows the electronic device to generate dubbing audio with the user's timbre and the video's prosodic rhythm, thereby improving the flexibility of video dubbing methods. Furthermore, it eliminates the need for users to continuously record the audio of the entire video, thus improving the efficiency of video dubbing.

[0079] In some embodiments of this application, the aforementioned dubbing timbre features are video timbre features, and the aforementioned dubbing prosody features are user prosody features; combined with Figure 1 ,like Figure 4 As shown, step 103 above can be achieved through the following steps 103b1 and 103b2:

[0080] Step 103b1: Extract user timbre features and user prosodic features from user audio.

[0081] In some embodiments of this application, prosodic features are extracted from user audio to characterize the prosody of the user's voice.

[0082] It should be noted that the process of the electronic device extracting user timbre features and user prosodic features from user audio can refer to the process of extracting user timbre features from user audio and extracting video prosodic features from the first audio. To avoid repetition, this embodiment will not repeat the process here.

[0083] Step 103b2: The electronic device generates dubbing audio based on the user's timbre and prosodic features.

[0084] It should be noted that the implementation process of step 103b2 above can refer to step A1 below. To avoid repetition, this embodiment will not repeat it here.

[0085] In this way, electronic devices extract user timbre and prosodic features from user audio; based on these features, they generate dubbing audio, enabling the devices to produce dubbing audio with the user's timbre and prosodic pattern, thus improving the flexibility of video dubbing methods. Furthermore, it eliminates the need for users to continuously record the audio of the entire video, thereby increasing the efficiency of video dubbing.

[0086] In some embodiments of this application, the aforementioned dubbing timbre features are video timbre features, and the aforementioned dubbing prosody features are user prosody features; combined with Figure 1 ,like Figure 5 As shown, step 103 above can be achieved through the following steps 103c1 and 103c2:

[0087] Step 103c1: The electronic device extracts video timbre features from the first audio and extracts user prosodic features from the user audio.

[0088] In some embodiments of this application, timbre features are extracted from the first audio to characterize the timbre of the video sound.

[0089] It should be noted that the process of the electronic device extracting video timbre features from the first audio and extracting user prosodic features from the user audio can refer to the process of extracting user timbre features from the user audio and extracting video prosodic features from the first audio. To avoid repetition, this embodiment will not repeat the process here.

[0090] Step 103c2: The electronic device generates dubbing audio based on the video timbre features and user prosodic features.

[0091] It should be noted that the implementation process of step 103c2 above can refer to step A1 below. To avoid repetition, this embodiment will not repeat it here.

[0092] In this way, the electronic device extracts video timbre features from the first audio and user prosodic features from the user's audio; based on the video timbre features and user prosodic features, it generates dubbing audio, enabling the electronic device to generate dubbing audio with the timbre of the video and the prosodic of the user, thereby improving the flexibility of video dubbing methods. Furthermore, it eliminates the need for users to continuously record the audio of the entire video, thus improving the efficiency of video dubbing.

[0093] In some embodiments of this application, the first interface further includes a text acquisition control; combined with Figure 1 ,like Figure 6 As shown, prior to step 103 above, the video processing method provided in this application embodiment may further include the following steps 105a and 105b, and step 103 above can be implemented through the following step A1:

[0094] Step 105a: The electronic device receives a third input to the voiceover text acquisition control.

[0095] In some embodiments of this application, the above-mentioned dubbing text acquisition control is used to instruct the acquisition of the dubbing text of the first video.

[0096] In some embodiments of this application, the shape of the above-mentioned voice-over text acquisition control can be any possible shape such as a circle, rectangle, triangle, rhombus, annulus or polygon, which can be determined according to actual usage requirements. This embodiment of the invention does not limit the shape.

[0097] In some embodiments of this application, the aforementioned third input may include any of the following: user click input, swipe input, press input, voice input, gesture input, or other feasible inputs, which are not limited in this application embodiment.

[0098] Step 105b: The electronic device responds to the third input and displays the voiceover text.

[0099] In some embodiments of this application, the electronic device may display a text prompt box in response to the fifth input described above, the text prompt box including the voice-over text described above.

[0100] In some embodiments of this application, the aforementioned dubbing text is the text corresponding to the aforementioned first video or the text input by the user.

[0101] In some embodiments of this application, the aforementioned dubbing text is the text corresponding to the aforementioned first video; step 105b can be implemented through the following step B1:

[0102] Step B1: In response to the third input, if the first audio is detected in the first video, the electronic device displays the dubbing text.

[0103] In some embodiments of this application, the aforementioned dubbing text is the text corresponding to the first audio in the aforementioned first video.

[0104] In some embodiments of this application, the electronic device can detect a first audio in the first video in response to the third input described above. Upon detecting the first audio, it acquires the text corresponding to the first audio and then displays the text corresponding to the first audio as the dubbing text in the text prompt box.

[0105] For example Figure 2B The dubbing window 23 may also include a dubbing text acquisition control 232. When the mobile phone receives a user's click input on the dubbing text acquisition control 232, the mobile phone can respond to the user's click input on the dubbing text acquisition control 232, and, if the first audio is detected in the first video, combine it with... Figure 2B ,like Figure 2D As shown, a text prompt box 25 is displayed on the first interface 20. The text prompt box 25 displays the voice-over text 251 and a fourth prompt message 252. The fourth prompt message 252 is used to indicate that the voice-over text 251 is the text obtained from the first video, that is, to indicate that the voice-over text 251 is the text corresponding to the first audio. The user's click input on the voice-over text acquisition control 232 is the aforementioned third input.

[0106] In this way, the electronic device, in response to the third input, displays the dubbing text when the first audio is detected in the first video, enabling the electronic device to generate dubbing audio with corresponding timbre and rhythm based on the text corresponding to the first audio detected in the first video, thereby improving the flexibility of video dubbing.

[0107] In some embodiments of this application, the aforementioned voice-over text is text input by the user, and the aforementioned third input includes a first sub-input and a second sub-input; the aforementioned step 105b can be implemented through the following steps C1 and C2:

[0108] Step C1: The electronic device responds to the first sub-input and displays a text input box on the first interface.

[0109] In some embodiments of this application, the first sub-input may include any of the following: user click input, swipe input, press input, voice input, gesture input, or other feasible inputs, which are not limited in this application embodiment.

[0110] In some embodiments of this application, the shape of the text input box can be any possible shape such as a circle, rectangle, triangle, rhombus, ring, or polygon, which can be determined according to actual usage requirements. This embodiment of the invention does not limit the shape.

[0111] Step C2: In response to the second sub-input to the text input box, the electronic device displays the dubbing text in the text input box.

[0112] In some embodiments of this application, the aforementioned second sub-input may include any of the following: user click input, swipe input, press input, voice input, gesture input, or other feasible inputs, which are not limited in this application embodiment.

[0113] In some embodiments of this application, the aforementioned dubbing text is the text corresponding to the aforementioned second sub-input.

[0114] Combination Figure 2BAs shown, when the mobile phone receives a user's click input on the voiceover text acquisition control 232, the mobile phone can respond to the user's click input on the voiceover text acquisition control 232. If the first audio is not detected in the first video, as... Figure 2E As shown, a text prompt box 25 is displayed on the first interface 20. The text prompt box 25 displays a fifth prompt message 253, which prompts the user to input voice-over text. When the mobile phone receives the user's second sub-input to the text prompt box 25 based on the fifth prompt message 253, the text corresponding to the fourth sub-input is displayed as voice-over text 251 in the text prompt box 25. The user's click on the voice-over text acquisition control 232 is the aforementioned first sub-input. For example, if the user inputs the text "Hello, nice to meet you", then "Hello, nice to meet you" will be displayed as voice-over text 251 in the text prompt box 25.

[0115] In this way, the electronic device displays a text input box on the first interface in response to the first sub-input, and displays the dubbing text in the text input box in response to the second sub-input, enabling the electronic device to generate dubbing audio with corresponding timbre and rhythm based on the text entered by the user, thereby improving the flexibility of video dubbing.

[0116] Step A1: The electronic device generates dubbing audio based on timbre features, rhythmic features, and dubbing text.

[0117] In this way, the electronic device receives a third input to the voice-over text acquisition control; in response to the third input, displays the voice-over text; and generates voice-over audio based on timbre features, rhythmic features, and the voice-over text. This enables the electronic device to generate voice-over audio with corresponding timbre and rhythm based on the voice-over text obtained by the user's operation of the voice-over text acquisition control, thereby improving the flexibility of video dubbing.

[0118] In some embodiments of this application, step A1 can be implemented by steps D1 to D4 as follows:

[0119] Step D1: The electronic device acquires the first phoneme sequence corresponding to the dubbed text;

[0120] In some embodiments of this application, step D1 can be implemented by steps a1 to a3 as follows:

[0121] Step a1: The electronic device acquires at least one first phoneme of the dubbed text and the duration of at least one first phoneme.

[0122] In some embodiments of this application, the duration of each first phoneme is the duration of a first phoneme.

[0123] In some embodiments of the present application, an electronic device may obtain the pronunciation of each character in the dubbed text, and then perform phoneme segmentation on the pronunciation of each character in the dubbed text to obtain at least one first phoneme of the above-mentioned dubbed text.

[0124] For example, the dubbed text may be "Hello", and the corresponding pronunciations are "ni" and "hao" respectively. Then, phoneme segmentation is performed on the pronunciation of each character in "Hello", and the at least one first phoneme obtained may be "sil n i3 h ao3 sil", where the "sil" at the beginning and end represents a silent phoneme, "i3" represents that "ni" is pronounced in the third tone, and "ao3" represents that "hao" is pronounced in the third tone.

[0125] In some embodiments of the present application, the electronic device may input the above at least one phoneme into a duration prediction model to obtain the at least one first phoneme duration. The above duration prediction model may be a model trained using a training set, and the training set includes multiple phoneme samples and multiple labels, and each label represents the duration corresponding to a phoneme sample.

[0126] Step a2: The electronic device determines the first quantity of each first phoneme in the at least one first phoneme based on the at least one first phoneme duration.

[0127] In some embodiments of the present application, the electronic device may calculate the ratio of the at least one first phoneme duration and the second duration to obtain the first quantity of each first phoneme in the at least one first phoneme, and the second duration is the duration of a single-frame audio.

[0128] For example, the first phoneme time corresponding to the "n" phoneme is 0.11 s, and the first duration is 10 ms. Then, the first quantity corresponding to the "n" phoneme is 0.11 s / 10 ms, that is, the first quantity corresponding to the "j" phoneme is 10.

[0129] Step a3: The electronic device sorts the at least one first phoneme based on the first quantity to obtain a first phoneme sequence.

[0130] In some embodiments of the present application, the first quantity corresponding to each first phoneme is the quantity of each first phoneme in the first phoneme sequence.

[0131] For example, the first number of "siln i3 h ao3 sil" is [22, 6, 16, 11, 22, 25], then the first phoneme sequence is "sil sil...sil nn...ni3 i3...i3 hh...h ao3 ao3...ao3 silsil...sil". In this first phoneme sequence, the number of the first "sil" is 22, the number of "n" is 6, the number of "i3" is 16, the number of "h" is 11, the number of "ao3" is 22, and the number of "sil" at the end of the first phoneme sequence is 25.

[0132] Thus, the electronic device obtains the first timbre sequence accurately and quickly by using at least one first phoneme and the duration of at least one first phoneme in the dubbed text; determining the first quantity of each first phoneme in the at least one first phoneme based on the duration of at least one first phoneme; and sorting the at least one first phoneme based on the first quantity.

[0133] Step D2: The electronic device acquires the second phoneme sequence corresponding to the target audio.

[0134] In some embodiments of this application, the target audio is the audio in the first video or the user's audio.

[0135] In some embodiments of this application, step D2 can be implemented by steps b1 to b4 as follows:

[0136] Step b1: The electronic device acquires the target text corresponding to the target audio.

[0137] In some embodiments of this application, the electronic device may first perform noise reduction processing on the target audio to obtain the noise-reduced target audio, and then perform recognition processing on the noise-reduced target audio to obtain the target text corresponding to the noise-reduced target audio.

[0138] Step b2: The electronic device acquires at least one second phoneme of the target text and the duration of at least one second phoneme.

[0139] In some embodiments of this application, the duration of each second phoneme is the duration of one second phoneme.

[0140] It should be noted that the specific process by which the electronic device acquires at least one second phoneme of the target text can refer to the specific process by which the electronic device acquires at least one first phoneme of the dubbing text. To avoid repetition, this embodiment will not repeat the process here.

[0141] In some embodiments of this application, the electronic device can obtain the third duration of the target audio, and then align the target audio with at least one second phoneme based on the third duration to obtain the second phoneme duration of each second phoneme.

[0142] Step b3: The electronic device determines a second quantity of each first phoneme in at least one second phoneme based on the duration of at least one second phoneme;

[0143] Step b4: The electronic device sorts at least one second phoneme based on the second quantity to obtain a second phoneme sequence.

[0144] It should be noted that the specific implementation process of steps b3 and b4 above can refer to steps a2 and a3 above. To avoid repetition, this embodiment will not repeat them here.

[0145] Thus, the electronic device can accurately and quickly acquire the first timbre sequence by acquiring the target text corresponding to the target audio; acquiring at least one second phoneme and the duration of at least one second phoneme in the target text; determining the second quantity of each first phoneme in the at least one second phoneme based on the duration of at least one second phoneme; and sorting the at least one second phoneme based on the second quantity to obtain the second phoneme sequence.

[0146] Step D3: The electronic device generates audio features of the dubbing audio based on the first phoneme sequence, the second phoneme sequence, the dubbing timbre features, and the dubbing rhythm features.

[0147] In some embodiments of this application, step D3 can be implemented by steps c1 to c5 as follows:

[0148] Step c1: The electronic device obtains the first splicing feature based on the second phoneme sequence and the dubbing prosody features.

[0149] In some embodiments of this application, the electronic device may employ a first embedding layer and a second embedding layer to convert the second phoneme sequence and the dubbing prosody features into a first feature vector and a second feature vector, and then concatenate the first feature vector and the second feature vector to obtain the first concatenated feature.

[0150] Step c2: The electronic device inputs the first splicing feature into the prosody prediction model to obtain the predicted prosody feature of the i-th phoneme in the first phoneme sequence.

[0151] In some embodiments of this application, the predicted prosodic features can be the prosodic features of the i-th phoneme predicted by the prosodic prediction model.

[0152] In some embodiments of this application, the first phoneme sequence includes N phonemes, where N is a positive integer and i is an integer ranging from 1 to N.

[0153] Step c3: The electronic device obtains the second splicing feature based on the predicted prosodic features and the i-th phoneme.

[0154] In some embodiments of this application, the electronic device may employ a first embedding layer and a second embedding layer respectively, using the predicted prosodic features and the i-th phoneme as a third feature vector and a fourth feature vector, and then concatenating the third feature vector and the fourth feature vector to obtain the second concatenated feature.

[0155] Step c4: The electronic device splices the second splicing feature and the first splicing feature to obtain the first splicing feature again, until i=N, to obtain at least one predicted prosodic feature, and each predicted prosodic feature corresponds to a phoneme in the first phoneme sequence;

[0156] Step c5: The electronic device obtains a third splicing feature based on at least one predicted prosodic feature and a dubbing timbre feature, and decodes the third splicing feature to generate the audio features of the dubbing text.

[0157] In some embodiments of this application, the electronic device can use a first embedding layer and a second embedding layer for each predicted prosodic feature and a dubbing timbre feature, respectively, to convert each predicted prosodic feature and dubbing timbre feature into a fifth feature vector and a sixth feature vector, and then concatenate the fifth feature vector and the sixth feature vector to obtain a third concatenated feature corresponding to each predicted prosodic feature.

[0158] In some embodiments of this application, the electronic device can decode the third splicing feature to obtain the audio feature.

[0159] Thus, the electronic device obtains a first splicing feature based on the second phoneme sequence and the dubbing prosodic features; inputs the first splicing feature into the prosodic prediction model to obtain the predicted prosodic feature of the i-th phoneme in the first phoneme sequence; based on the predicted prosodic feature and the i-th phoneme, a second splicing feature is obtained; the electronic device splices the second splicing feature and the first splicing feature to obtain the first splicing feature again, until i=N, obtaining at least one predicted prosodic feature, each predicted prosodic feature corresponding to a phoneme in the first phoneme sequence; based on at least one predicted prosodic feature and the dubbing timbre feature, a third splicing feature is obtained, and the third splicing feature is decoded to generate the audio features of the dubbing text, which can accurately and quickly obtain the audio features of the dubbing text.

[0160] Step D4: The electronic device encodes the audio features to obtain the dubbing audio.

[0161] In some embodiments of this application, the electronic device may use a vocoder to encode the above-mentioned audio features to obtain the above-mentioned dubbing audio.

[0162] In this way, the electronic device can obtain the first phoneme sequence corresponding to the dubbing text; obtain the second phoneme sequence corresponding to the target audio; generate audio features of the dubbing audio based on the first phoneme sequence, the second phoneme sequence, the dubbing timbre features and the dubbing prosody features; encode the audio features to obtain the dubbing audio, thus accurately and quickly obtaining the dubbing audio, thereby improving the efficiency of video dubbing.

[0163] Step 104: The electronic device generates a second video based on the dubbing audio and the first video.

[0164] In some embodiments of this application, the electronic device can synchronize the aforementioned dubbing audio and the first video to obtain the aforementioned second video.

[0165] In some embodiments of this application, step 104 above can be implemented by step 104a as follows:

[0166] Step 104a: In response to the fourth input to the video generation control, the electronic device generates a second video based on the dubbing audio and the first video.

[0167] In some embodiments of this application, the fourth input may include any of the following: user click input, swipe input, press input, voice input, gesture input, or other feasible inputs, which are not limited in this application.

[0168] In some embodiments of this application, the video generation control described above is used to instruct the generation of the second video described above.

[0169] In some embodiments of this application, the shape of the video generation control can be any possible shape such as a circle, rectangle, triangle, rhombus, annulus or polygon, which can be determined according to actual usage requirements. This embodiment of the invention does not limit the shape.

[0170] For example Figure 2B As shown, the dubbing window 23 may also include a video generation control 233. When the mobile phone receives a user's click input on the video generation control 233, the mobile phone can respond to the user's click input on the video generation control 233 and generate a second video based on the dubbing audio and the first video. The user's click input on the video generation control 233 is the aforementioned fourth input.

[0171] In the video processing method provided in this application embodiment, a first input is received to a first interface, the first interface including the screen of a first video; in response to the first input, dubbing timbre features and dubbing rhythm features are acquired; dubbing audio is acquired based on the dubbing timbre features and dubbing rhythm features; and a second video is generated based on the dubbing audio and the first video. Thus, the electronic device can acquire corresponding dubbing timbre features and dubbing rhythm features according to the user's input, and then generate dubbing audio based on the dubbing timbre features and dubbing rhythm features, thereby improving the flexibility of video dubbing methods. Furthermore, it eliminates the need for the user to continuously record the audio of the entire video, thereby improving the efficiency of video dubbing.

[0172] In some embodiments of this application, the first interface further includes an audio acquisition control; prior to step 101, the video processing method provided in this application may further include the following steps 106a and 106b:

[0173] Step 106a: The electronic device receives a second input to the audio acquisition control.

[0174] In some embodiments of this application, the second input may include any of the following: user click input, swipe input, press input, voice input, gesture input, or other feasible inputs, which are not limited in this application.

[0175] In some embodiments of this application, the audio acquisition control can be a video audio acquisition control or a user audio acquisition control.

[0176] In some embodiments of this application, the video audio acquisition control is used to indicate the acquisition of video audio, and the user audio acquisition control is used to indicate the acquisition of user audio.

[0177] In some embodiments of this application, the shape of the audio acquisition control can be any possible shape such as a circle, rectangle, triangle, rhombus, ring, or polygon, which can be determined according to actual usage requirements. This embodiment of the invention does not limit the shape.

[0178] Step 106b: The electronic device responds to the second input and displays the first prompt message.

[0179] In some embodiments of this application, the second prompt information is used to indicate that the target audio has been obtained, and the target audio is the audio in the first video or the user's audio.

[0180] In some embodiments of this application, when the audio acquisition control is a video audio acquisition control, the electronic device can respond to the second input, extract audio from the first video, and display a first prompt message indicating that the video audio has been acquired.

[0181] For example Figure 2B As shown, the dubbing window 23 may also include a video and audio acquisition control 234. When the mobile phone receives a user's click input on the video and audio acquisition control 234, the mobile phone can respond to the user's click input on the video and audio acquisition control 234, acquire video and audio, and, when the video and audio are acquired, combine them with... Figure 2B ,like Figure 2F As shown, the first prompt message 261, which indicates that the audio has been extracted from the first video, is displayed on the first interface 20. The user's click input on the video audio acquisition control 234 is the second input mentioned above.

[0182] It should be noted that if the phone does not capture video and audio, combined with... Figure 2B ,like Figure 2G As shown, the sixth prompt message 27 is displayed on the first interface 20. The sixth prompt message 27 is used to indicate that no video or audio has been captured.

[0183] In some embodiments of this application, when the audio acquisition control is a user audio acquisition control, the electronic device can respond to the second input, acquire user audio, and display a first prompt message indicating that user audio has been acquired.

[0184] For example Figure 2F As shown, when the sixth prompt message 27 is displayed on the first interface 20, the user can click on the user audio acquisition control 235 in the dubbing window 23 to input data. The mobile phone can respond to the user's click on the user audio acquisition control 235, combined with... Figure 2F ,like Figure 2H As shown, a guide text prompt box 28 is displayed, which includes guide text 281 and an audio recording control 282. After the user clicks and inputs information on the audio recording control 282, the mobile phone can capture the audio recorded by the user based on the guide text 281, and then combine the captured audio with the audio recorded by the user based on the guide text 281. Figure 2G ,like Figure 2I As shown, the first prompt message 262 indicating that the user's audio has been extracted is displayed on the first interface 20.

[0185] In this way, by receiving a second input to the audio acquisition control, the electronic device responds to the second input by displaying a first prompt, enabling the electronic device to display a prompt about the target audio acquired by the electronic device based on the user's operation of the audio acquisition control, thereby improving the flexibility of video dubbing.

[0186] In some embodiments of this application, after "obtaining the duration of at least one first phoneme of the dubbing text" in step a1 above, the video processing method provided in this application embodiment may further include the following steps 108a to 108e:

[0187] Step 108a: The electronic device calculates the sum of the durations of at least one first phoneme to obtain the first duration of the dubbing audio.

[0188] For example, if the duration of the first phoneme corresponding to at least one "siln i3 h ao3 sil" is 0.22s, 0.06s, 0.16s, 0.11s, 0.22s and 0.25s respectively, then the first duration of the dubbing audio is 0.22s+0.06s+0.16s, 0.11s+0.22s+0.25s, that is, the first duration of the dubbing audio is 1.04s.

[0189] Step 108b: The electronic device calculates the ratio of the first duration to the video duration of the first video.

[0190] In some embodiments of this application, the above ratio is the interpolation coefficient of the above audio feature, and the electronic device can calculate the above interpolation coefficient using the following formula:

[0191] α=L o / L t (1)

[0192] Where α represents the interpolation coefficient mentioned above, L o L represents the first duration mentioned above. t This indicates the duration of the aforementioned video.

[0193] Step 108c: If the ratio is within the first numerical range, the electronic device performs interpolation processing on the audio features based on the ratio to obtain the processed audio features.

[0194] In some embodiments of this application, the electronic device may use the following formula to interpolate audio features:

[0195] feat n =acoustic_feat [n*α] (2)

[0196] Where feat represents the processed audio feature, acoustic_feat represents the aforementioned audio feature, n is an integer from 0 to M, M is the ratio of the aforementioned video duration to the first duration, and [n*α] represents rounding n*α to the nearest integer.

[0197] It should be noted that after obtaining the processed audio features, the electronic device can use a vocoder to decode the processed audio features, obtain the dubbing audio again, and then synchronize the obtained dubbing audio with the first video to obtain the second video.

[0198] In this way, the electronic device obtains the first duration of the dubbing audio by calculating the sum of the durations of at least one first phoneme; it calculates the ratio of the first duration to the video duration of the first video, and when the ratio is within a first numerical range, it performs interpolation processing on the audio features based on the ratio to obtain the processed audio features. The audio features can be flexibly adjusted according to the ratio of the first duration to the video duration of the first video, and then the adjusted audio features can be decoded to adjust the dubbing audio, thereby improving the flexibility of video dubbing.

[0199] Step 108d: If the ratio is outside the first numerical range and the ratio is less than the minimum value of the first numerical range, the electronic device displays a second prompt message.

[0200] In some embodiments of this application, the second prompt information is used to prompt for expansion of the dubbing text.

[0201] In some embodiments of this application, if the speech rate of the dubbing audio is too fast or too slow, it will cause the sound and picture to be out of sync during dubbing. If the ratio is outside the first numerical range and the ratio is less than the minimum value of the first numerical range, it indicates that the speech rate of the dubbing audio is too fast. In this case, the electronic device can display a third prompt message to prompt the user to expand the dubbing text to generate a dubbing audio with a duration close to the first duration.

[0202] For example, the first numerical range can be 0.85 to 1.1. When the mobile phone determines that the ratio is 0.6, which is outside the first numerical range of 0.85 to 1.1, and the ratio is less than the minimum value of the first numerical range of 0.85, such as... Figure 2J As shown, a second prompt message 291 for supplementing the above-mentioned dubbing text can be displayed on the first interface 20.

[0203] Step 108e: If the ratio is outside the first numerical range and the ratio is greater than the maximum value of the first numerical range, the electronic device displays a third prompt message.

[0204] In some embodiments of this application, the aforementioned third prompt information is used to prompt for shortening the aforementioned dubbing text.

[0205] In some embodiments of this application, if the ratio is outside the first numerical range and the ratio is less than the minimum value of the first numerical range, it indicates that the speech rate of the dubbing audio is too fast. In this case, the electronic device can display a fourth prompt message to prompt the user to shorten the dubbing text in order to generate a dubbing audio with a duration close to the first duration.

[0206] For example, the first numerical range can be 0.85 to 1.1. When the mobile phone determines that the ratio is 1.5, which is outside the first numerical range of 0.85 to 1.1, and the ratio is greater than the maximum value of the first numerical range of 1.1, such as... Figure 2KAs shown, a third prompt message 292 for prompting the shortening of the aforementioned dubbing text can be displayed on the first interface 20.

[0207] Thus, when the ratio is outside the first numerical range and less than the minimum value of the first numerical range, the electronic device displays a second prompt message; when the ratio is outside the first numerical range and greater than the maximum value of the first numerical range, it displays a third prompt message. This allows the electronic device to adjust the dubbing audio based on the dubbing text re-entered by the user by viewing the second and third prompt messages, thereby improving the flexibility of the video dubbing method.

[0208] Below, in conjunction with Figure 7 This explains how electronic devices generate voice-over audio. Figure 7 This is a flowchart illustrating the dubbing audio generation method provided in an embodiment of this application. Figure 7 As shown, the method includes the following steps 701 to 707.

[0209] Step 701: Obtain the target text corresponding to the target audio, at least one second phoneme corresponding to the target text, and the duration of at least one second phoneme.

[0210] For example, an electronic device can first denoise the target audio. For audio generation tasks, excessive noise in the source audio can affect the final result. Therefore, the electronic device can use the open-source project DeepFilterNet3 to denoise the target audio and obtain clean target audio.

[0211] For example, electronic devices can use many speech recognition models. This invention uses Whisper to recognize the denoised target audio and obtain the target text corresponding to the target audio.

[0212] For example, after obtaining the target text, the electronic device can process the text: convert the numbers, symbols, etc. in the target text into corresponding Chinese / English characters; at the same time, convert polyphonic characters into correct pronunciations, and finally convert them into at least one corresponding second phoneme. For example, when the target text is "Is my mood OK today?", the at least one second phoneme corresponding to the target text is "siljin1 t ian1 x in1 q ing2 OW0 K EY1 m a5sil", where the "sil" at the beginning and end represents the silent phoneme.

[0213] For example, an electronic device may use an MFA forced alignment tool to align the target audio with at least one second phoneme according to the duration of the target audio, to obtain the duration of the second phoneme corresponding to each second phoneme, and then align the target audio with the duration of the second phoneme corresponding to each second phoneme.

[0214] Step 702: Extract timbre features from the target audio.

[0215] For example, an electronic device can perform Mel feature extraction on the target audio to obtain Mel features. The number of Mel features is consistent with the total number of frames in the target audio. For example, if the target audio duration is 1.01s, taking 10ms / frame as an example, the total number of frames in the target audio is 104, and the number of Mel features is also 104. For example, the electronic device can input the Mel features into the timbre extraction module to extract a timbre feature vector timbre_feat with dimensions (8, 246) from the Mel features; the main body of the timbre extraction module is a convolutional structure, and the structure of the timbre extraction module can be as follows: Figure 8 As shown, the timbre extraction module 80 includes a first linear and normalization module (linear&layernorm801), N cascaded convolution modules (ConvBlock802), and a second linear&layernorm803. Where N = 246.

[0216] Step 703: Extract prosodic features from the target audio.

[0217] For example, an electronic device can employ a prosody extraction module to extract prosodic features from target audio. The main structure of the prosody extraction module is similar to that of the timbre extraction module. The structure of the prosody extraction module can be described as follows: Figure 9 As shown, the prosody extraction module 90 includes a first convolutional model (ConvModule 91), a max pooling module (Max Pool 92), a second convolutional model (ConvModule 93), and a vector quantizer module (Vector Quantizer 94). The first and second convolutional models (ConvModule 91 and ConvModule 93) have the same structure, similar to the timbre extraction module. The first convolutional model (ConvModule 91) may include a third linear and normalization module (linear&layernorm 911), M cascaded convolutional modules (ConvBlock 912), and a fourth linear&layernorm 913, where M = 4. The input to the prosody extraction module is the same as that of the timbre extraction module, both being Mel features. The number of prosodic features in the output target audio is the same as the number of frames in the target audio. For example, if the target audio duration is 1.04 seconds, taking 10ms / frame as an example, the total number of frames in the target audio is 104, and the number of prosodic features in the target audio is also 104.

[0218] Step 704: Obtain at least one first phoneme sequence corresponding to the dubbing text.

[0219] It should be noted that the process of an electronic device acquiring at least one first phoneme corresponding to the dubbed text can refer to the process of acquiring at least one second phoneme corresponding to the target text. To avoid repetition, it will not be described again here.

[0220] For example, after an electronic device acquires at least one first phoneme corresponding to a dubbed text, it can input the at least one first phoneme into a duration prediction model to predict the duration of each first phoneme. Based on the duration of at least one first timbre, it determines the first quantity corresponding to each first timbre. The structure of the duration prediction is similar to that of the timbre extraction module. For example, the dubbed text can be "sound reproduction," and the at least one first phoneme corresponding to the dubbed text is "sil sh eng1 in1fu4 k e4 sil." The duration predicted by the duration model for each first phoneme is [0.24s, 0.12s, 0.17s, 0.21s, 0.08s, 0.11s, 0.06s, 0.17s, 0.27s]. Assuming the duration of a single frame of audio is 10ms, the first quantity corresponding to each first phoneme is [24, 12, 17, 21, 8, 11, 6, 17, 27]. The first phoneme sequence obtained by expanding and sorting the at least one first phoneme according to the first quantity can be as follows: Figure 10 As shown, the number below each phoneme indicates the first quantity of that phoneme.

[0221] Step 705: The electronic device determines the second phoneme sequence corresponding to the target text based on the duration of at least one second phoneme, and performs prosodic prediction on the dubbing text based on the second phoneme sequence.

[0222] For example, electronic devices can employ a prosody prediction model to predict the prosody of dubbed text. The prosody prediction model is a transformer decoder-only structure; the structure used in this invention is Llama. This model can predict the result of the nth step based on the results of the previous n-1 steps. This model structure can fully utilize the preceding context information to guide the generation of the current result.

[0223] For example, the electronic device can determine the second quantity corresponding to each second phoneme according to the duration of the at least one second phoneme, and then expand and sort the at least one second phoneme according to the second quantity to obtain the second phoneme sequence.

[0224] For example, an electronic device can extract features from the second phoneme sequence and prosodic features using an embedding module, respectively, to obtain a first feature and a second feature. The first and second features are then concatenated to obtain a first concatenated feature. The second phoneme sequence has a dimension of 1*104, and after embedding feature extraction, a first feature with a dimension of 104*127 can be extracted. The prosodic feature has a dimension of 1*104, and after embedding feature extraction, a second feature with a dimension of 104*246 can be extracted. Concatenating the first and second features yields a first concatenated feature, prompt_feat, with a dimension of 104*384.

[0225] The prosody prediction model can predict the result of the next step based on the input at the current time step. For example, in the application of a text large model, when the input is "The weather is nice today. Let's go out", the model can predict the next word "stroll". Then, "stroll" is also input into the model, that is, "The weather is nice today. Let's go out stroll", and the model can predict "stroll". Continuing to input "stroll" into the model, and finally inputting "bar" into the model.至此, the model has predicted the complete sentence "The weather is nice today. Let's go out for a stroll". Similarly, the first concatenated feature is input into the prosody prediction model to predict the first prosody feature, which is used as the prosody feature target_prosody_1 of the first phoneme in the first phoneme sequence. The target_prosody_1 and the first phoneme feature are respectively subjected to feature extraction through embedding and then concatenated to obtain the second concatenated feature target_feat. The prompt_feat and the target_feat are concatenated to obtain the first concatenated feature prompt_feat again. Here, the concatenation is performed in the dimension of the number of features. For example, the dimension of prompt_feat is [104, 384], and the dimension of target_feat is [1, 384]. After concatenation, the dimension of prompt_feat is [105, 384]. Continuing to send the concatenated prompt_feat into the prosody prediction module to predict the prosody feature target_prosody_2 of the second phoneme in the first phoneme sequence. The prosody feature and the second phoneme of 104.2 are respectively subjected to feature extraction through embedding and then concatenated to obtain a new target feat, which is then concatenated with prompt_feat and sent into the prosody prediction module for prosody prediction. This process is repeated until the prosody feature target_prosody_n of the last phoneme in the first phoneme sequence is predicted, and finally at least one predicted prosody feature t_prosody_feats = [trager_prosody_1, target_prosody_2,..., target_prosody_n] is obtained.

[0226] Step 706: Obtain a third concatenated feature based on at least one predicted prosody feature and a timbre feature, and use a decoder to process the third concatenated feature to obtain the acoustic feature of the target audio.

[0227] Exemplarily, the electronic device can concatenate each predicted prosody feature in at least one predicted prosody feature and the timbre feature to obtain at least one third concatenated feature.至此, the dubbed text can be transformed into a feature that simultaneously has the timbre and prosody of the target audio. Then, the decoder processes the third concatenated feature to decode the acoustic feature (i.e., the above audio feature) acoustic_feats.

[0228] For example, each predicted prosodic feature can have a dimension of 1*384, and the timbre feature has a dimension of 8,246. The dimension of each third concatenation feature obtained by combining each predicted prosodic feature and the timbre feature can be 8*640.

[0229] Step 707: The vocoder encodes the acoustic features to obtain the dubbing audio.

[0230] For example, electronic devices can use the open-source vocoder Vocos to encode acoustic features. Instead of directly modeling audio samples in the time domain, the Vocos vocoder generates spectral coefficients and rapidly reconstructs the audio using inverse Fourier transform, thus achieving a balance between high efficiency and high quality. The acoustic_feats obtained in step 706 are then synthesized using this vocoder to create the final dubbing audio.

[0231] The following is combined with Figure 11 This section explains how to generate dubbed videos using a mobile phone as the electronic device. Figure 11 As shown, the video processing method provided in this application may include the following steps 1101 to 1106:

[0232] Step 1101: The phone displays the first screen.

[0233] For example, the first interface can be the interface of a video dubbing application (APP), or the first interface can be the interface of the video to be dubbed (i.e., the first video mentioned above), and the first interface can include the screen to be dubbed.

[0234] For example, after opening a video dubbing app on a mobile phone and displaying the first screen, such as Figure 2B As shown, a dubbing window 23 can be displayed on the first interface 20, and the dubbing window may include a video import control 236. When the user clicks on the video import control 236, the electronic device can respond to the user's click on the video import control 236 and display at least one video to be imported on the first interface 20. When the user clicks on the first video among the at least one videos, the mobile phone can respond to the user's click on the first video and display the video screen of the first video on the first interface 20.

[0235] Step 1102: The mobile phone extracts human voice from the first video (i.e., extracts the aforementioned first audio).

[0236] For example, such as Figure 2BAs shown, when a user clicks on the video audio acquisition control 234, the mobile phone can respond to the user's click on the control 234, extract the human voice (i.e., the aforementioned first audio) from the first video, save the extracted human voice audio, and record the start and end times of the human voice. Simultaneously, as... Figure 2F As shown, the first prompt message 261 is displayed on the first interface 20, prompting the user that the extraction of the human voice from the video is complete; otherwise... Figure 2G As shown, the sixth prompt message 27 is displayed on the first interface 20, indicating that no human voice was extracted.

[0237] Step 1103: The mobile phone acquires the user's audio.

[0238] For example, such as Figure 2B As shown, when a user taps the user audio acquisition control 235 (i.e., the microphone-shaped control in Figure 2), the mobile phone can respond to the user's tap input on the user audio acquisition control 235. Figure 2F ,like Figure 2H As shown, a guide text prompt box 28 is displayed, which includes guide text 281 and an audio recording control 282. After the user clicks and inputs information on the audio recording control 282, the mobile phone can capture and save the audio recorded by the user based on the guide text 281. After saving the audio recorded by the user based on the guide text 281, it is combined with... Figure 2H ,like Figure 2I As shown, the first prompt message 262 indicating that the user's audio has been extracted is displayed on the first interface 20.

[0239] Step 1104: The mobile phone responds to the user's first input and selects the dubbing mode.

[0240] For example, the mobile phone can provide three dubbing modes, where the video voice is obtained in step 202 and the user audio is obtained in step 203.

[0241] Mode 1: User prosody + User vocal timbre

[0242] If a user likes their own voice and wants to use their own voice and pronunciation style to generate dubbing audio, then step 702 extracts the user's voice timbre, step 703 extracts the user's prosody, and uses the prosody as input to the prosody prediction model in step 705 to predict the prosody of the dubbing text, ultimately generating dubbing audio that matches the user's voice timbre and prosody.

[0243] Mode 2: User's voice + video vocal rhythm

[0244] If a user likes the pronunciation style of the voice in the video and only wants to change the timbre, then: step 702 extracts the timbre of the user's audio, step 703 extracts the prosody of the video's voice audio and uses this prosody as input to the prosody prediction module in step 705 to predict the prosody of the dubbing text, and finally generates dubbing audio that matches the user's timbre and prosody.

[0245] Mode 3: User rhythm + video voice timbre

[0246] If a user likes the timbre of the voice in the video and wants to use the timbre of the voice in the video and their own pronunciation style to generate audio, then step 702 extracts the timbre of the voice audio in the video, step 703 extracts the prosody of the user's audio, and uses the prosody as the input of the prosody prediction model in step 705 to predict the prosody of the dubbing text, and finally generates dubbing audio that matches the user's timbre and prosody.

[0247] For example, such as Figure 2B As shown, when a user clicks on the dubbing mode selection control 231, the phone responds by displaying the dubbing mode selection window 23. The displayed dubbing mode selection window 24 includes dubbing mode 1, dubbing mode 2, and dubbing mode 3. The user's click on the dubbing mode selection control 231 is the first input mentioned above. When no human voice is detected in step 702, the dubbing mode 2 and mode 3 option buttons become lighter in color and cannot be clicked; when human voice is detected, the three mode option buttons are displayed normally and can all be selected.

[0248] Step 1105: Obtain the voiceover text on your mobile phone.

[0249] For example, such as Figure 2B As shown, when the mobile phone receives a user's click input on the voiceover text acquisition control 233, it can respond to the user's click input on the voiceover text acquisition control 233, such as... Figure 2D As shown, a text prompt box 25 is displayed on the first interface 20. If human voice is extracted in step 702, the text prompt box will be filled with the extracted human voice text by default, which is obtained by the ASR tool. The user can also change it to their desired content. If no human voice is detected in step 702, then... Figure 2E As shown, text prompt box 25 displays a text prompting the user to enter.

[0250] Step 1106: Generate a dubbed video on the mobile phone (i.e., the second video mentioned above).

[0251] For example, such as Figure 2B As shown, when the mobile phone receives a user's click input on the video generation control 233, the mobile phone can respond to the user's click input on the video generation control 233 and generate a dubbed video based on the dubbed audio and the first video.

[0252] For example, when the duration of the generated dubbing audio is inconsistent with the duration of the first video or the duration of the video's human voice audio, it is necessary to extend or shorten the duration of the generated dubbing audio. The mobile phone can use nearest neighbor interpolation to adjust the duration of the dubbing audio. For example, the mobile phone can change the number of audio features by using interpolation coefficients, thereby adjusting the duration of the dubbing audio. See steps 108a to 108c above for details; to avoid repetition, they will not be repeated here.

[0253] The generated audio needs to be inserted into the video. Interpolation adjusts the speech rate; too fast or too slow a speech rate will cause the sound to be out of sync with the video. Step 704 yields the final duration of the synthesized audio. If the calculated interpolation coefficient α is outside the set range, the user is prompted to provide a new text to be synthesized to generate audio that is close to the video duration or the duration of the voiceover in the video. Research shows that an interpolation coefficient α between 0.85 and 1.1 yields the best results. When α < 0.85, the user is prompted to expand the audio text; when α > 1.1, the user is prompted to shorten the audio text. For details, please refer to [reference needed]. Figure 2J and Figure 2K This will not be elaborated upon here.

[0254] For example, a mobile phone can use the audio / video processing tool ffmpeg to merge audio and video.

[0255] For example, when there is no voice in the first video, the audio and video are directly combined using the ffmpeg tool; when there is voice in the first source video, the synthesized audio is combined with the video according to the start and end times of the voice in the video using the ffmpeg tool.

[0256] For example, to enhance the entertainment value, if the first video contains lip-syncing actions, such as interviews, speeches, or talk shows, the phone can use the MyHeyGen open-source tool to align the synthesized audio with the lip movements in the video content, generating a vivid and interesting re-dubbed video.

[0257] It should be noted that the specific implementation process of the video processing method in the above steps can be found in the relevant description of the above embodiments. To avoid repetition, this embodiment will not repeat the details here.

[0258] It should be noted that each of the above method embodiments, or various possible implementations of each method embodiment, can be executed individually or in combination of any two or more. The specific implementation can be determined according to actual usage requirements, and this application embodiment does not impose any restrictions on this.

[0259] The video processing method provided in this application can be executed by a video processing device. This application uses a video processing device executing the video processing method as an example to illustrate the video processing device provided in this application.

[0260] Figure 12 This is a schematic diagram of the structure of a video processing device 1200 provided in an embodiment of this application. The video processing device 1200 includes a receiving module 1201 and a processing module 1202.

[0261] The receiving module 1201 is used to receive a first input to the first interface, the first interface including the screen of the first video.

[0262] Processing module 1202 is used to obtain dubbing timbre features and dubbing prosody features in response to the first input;

[0263] Based on the dubbing timbre features and the dubbing rhythm features, the dubbing audio is obtained;

[0264] A second video is generated based on the dubbing audio and the first video.

[0265] In some embodiments of this application, the voice-over timbre feature is a user timbre feature, and the voice-over prosody feature is a video prosody feature;

[0266] The processing module 1202 is specifically used for:

[0267] Extract the user timbre features from the user's audio and extract the video prosodic features from the first audio, wherein the first audio is the audio in the first video;

[0268] The dubbing audio is generated based on the user's timbre features and the video prosodic features.

[0269] In some embodiments of this application, the voice-over timbre feature is a user timbre feature, and the voice-over prosody feature is a user prosody feature;

[0270] The processing module 1202 is specifically used for:

[0271] Extract the user's timbre features and prosodic features from the user's audio;

[0272] The dubbing audio is generated based on the user's timbre features and prosodic features.

[0273] In some embodiments of this application, the voice-over timbre features are video timbre features, and the voice-over prosody features are user prosody features;

[0274] The processing module 1202 is specifically used for:

[0275] Extract the video timbre features from the first audio and extract the user prosodic features from the user audio, wherein the first audio is the audio in the first video;

[0276] The dubbing audio is generated based on the video timbre features and the user prosodic features.

[0277] In some embodiments of this application, the first interface further includes an audio acquisition control;

[0278] The receiving module 1201 is also used to receive a second input to the audio acquisition control;

[0279] Combination Figure 12 ,like Figure 13 As shown, the device 1200 further includes a display module 1203, which is used to display a first prompt message in response to the second input. The first prompt message is used to prompt that a target audio has been obtained, and the target audio is the audio in the first video or the user's audio.

[0280] In some embodiments of this application, the first interface further includes a voiceover text acquisition control;

[0281] The receiving module 1201 is further configured to receive a third input to the dubbing text acquisition control before generating the dubbing audio based on the timbre features and the prosody features;

[0282] Display module 1203 is used to display dubbing text in response to the third input, wherein the dubbing text is the text corresponding to the first video or the text input by the user;

[0283] The processing module 1202 is specifically used for:

[0284] The dubbing audio is generated based on the timbre features, the prosodic features, and the dubbing text.

[0285] In some embodiments of this application, the dubbing text is the text corresponding to the first video;

[0286] The display module 1203 is specifically used for:

[0287] In response to the third input, if the first audio is detected in the first video, the dubbing text is displayed, wherein the dubbing text is the text corresponding to the first audio in the first video.

[0288] In some embodiments of this application, the voiceover text is text input by the user, and the third input includes a first sub-input and a second sub-input;

[0289] The display module 1203 is specifically used for:

[0290] In response to the first sub-input, a text input box is displayed on the first interface;

[0291] In response to a second sub-input to the text input box, the voice-over text is displayed in the text input box, the voice-over text being the text corresponding to the second sub-input.

[0292] In some embodiments of this application, the processing module 1202 is specifically used for:

[0293] Obtain the first phoneme sequence corresponding to the dubbed text;

[0294] Obtain the second phoneme sequence corresponding to the target audio, wherein the target audio is the audio in the first video or the user audio;

[0295] Based on the first phoneme sequence, the second phoneme sequence, the dubbing timbre features, and the dubbing prosody features, the audio features of the dubbing audio are generated;

[0296] The audio features are encoded to obtain the dubbing audio.

[0297] In some embodiments of this application, the processing module 1202 is specifically used for:

[0298] Obtain at least one first phoneme and at least one duration of the first phoneme from the dubbed text, wherein the duration of each first phoneme is the duration of one first phoneme.

[0299] Based on the duration of the at least one first phoneme, determine the first quantity of each first phoneme in the at least one first phoneme;

[0300] Based on the first quantity, the at least one first phoneme is sorted to obtain the first phoneme sequence.

[0301] In some embodiments of this application, the processing module 1202 is further configured to obtain the duration of at least one first phoneme of the dubbing text, and then calculate the sum of the durations of the at least one first phoneme to obtain the first duration of the dubbing audio.

[0302] Calculate the ratio of the first duration to the video duration of the first video;

[0303] If the ratio is within the first numerical range, then the audio features are interpolated based on the ratio to obtain the processed audio features;

[0304] The display module 1203 is further configured to display a second prompt message if the ratio is outside the first numerical range and the ratio is less than the minimum value of the first numerical range, the second prompt message being used to prompt for expansion of the dubbing text;

[0305] If the ratio is outside the first numerical range and the ratio is greater than the maximum value of the first numerical range, a third prompt message is displayed, which is used to prompt the user to shorten the dubbing text.

[0306] In some embodiments of this application, the processing module 1202 is specifically used for:

[0307] Based on the second phoneme sequence and the dubbing prosodic features, the first splicing feature is obtained;

[0308] The first splicing feature is input into the prosody prediction model to obtain the predicted prosody feature of the i-th phoneme in the first phoneme sequence. The first phoneme sequence includes N phonemes, where N is a positive integer and i is an integer from 1 to N.

[0309] Based on the predicted prosodic features and the i-th phoneme, the second splicing feature is obtained;

[0310] The second splicing feature and the first splicing feature are spliced ​​together to obtain the first splicing feature again, until i = N, to obtain at least one predicted prosodic feature, each predicted prosodic feature corresponding to a phoneme in the first phoneme sequence;

[0311] Based on the at least one predicted prosodic feature and the dubbing timbre feature, a third splicing feature is obtained, and the third splicing feature is decoded to generate the audio features of the dubbing text.

[0312] In the video processing apparatus provided in this application embodiment, a first input is received to a first interface, the first interface including a frame of a first video; in response to the first input, dubbing timbre features and dubbing rhythm features are acquired; dubbing audio is acquired based on the dubbing timbre features and dubbing rhythm features; and a second video is generated based on the dubbing audio and the first video. Thus, the electronic device can acquire corresponding dubbing timbre features and dubbing rhythm features according to the user's input, and then generate dubbing audio based on the dubbing timbre features and dubbing rhythm features, thereby improving the flexibility of the video dubbing method. Furthermore, it eliminates the need for the user to continuously record the audio of the entire video, thereby improving the efficiency of video dubbing.

[0313] The video processing device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device, augmented reality / virtual reality device, robot, wearable device, super mobile personal computer, netbook, or personal digital assistant, etc. It can also be a server, network attached storage (NAS), personal computer (PC), television set (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the specific devices.

[0314] The video processing device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system used.

[0315] The video processing apparatus provided in this application can implement the various processes implemented in the various embodiments of the above-described video processing method. To avoid repetition, it will not be described again here.

[0316] Optionally, such as Figure 14 As shown, this application embodiment also provides an electronic device 1400, including a processor 1401 and a memory 1402. The memory 1402 stores a program or instructions that can run on the processor 1401. When the program or instructions are executed by the processor 1401, they implement the various steps of the above-described video processing method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.

[0317] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.

[0318] Figure 15 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application.

[0319] The electronic device 1500 includes, but is not limited to, components such as: radio frequency unit 1501, network module 1502, audio output unit 1503, input unit 1504, sensor 1505, display unit 1506, user input unit 1507, interface unit 1508, memory 1509, and processor 1510.

[0320] Those skilled in the art will understand that the electronic device 1500 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 1510 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 15 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0321] The user input unit 1507 is used to receive a first input to the first interface, which includes the image of the first video.

[0322] Processor 1510 is configured to, in response to the first input, acquire dubbing timbre features and dubbing prosody features;

[0323] Based on the dubbing timbre features and the dubbing rhythm features, the dubbing audio is obtained;

[0324] A second video is generated based on the dubbing audio and the first video.

[0325] In some embodiments of this application, the voice-over timbre feature is a user timbre feature, and the voice-over prosody feature is a video prosody feature;

[0326] Processor 1510, specifically used for:

[0327] Extract the user timbre features from the user's audio and extract the video prosodic features from the first audio, wherein the first audio is the audio in the first video;

[0328] The dubbing audio is generated based on the user's timbre features and the video prosodic features.

[0329] In some embodiments of this application, the voice-over timbre feature is a user timbre feature, and the voice-over prosody feature is a user prosody feature;

[0330] Processor 1510, specifically used for:

[0331] Extract the user's timbre features and prosodic features from the user's audio;

[0332] The dubbing audio is generated based on the user's timbre features and prosodic features.

[0333] In some embodiments of this application, the voice-over timbre features are video timbre features, and the voice-over prosody features are user prosody features;

[0334] Processor 1510, specifically used for:

[0335] Extract the video timbre features from the first audio and extract the user prosodic features from the user audio, wherein the first audio is the audio in the first video;

[0336] The dubbing audio is generated based on the video timbre features and the user prosodic features.

[0337] In some embodiments of this application, the first interface further includes an audio acquisition control;

[0338] The user input unit 1507 is also configured to receive a second input to the audio acquisition control;

[0339] Display unit 1506 is configured to respond to the second input by displaying a first prompt message, the first prompt message being used to indicate that a target audio has been obtained, the target audio being the audio in the first video or the user's audio.

[0340] In some embodiments of this application, the first interface further includes a voiceover text acquisition control;

[0341] The user input unit 1507 is further configured to receive a third input to the dubbing text acquisition control before generating the dubbing audio based on the timbre features and the prosody features;

[0342] Display unit 1506 is configured to display dubbing text in response to the third input, wherein the dubbing text is the text corresponding to the first video or the text input by the user;

[0343] Processor 1510, specifically used for:

[0344] The dubbing audio is generated based on the timbre features, the prosodic features, and the dubbing text.

[0345] In some embodiments of this application, the dubbing text is the text corresponding to the first video;

[0346] Display unit 1506 is specifically used for:

[0347] In response to the third input, if the first audio is detected in the first video, the dubbing text is displayed, wherein the dubbing text is the text corresponding to the first audio in the first video.

[0348] In some embodiments of this application, the voiceover text is text input by the user, and the third input includes a first sub-input and a second sub-input;

[0349] Display unit 1506 is specifically used for:

[0350] In response to the first sub-input, a text input box is displayed on the first interface;

[0351] In response to a second sub-input to the text input box, the voice-over text is displayed in the text input box, the voice-over text being the text corresponding to the second sub-input.

[0352] In some embodiments of this application, the processor 1510 is specifically used for:

[0353] Obtain the first phoneme sequence corresponding to the dubbed text;

[0354] Obtain the second phoneme sequence corresponding to the target audio, wherein the target audio is the audio in the first video or the user audio;

[0355] Based on the first phoneme sequence, the second phoneme sequence, the dubbing timbre features, and the dubbing prosody features, the audio features of the dubbing audio are generated;

[0356] The audio features are encoded to obtain the dubbing audio.

[0357] In some embodiments of this application, the processor 1510 is specifically used for:

[0358] Obtain at least one first phoneme and at least one duration of the first phoneme from the dubbed text, wherein the duration of each first phoneme is the duration of one first phoneme.

[0359] Based on the duration of the at least one first phoneme, determine the first quantity of each first phoneme in the at least one first phoneme;

[0360] Based on the first quantity, the at least one first phoneme is sorted to obtain the first phoneme sequence.

[0361] In some embodiments of this application, the processor 1510 is further configured to, after obtaining the duration of at least one first phoneme of the dubbing text, calculate the sum of the durations of the at least one first phoneme to obtain the first duration of the dubbing audio.

[0362] Calculate the ratio of the first duration to the video duration of the first video;

[0363] If the ratio is within the first numerical range, then the audio features are interpolated based on the ratio to obtain the processed audio features;

[0364] The display unit 1506 is further configured to display a second prompt message if the ratio is outside the first numerical range and the ratio is less than the minimum value of the first numerical range, the second prompt message being used to prompt for expansion of the dubbing text;

[0365] If the ratio is outside the first numerical range and the ratio is greater than the maximum value of the first numerical range, a third prompt message is displayed, which is used to prompt the user to shorten the dubbing text.

[0366] In some embodiments of this application, the processor 1510 is specifically used for:

[0367] Based on the second phoneme sequence and the dubbing prosodic features, the first splicing feature is obtained;

[0368] The first splicing feature is input into the prosody prediction model to obtain the predicted prosody feature of the i-th phoneme in the first phoneme sequence. The first phoneme sequence includes N phonemes, where N is a positive integer and i is an integer from 1 to N.

[0369] Based on the predicted prosodic features and the i-th phoneme, the second splicing feature is obtained;

[0370] The second splicing feature and the first splicing feature are spliced ​​together to obtain the first splicing feature again, until i = N, to obtain at least one predicted prosodic feature, each predicted prosodic feature corresponding to a phoneme in the first phoneme sequence;

[0371] Based on the at least one predicted prosodic feature and the dubbing timbre feature, a third splicing feature is obtained, and the third splicing feature is decoded to generate the audio features of the dubbing text.

[0372] In the electronic device provided in this application embodiment, a first input is received to a first interface, the first interface including a frame of a first video; in response to the first input, dubbing timbre features and dubbing rhythm features are acquired; dubbing audio is acquired based on the dubbing timbre features and dubbing rhythm features; and a second video is generated based on the dubbing audio and the first video. Thus, the electronic device can acquire corresponding dubbing timbre features and dubbing rhythm features according to the user's input, and then generate dubbing audio based on the dubbing timbre features and dubbing rhythm features, thereby improving the flexibility of the video dubbing method. Furthermore, it eliminates the need for the user to continuously record the audio of the entire video, thereby improving the efficiency of video dubbing.

[0373] It should be understood that, in this embodiment, the input unit 1504 may include a graphics processing unit (GPU) 15041 and a microphone 15042. The GPU 15041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 1506 may include a display panel 15061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 1507 includes a touch panel 15071 and at least one of other input devices 15072. The touch panel 15071 is also called a touch screen. The touch panel 15071 may include a touch detection device and a touch controller. Other input devices 15072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.

[0374] The memory 1509 can be used to store software programs and various data. The memory 1509 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 1509 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 1509 in this embodiment includes, but is not limited to, these and any other suitable types of memory.

[0375] Processor 1510 may include one or more processing units; optionally, processor 1510 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 1510.

[0376] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described video processing method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0377] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0378] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above video processing method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0379] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0380] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the video processing method embodiments described above, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0381] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0382] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0383] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A video processing method, characterized in that, The method includes: Receive a first input to a first interface, the first interface including a frame of a first video; In response to the first input, the voice-over timbre features and voice-over prosody features are obtained; Based on the dubbing timbre features and the dubbing rhythm features, the dubbing audio is obtained; A second video is generated based on the dubbing audio and the first video; The step of obtaining the dubbing audio based on the dubbing timbre features and the dubbing rhythm features includes: Based on the second phoneme sequence and the dubbing prosody features, a first splicing feature is obtained; wherein, the second phoneme sequence is the phoneme sequence corresponding to the target audio, and the target audio is the audio in the first video or the user's audio; The first splicing feature is input into the prosody prediction model to obtain the predicted prosody feature of the i-th phoneme in the first phoneme sequence. The first phoneme sequence includes N phonemes, where N is a positive integer and i is an integer from 1 to N. The first phoneme sequence is the phoneme sequence corresponding to the dubbing text, and the dubbing text is the text corresponding to the first video or the text input by the user. Based on the predicted prosodic features and the i-th phoneme, the second splicing feature is obtained; The second splicing feature and the first splicing feature are spliced ​​together to obtain the first splicing feature again, until i=N, to obtain at least one predicted prosodic feature, each predicted prosodic feature corresponding to a phoneme in the first phoneme sequence; Based on the at least one predicted prosodic feature and the dubbing timbre feature, a third splicing feature is obtained, and the third splicing feature is decoded to generate the audio features of the dubbing text; The audio features are encoded to obtain the dubbing audio.

2. The method according to claim 1, characterized in that, The voice-over timbre feature is a user timbre feature, and the voice-over prosody feature is a video prosody feature; The step of obtaining the dubbing audio based on the dubbing timbre features and the dubbing rhythm features includes: Extract the user timbre features from the user's audio and extract the video prosodic features from the first audio, wherein the first audio is the audio in the first video; The dubbing audio is generated based on the user's timbre features and the video prosodic features.

3. The method according to claim 1, characterized in that, The voice-over timbre feature is the user's timbre feature, and the voice-over prosody feature is the user's prosody feature; The step of obtaining the dubbing audio based on the dubbing timbre features and the dubbing rhythm features includes: Extract the user's timbre features and prosodic features from the user's audio; The dubbing audio is generated based on the user's timbre features and prosodic features.

4. The method according to claim 1, characterized in that, The voice-over timbre features are video timbre features, and the voice-over prosody features are user prosody features; The step of obtaining the dubbing audio based on the dubbing timbre features and the dubbing rhythm features includes: Extract the video timbre features from the first audio and extract the user prosodic features from the user audio, wherein the first audio is the audio in the first video; The dubbing audio is generated based on the video timbre features and the user prosodic features.

5. The method according to any one of claims 2 to 4, characterized in that, The first interface also includes an audio acquisition control; The method further includes: Receive a second input to the audio acquisition control; In response to the second input, a first prompt message is displayed, which indicates that the target audio has been obtained. The target audio is either the audio in the first video or the user's audio.

6. The method according to claim 1, characterized in that, The first interface also includes a voiceover text acquisition control; Before obtaining the dubbing audio based on the dubbing timbre features and the dubbing prosody features, the method further includes: Receive a third input to the voiceover text acquisition control; In response to the third input, the dubbing text is displayed; The step of obtaining the dubbing audio based on the dubbing timbre features and the dubbing rhythm features includes: The dubbing audio is generated based on the dubbing timbre features, the dubbing rhythm features, and the dubbing text.

7. The method according to claim 6, characterized in that, The dubbing text is the text corresponding to the first video; The response to the third input, displaying the dubbing text, includes: In response to the third input, if the first audio is detected in the first video, the dubbing text is displayed, wherein the dubbing text is the text corresponding to the first audio in the first video.

8. The method according to claim 6, characterized in that, The voiceover text is text input by the user, and the third input includes a first sub-input and a second sub-input; The response to the third input, displaying the dubbing text, includes: In response to the first sub-input, a text input box is displayed on the first interface; In response to a second sub-input to the text input box, the voice-over text is displayed in the text input box, the voice-over text being the text corresponding to the second sub-input.

9. The method according to claim 6, characterized in that, The method further includes: Obtain the first phoneme sequence corresponding to the dubbed text; Obtain the second phoneme sequence corresponding to the target audio.

10. The method according to claim 9, characterized in that, The step of obtaining the first phoneme sequence corresponding to the dubbed text includes: Obtain at least one first phoneme and at least one duration of the first phoneme from the dubbed text, wherein the duration of each first phoneme is the duration of one first phoneme. Based on the duration of the at least one first phoneme, determine the first quantity of each first phoneme in the at least one first phoneme; Based on the first quantity, the at least one first phoneme is sorted to obtain the first phoneme sequence.

11. The method according to claim 10, characterized in that, After obtaining the duration of at least one first phoneme of the dubbed text, the method further includes: The first duration of the dubbing audio is obtained by summing the durations of the at least one first phoneme. Calculate the ratio of the first duration to the video duration of the first video; If the ratio is within the first numerical range, then the audio features are interpolated based on the ratio to obtain the processed audio features; If the ratio is outside the first numerical range and the ratio is less than the minimum value of the first numerical range, a second prompt message is displayed, which is used to prompt for expansion of the dubbing text; If the ratio is outside the first numerical range and the ratio is greater than the maximum value of the first numerical range, a third prompt message is displayed, which is used to prompt the user to shorten the dubbing text.

12. A video processing apparatus, characterized in that, The device includes: A receiving module is used to receive a first input to a first interface, the first interface including a frame of a first video. The processing module is used to obtain the dubbing timbre features and dubbing prosody features in response to the first input; Based on the second phoneme sequence and the dubbing prosody features, a first splicing feature is obtained; wherein, the second phoneme sequence is the phoneme sequence corresponding to the target audio, and the target audio is the audio in the first video or the user's audio; The first splicing feature is input into the prosody prediction model to obtain the predicted prosody feature of the i-th phoneme in the first phoneme sequence. The first phoneme sequence includes N phonemes, where N is a positive integer and i is an integer from 1 to N. The first phoneme sequence is the phoneme sequence corresponding to the dubbing text, and the dubbing text is the text corresponding to the first video or the text input by the user. Based on the predicted prosodic features and the i-th phoneme, the second splicing feature is obtained; The second splicing feature and the first splicing feature are spliced ​​together to obtain the first splicing feature again, until i=N, to obtain at least one predicted prosodic feature, each predicted prosodic feature corresponding to a phoneme in the first phoneme sequence; Based on the at least one predicted prosodic feature and the dubbing timbre feature, a third splicing feature is obtained, and the third splicing feature is decoded to generate the audio features of the dubbing text; The audio features are encoded to obtain the dubbing audio; A second video is generated based on the dubbing audio and the first video.

13. The apparatus according to claim 12, characterized in that, The voice-over timbre feature is a user timbre feature, and the voice-over prosody feature is a video prosody feature; The processing module is specifically used for: Extract the user timbre features from the user's audio and extract the video prosodic features from the first audio, wherein the first audio is the audio in the first video; The dubbing audio is generated based on the user's timbre features and the video prosodic features.

14. The apparatus according to claim 12, characterized in that, The voice-over timbre feature is the user's timbre feature, and the voice-over prosody feature is the user's prosody feature; The processing module is specifically used for: Extract the user's timbre features and prosodic features from the user's audio; The dubbing audio is generated based on the user's timbre features and prosodic features.

15. The apparatus according to claim 12, characterized in that, The voice-over timbre features are video timbre features, and the voice-over prosody features are user prosody features; The processing module is specifically used for: Extract the video timbre features from the first audio and extract the user prosodic features from the user audio, wherein the first audio is the audio in the first video; The dubbing audio is generated based on the video timbre features and the user prosodic features.

16. The apparatus according to claim 12, characterized in that, The first interface also includes a voiceover text acquisition control; The receiving module is further configured to receive a third input to the dubbing text acquisition control before generating the dubbing audio based on the timbre features and the dubbing rhythm features; The device further includes: a display module for responding to the third input; The processing module is specifically used for: The dubbing audio is generated based on the timbre features, the dubbing rhythm features, and the dubbing text.

17. The apparatus according to claim 16, characterized in that, The processing module is further configured to: Obtain the first phoneme sequence corresponding to the dubbed text; Obtain the second phoneme sequence corresponding to the target audio.

18. An electronic device, characterized in that, It includes a processor and a memory, the memory storing a program or instructions that can run on the processor, the program or instructions being executed by the processor to implement the steps of the video processing method as described in any one of claims 1 to 11.

19. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the video processing method as described in any one of claims 1 to 11.