Video data processing methods, electronic devices and readable media

By recognizing and fusing video audio with background music, the problem of audio loss in the one-click video creation function is solved, ensuring that the original video audio is preserved and highlighted in short videos, thus improving the user experience.

CN119233037BActive Publication Date: 2025-10-28HONOR DEVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310794342.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-29
Publication Date
2025-10-28
Estimated Expiration
2043-06-29

AI Technical Summary

Technical Problem

In existing technologies, short videos generated by electronic devices using the one-click video creation function do not retain the original audio selected by the user, resulting in audio loss.

Method used

By recognizing video and audio through electronic devices and performing scene and sound event recognition, combined with background music, video and audio are merged to generate short videos with one click, while retaining and highlighting some audio in the video.

Benefits of technology

It enables the preservation and highlighting of some audio portions of the video during the one-click short video generation process, thus improving the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119233037B_ABST
    Figure CN119233037B_ABST
Patent Text Reader

Abstract

This application provides a video data processing method, an electronic device, and a readable medium. In the video data processing method, the electronic device displays a first interface showing a thumbnail of a first video cover and a first button. The first button can be understood as a trigger button for the electronic device to generate a short video using a one-click video-to-video function. After the user clicks the first button on the first interface, the electronic device can play a second video. The second video includes at least a portion of the audio and background music of the first video, thus enabling the user to retain at least a portion of the original video's audio in the short video generated using the one-click video-to-video function. After the user clicks the first button, the electronic device obtains the second video, which includes at least a portion of the first video's audio and background music. Furthermore, the application enables the user to automatically mix the audio and background music of the first video with a single click.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a method for processing video data, an electronic device, a computer program product, and a computer-readable storage medium. Background Technology

[0002] The one-click video creation function means that after a user selects a photo or video, they can click the one-click video creation button, and the electronic device will automatically combine the user's selected photo or video with the special effects and background music configured on the electronic device to create a short video.

[0003] However, the short videos generated by electronic devices using the one-click video creation function do not retain the original audio of the user-selected video, resulting in the loss of audio from the user-selected video. Summary of the Invention

[0004] This application provides a video data processing method, electronic device, computer program product, and computer-readable storage medium, with the aim of enabling the generation of short videos using the one-click blockbuster function to retain at least some of the audio from the original video.

[0005] To achieve the above objectives, this application provides the following technical solution:

[0006] In a first aspect, this application provides a method for processing video data, comprising: an electronic device displaying a first interface, the first interface displaying a thumbnail of a first video cover and a first button; the electronic device, in response to a user clicking the first button, displaying a second interface, the second interface being a playback interface for a second video, the second video including at least a portion of the audio and background music of the first video.

[0007] As can be seen from the above: The electronic device displays a first interface, which shows a thumbnail of the cover of the first video and a first button. This first button can be understood as a trigger button for the electronic device to generate a short video using the one-click video-to-video function. After the user clicks the first button on the first interface, the electronic device can play a second video. The second video includes at least part of the audio and background music of the first video, thus enabling the user to retain at least some of the original video's audio in the short video generated using the one-click video-to-video function. Clicking the first button also allows the electronic device to automatically mix the audio and background music of the first video with a single click.

[0008] In one possible implementation, before displaying the second interface, the device further includes displaying a third interface, which is a buffer interface for generating the second video.

[0009] In one possible implementation, before displaying the second interface, the process further includes: configuring background music on the electronic device, the background music matching the first video; if the electronic device recognizes that the audio of the first video includes human voices, then it merges the audio of the first video and the background music to obtain merged audio; and the electronic device merges the image of the first video and the merged audio to obtain the second video.

[0010] In this possible implementation, the electronic device configures background music that matches the style of the first video. This ensures that the audio and background music styles are consistent during the fusion process, avoiding any incongruity. If the electronic device recognizes that the first video retains human voices in its audio, then fusing the audio and background music of the first video allows the second video to retain the valid audio from the first video, preventing invalid audio from being retained in the second video.

[0011] In one possible implementation, before the electronic device merges the audio and background music of the first video to obtain merged audio, the method further includes: the electronic device performing noise reduction processing on the audio of the first video to obtain noise-reduced audio; wherein, the electronic device merging the audio and background music of the first video to obtain merged audio includes: the electronic device merging the noise-reduced audio and background music of the first video to obtain merged audio.

[0012] In this possible implementation, before the electronic device merges the audio and background music of the first video, it first performs noise reduction processing on the music of the first video to obtain clean vocals in the first video, thus ensuring the effect after the audio and background music of the first video are merged.

[0013] In one possible implementation, the electronic device identifies that the audio of the first video includes human voice, including: the electronic device performs sound event recognition on the audio of the first video to obtain sound event recognition results; the electronic device determines that the sound event recognition results include human voice events.

[0014] In one possible implementation, sound event recognition includes speech recognition.

[0015] In one possible implementation, after configuring background music, the electronic device further includes: if the electronic device recognizes that the audio of the first video includes key audio, then it merges the audio of the first video and the background music to obtain merged audio; the electronic device merges the image of the first video and the merged audio to obtain a second video.

[0016] In this possible implementation, if the audio of the first video includes key audio, then fusing the audio of the first video with the background music can ensure that the key audio of the first video is preserved in the second video.

[0017] In one possible implementation, the process of the electronic device recognizing that the audio of the first video includes key audio further includes: the electronic device extracting features from the audio of the first video to obtain long-term audio features, and performing scene recognition on the long-term audio features to obtain a scene recognition result, wherein the scene recognition result indicates that the first video belongs to a first scene; the electronic device recognizing that the audio of the first video includes key audio includes: the electronic device recognizing that the audio of the first video includes key audio of the first scene.

[0018] In one possible implementation, the electronic device is configured with multiple scenarios, and the electronic device identifies key audio in the audio of the first video, including: the electronic device identifies key audio belonging to the first scenario in the audio of the first video.

[0019] In one possible implementation, the electronic device identifies that the audio of the first video includes key audio, including: the electronic device performs sound event recognition on the audio of the first video to obtain sound event recognition results; the electronic device determines that the sound event recognition results include sound events of key audio.

[0020] In one possible implementation, the key audio may include at least human voice, or the key audio may not include human voice.

[0021] In one possible implementation, the audio and background music of the first video are merged to obtain merged audio, which includes: merging the audio and background music of the first video according to the fusion ratio of audio amplitude in corresponding time periods to obtain merged audio.

[0022] In one possible implementation, the audio amplitude blending ratio is used to indicate that the amplitude of the audio from the first video is greater than the amplitude of the background music in the blended audio. In this possible implementation, the blending ratio of the audio from the first video and the background music indicates that the amplitude of the audio from the first video is greater than the amplitude of the background music, which can make the audio from the first video stand out more relative to the background music in the second video.

[0023] In one possible implementation, the audio and background music of the first video are merged to obtain merged audio, which includes: merging the audio and background music of the first video according to the principle that the energy of the audio of the first video is substantially the same as the energy of the background music, thereby achieving that the volume of the background music and the audio of the first video are maintained at substantially the same level in the second video.

[0024] In one possible implementation, the audio and background music of the first video are fused according to the principle that the energy of the audio of the first video and the energy of the background music are substantially the same to obtain fused audio. This includes: taking the product of the audio energy of the first video and a first coefficient as the audio energy of the first video in the fused audio, and taking the product of the background music energy and a second coefficient as the background music energy in the fused audio. The first coefficient is the proportion of the background music energy to the sum of the audio energy of the first video and the energy of the background music, and the second coefficient is the proportion of the audio energy of the first video to the sum of the audio energy of the first video and the energy of the background music.

[0025] In one possible implementation, the audio and background music of the first video are merged to obtain merged audio, including: the key audio includes at least human voices, and the audio and background music of the first video are merged according to a first ratio of audio amplitude in a corresponding time period to obtain merged audio; or, the key audio does not include human voices, and the audio and background music of the first video are merged according to a second ratio of audio amplitude in a corresponding time period to obtain merged audio; wherein, both the first ratio and the second ratio of audio amplitude are used to indicate the proportional relationship between the audio and background music of the first video; the first ratio of audio amplitude is greater than the second ratio of audio amplitude, so that the audio of the first video includes human voices, which can better highlight human voices.

[0026] In a second aspect, this application provides an electronic device, including: one or more processors, a memory, and a display screen; the memory and the display screen are coupled to the one or more processors, the memory is used to store a computer program, the computer program including computer instructions, and when the one or more processors execute the computer instructions, the electronic device performs a video data processing method as described in any one of the first aspects.

[0027] Thirdly, this application provides a computer-readable storage medium for storing a computer program, which, when executed, is specifically used to implement the video data processing method as described in any one of the first aspects.

[0028] Fourthly, this application provides a computer program product that, when run on a computer, causes the computer to perform the video data processing method as described in any one of the first aspects. Attached Figure Description

[0029] Figure 1 A diagram illustrating the one-click generation process of short videos provided in this application embodiment;

[0030] Figure 2 A hardware structure diagram of the electronic device provided in the embodiments of this application;

[0031] Figure 3 A flowchart illustrating a video data processing method provided in an embodiment of this application;

[0032] Figure 4 An illustration of scene recognition using a neural network model provided in an embodiment of this application;

[0033] Figure 5 The illustration shows the scenarios and corresponding target sound events provided in the embodiments of this application;

[0034] Figure 6 A flowchart of a video data processing method provided in another embodiment of this application. Detailed Implementation

[0035] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. The terminology used in the following embodiments is for the purpose of describing specific embodiments only and is not intended to be a limitation of this application. As used in the specification and appended claims of this application, the singular expressions "a," "an," "the," "the," "the," and "this" are intended to also include expressions such as "one or more," unless the context clearly indicates otherwise. It should also be understood that in the embodiments of this application, "one or more" refers to one, two, or more; "and / or" describes the relationship between related objects, indicating that three relationships may exist; for example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship.

[0036] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0037] The "multiple" mentioned in the embodiments of this application refers to two or more. It should be noted that in the description of the embodiments of this application, terms such as "first" and "second" are used only for the purpose of distinguishing descriptions and should not be construed as indicating or implying relative importance, nor should they be construed as indicating or implying order.

[0038] The one-click video creation function, also known as the one-click blockbuster function, refers to a feature where, after a user selects a photo or video and clicks the one-click video creation button (or one-click blockbuster button), the electronic device automatically synthesizes the selected photo or video along with the device's configured effects and background music to create a short video. Effects refer to special effects that can be added to video frames, supported by the source material, such as snowflakes, fireworks, and other animation effects, as well as filters, stickers, and borders. In some embodiments, effects may also be referred to as styles or style themes.

[0039] However, the short videos generated by electronic devices using the one-click video creation function do not retain the original sound of the user-selected video, only including background music, resulting in the loss of audio from the user-selected video.

[0040] Based on this, the embodiments of this application provide a video data processing method that synthesizes data such as images, videos, background music and special effects to obtain one-click short videos. In the process, at least some of the audio in the video can be retained, and specific audio in the video can be highlighted in the one-click short videos.

[0041] The following combination Figure 1 Taking a mobile phone as an example, this application describes the process of generating short videos using the one-click blockbuster function provided in the embodiments.

[0042] When a user opens the Gallery app, the interface displays multiple locally saved videos and images. The user can select any image and / or video displayed to generate a short video with a single click. For example, the Gallery app interface might look like this: Figure 1 As shown in (a), this includes the video's cover thumbnail and image, which the user can view. Figure 1 In the interface shown in (a), long press the cover thumbnail 101, image 102 and image 103 of the video respectively to select video 101, image 102 and image 103 as the material for one-click short video creation. Figure 1 (b) shows the interface where the user selects video 101, image 102, and image 103. The interface also displays a "One-Click Blockbuster" button 104, which triggers the phone to generate a short video using video 101, image 102, image 103, and special effects.

[0043] Users Figure 1 As shown in (b), when the "One-Click Blockbuster" button 104 is clicked, the mobile phone will match the background music according to the style of the video 101, image 102 and image 103, and can also match special effects to synthesize the video 101, image 102 and image 103, as well as the special effects and background music configured by the mobile phone, to obtain a short video with one click. Figure 1Image (c) shows the waiting screen (or buffering interface) for generating a short video with one click on the phone. After generating the short video with one click, the phone can... Figure 1 As shown in (c) and (d), the short video plays automatically.

[0044] In some embodiments, the phone may not display [the message / instructions]. Figure 1 The screen shown in (c) shows that after the user clicks the "One-Click Blockbuster" button (104), the phone generates a short video with one click, and as shown... Figure 1 As shown in (c) and (d), the short video plays automatically.

[0045] The one-click short video generated by the mobile phone retains part of the audio in video 101. In some embodiments, the human voice in video 101 can also be highlighted in the one-click short video.

[0046] In some embodiments, users can select at least one video in the gallery application interface, and then use their mobile phones to combine the video, effects, and background music to create a short video with one click. In other embodiments, users can also select at least one image in the gallery application interface, and then use their mobile phones to combine the image, effects, and background music to create a short video with one click.

[0047] The video data processing method provided in this application embodiment involves synthesizing a user-selected image, video, and background music and effects configured on the mobile phone. In this process, the mobile phone needs to merge the audio and background music in the video. Therefore, to facilitate the explanation of the process of merging the audio and background music in the video, the following content of this application embodiment will be described using the example of a user selecting an image and at least one video in the gallery application interface.

[0048] The video data processing method provided in this application embodiment can also be applied to electronic devices such as tablet computers, personal digital assistants (PDAs), desktop, laptop, and notebook computers, ultra-mobile personal computers (UMPCs), handheld computers, netbooks, and wearable devices.

[0049] Taking mobile phones as an example, Figure 2 This is an example of the composition of an electronic device provided in an embodiment of this application. For example... Figure 2 As shown, the electronic device 100 may include a processor 110, an internal memory 120, a camera 130, a display screen 140, a mobile communication module 150, a wireless communication module 160, an audio module 170, and a sensor module 180, etc.

[0050] It is understood that the structure illustrated in this embodiment does not constitute a specific limitation on the electronic device 100. In other embodiments, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0051] Processor 110 may include one or more processing units, such as application processors (APs), modem processors, graphics processing units (GPUs), image signal processors (ISPs), controllers, video codecs, digital signal processors (DSPs), baseband processors, smart sensor hubs, and / or neural network processing units (NPUs). These different processing units may be independent devices or integrated into one or more processors.

[0052] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0053] Internal memory 120 can be used to store executable program code, including instructions. Processor 110 executes various functional applications and data processing of electronic device 100 by running the instructions stored in internal memory 120. Internal memory 120 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback, image playback, etc.), etc. The data storage area may store data created during the use of electronic device 100 (such as audio data, phonebook, etc.). Furthermore, internal memory 120 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc. Processor 110 executes various functional applications and data processing of electronic device 100 by running instructions stored in internal memory 120 and / or instructions stored in memory located within the processor.

[0054] In some embodiments, the internal memory 120 stores instructions for performing video data processing methods. The processor 110 can execute the instructions stored in the internal memory 120 to synthesize data such as images, videos, background music, and special effects to obtain a one-click short video. Furthermore, at least a portion of the audio from the video can be retained during the generation of the one-click short video, and specific audio from the video can be highlighted within the one-click short video.

[0055] Electronic devices implement display functions through a GPU, a display screen 140, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 140 and the application processor. The GPU performs mathematical and geometric calculations for image rendering. The processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.

[0056] The display screen 140 is used to display images, video interfaces, etc. The display screen 140 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device may include one or N displays 140, where N is a positive integer greater than 1.

[0057] The wireless communication function of electronic device 100 can be realized through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor, etc.

[0058] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover one or more communication frequency bands. Different antennas can also be multiplexed to improve antenna utilization. For example, antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in conjunction with tuning switches.

[0059] The mobile communication module 150 can provide solutions for wireless communication, including 2G / 3G / 4G / 5G, applied to the electronic device 100. The mobile communication module 150 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves via antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The mobile communication module 110 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via antenna 1. In some embodiments, at least some functional modules of the mobile communication module 150 may be housed in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 150 and at least some modules of the processor 110 may be housed in the same device.

[0060] The wireless communication module 160 can provide solutions for wireless communication applications on the electronic device 100, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 110. The wireless communication module 150 can also receive signals to be transmitted from processor 110, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.

[0061] Electronic devices can implement audio functions such as music playback and recording through audio modules 170, speakers 170A, receivers 170B, microphones 170C, headphone jacks 170D, and application processors.

[0062] The audio module 170 is used to convert digital audio information into analog audio signals for output, and also to convert analog audio input into digital audio signals. The audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 170 may be located in the processor 110, or some functional modules of the audio module 170 may be located in the processor 110.

[0063] The speaker 170A, also known as a "loudspeaker," is used to convert audio electrical signals into sound signals. Electronic devices can listen to music or make hands-free calls through the speaker 170A.

[0064] The receiver 170B, also known as the "earpiece," is used to convert audio electrical signals into sound signals. When an electronic device answers a phone call or voice message, the receiver 170B can be brought close to the ear to hear the voice.

[0065] Microphone 170C, also known as a "microphone" or "voice transducer," is used to convert sound signals into electrical signals. When making a phone call or sending a voice message, the user can speak by bringing their mouth close to microphone 170C, inputting the sound signal into microphone 170C. Electronic devices can have at least one microphone 170C. In some embodiments, electronic devices can have two microphones 170C, which, in addition to collecting sound signals, can also perform noise reduction. In other embodiments, electronic devices can have three, four, or more microphones 170C, enabling sound signal collection, noise reduction, sound source identification, and directional recording, among other functions.

[0066] The 170D headphone jack is used to connect wired headphones. The 170D headphone jack can be a USB interface or a 3.5mm Open Mobile Terminal Platform (OMTP) standard interface, a CTIA (Cellular Telecommunications Industry Association of the USA) standard interface.

[0067] In sensor module 180, pressure sensor 180A is used to sense pressure signals and convert them into electrical signals. In some embodiments, pressure sensor 180A can be disposed on display screen 140. There are many types of pressure sensors 180A, such as resistive pressure sensors, inductive pressure sensors, and capacitive pressure sensors. A capacitive pressure sensor may include at least two parallel plates with conductive material. When a force is applied to pressure sensor 180A, the capacitance between the electrodes changes. The electronic device determines the pressure intensity based on the change in capacitance. When a touch operation is applied to display screen 140, the electronic device detects the touch operation intensity based on pressure sensor 180A. The electronic device can also calculate the touch position based on the detection signal from pressure sensor 180A.

[0068] Touch sensor 180B, also known as a "touch device," can be disposed on display screen 140. The touch sensor 180B and display screen 140 together form a touchscreen, also known as a "touchscreen." Touch sensor 180B is used to detect touch operations applied to or near it. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through display screen 140. In other embodiments, touch sensor 180B may also be disposed on the surface of the electronic device, in a different location than display screen 140.

[0069] The 180C accelerometer can detect the magnitude of acceleration in various directions (typically three axes) of electronic devices. When the electronic device is stationary, it can detect the magnitude and direction of gravity, and can also be used to identify the attitude of the electronic device.

[0070] The gyroscope sensor 180D can be used to determine the motion attitude of an electronic device. In some embodiments, the angular velocity of the electronic device about three axes (i.e., the x, y, and z axes) can be determined by the gyroscope sensor 180D.

[0071] The technical solutions involved in the following embodiments can all be implemented in electronic devices with the above-described hardware architecture.

[0072] The following combination Figure 3 This paper introduces the process of generating short videos with one click using video data processing methods on electronic devices.

[0073] like Figure 3 As shown in the embodiment of this application, a method for processing video data includes:

[0074] S301. The electronic device acquires the video and image selected by the user.

[0075] When a user opens the Gallery app, their electronic device displays something like... Figure 1 As shown in (a), the interface of the gallery application is displayed, which includes multiple videos and multiple images. The user long-presses video 101, image 102, and image 103 to select the video and image.

[0076] In response to a user's long-press action on a video or image, the electronic device acquires the video or image selected by the user. In some embodiments, acquiring the video and image may refer to acquiring the video stream data of the video and the image data of the image.

[0077] S302. Background music that matches the style of images and videos in electronic devices.

[0078] The style of an image can also refer to its subject matter or category. Similarly, the style of a video can refer to its subject matter or category. For example, Figure 1 Images 102 and 103 shown in (a) belong to a football match scene, and their style is sports-themed. Similarly, video 101 also has a sports-themed style. As another example, if the images and videos selected by the user include food-related elements such as dishes and food, then the images and videos will have a food-themed style.

[0079] In some embodiments, the background music configured on the electronic device can be one or more.

[0080] In some embodiments, configuring background music that matches the style of images and videos in an electronic device can refer to the rhythm of the background music matching the style of the images and videos. In one possible implementation, if the style of the images and videos is lively and vibrant, then more energetic background music is configured; if the style of the images and videos is quiet and soothing, then more calming background music is configured. For example, Figure 1 In example (a), video 101, image 102, and image 103 are in a sporty style, belonging to the lively and bustling category, and the background music configured on the electronic device is also in a dynamic style. As another example, if the image and video selected by the user are in a food-themed style, belonging to the quiet category, the background music configured on the electronic device is also in a more soothing style.

[0081] In some embodiments, the electronic device identifies the style of the image and the style of the video, respectively, and matches background music that matches both the image style and the video style. In other embodiments, the electronic device may also match background music that matches either the image style or the video style. The background music matched by the electronic device in this step can be used as background music for a short video created with a single click.

[0082] In some embodiments, the electronic device can identify key features of an image and use these key features to obtain the image's style. In other embodiments, the electronic device can acquire multiple frames of a video, identify key features in each frame, use the key features of one frame to obtain the image's style, and then comprehensively evaluate the styles of the multiple frames to obtain the video's style. In one possible implementation, the electronic device is equipped with an image recognition model that can identify key features of an image and obtain the image's style based on these key features.

[0083] In some embodiments, in addition to configuring background music that matches the style of the images and videos, the electronic device may also configure special effects that match the style of the images and videos.

[0084] In some embodiments, the electronic device has multiple special effects templates. Each special effects template may include special effects content such as multiple audio effects and video effects, and also includes descriptive information indicating the time of addition of each effect and its position in the video frame. The special effects content and descriptive information of each special effects template are indicated by the template information of the special effects template. The template information of each special effects template may be stored in JSON format in a public resource path. Each special effects template configured by the electronic device corresponds to an ID, which can uniquely identify the special effects template. The template information of the special effects template can be determined by the ID of the special effects template. The special effects configured by the electronic device to match the style of images and videos can refer to the matching special effects templates.

[0085] S303, Electronic device acquires a video.

[0086] The electronic device acquires a video from the video obtained in step S301 and performs the processing described in steps S304 to S310. In some embodiments, if the electronic device acquires the image obtained in step S301, the electronic device skips that image and does not process it.

[0087] In some embodiments, "acquiring a video" by an electronic device can refer to acquiring video stream data of a video.

[0088] S304. The electronic device separates the image and audio of the video.

[0089] In some embodiments, the electronic device may utilize FFmpeg's multi-output mode to achieve audio-visual separation of video, and then the electronic device may extract object elements and timelines from the separated images and audio.

[0090] The electronic device acquires video stream data of a video in step S303, including an image frame sequence and an audio stream. The image frame sequence includes multiple images. The electronic device needs to perform steps S305 and S306 on the audio stream in the video; therefore, the electronic device needs to first separate the image frame sequence and the audio stream from the video stream data.

[0091] S305. The electronic device performs scene recognition on the audio separated from the video and obtains the scene recognition result.

[0092] In some embodiments, the electronic device may be configured with a neural network model, referred to as a scene recognition network model, which is used to perform scene recognition on the audio stream. The scene recognition network model can be trained. Figure 4 (a) shows the training process of the scene recognition network model.

[0093] like Figure 4 As shown in (a), feature extraction is performed on the sample audio to obtain audio features. To improve the training effect of the scene recognition network model, the audio features can be augmented. The augmented audio features are then input into the neural network model, and the scene recognition network model is trained based on the loss function. The trained scene recognition network model is able to identify the scene to which the input audio belongs.

[0094] In some embodiments, such as Figure 4 As shown in (b), after the electronic device performs step S304, it extracts features from the audio stream to obtain audio features, calls the trained scene recognition network model to process the audio features, and obtains the scene to which the audio belongs. The scene to which the audio belongs is the scene recognition result.

[0095] In some embodiments, feature extraction of an audio stream by an electronic device includes extracting long-term audio features and short-term audio features from the audio stream. A scene recognition network model processes the long-term and short-term audio features to determine the scene to which the audio stream belongs. Long-term audio features can be understood as extracting features from a relatively long period of audio in the audio stream, such as a segment of audio lasting several tens of seconds; short-term audio features can be understood as extracting features from a shorter period of audio in the audio stream or at a specific moment, such as a segment of audio lasting several seconds or a single second.

[0096] In one possible implementation, the electronic device can be configured with a scene recognition network model. This scene recognition network model can recognize scenes that are not limited to one; for example, it can identify whether the audio stream includes scenes such as a kitchen, park, square, or conference room. The electronic device recognizes different scenes through a single scene recognition network model, offering the advantage of a simple processing flow.

[0097] Electronic devices may also use other methods to perform scene recognition on audio and obtain scene recognition results, and this application embodiment does not limit this.

[0098] S306. The electronic device performs sound event recognition on the audio separated from the video and obtains the sound event recognition result.

[0099] A sound event can be understood as a specific audio event, i.e., a specific sound. Examples include human voices, chewing sounds, applause, and fireworks. In some embodiments, the scope of sound events that an electronic device can recognize can be configured.

[0100] In some embodiments, the electronic device may be configured with a neural network model, referred to as an acoustic event recognition network model, which is used to recognize acoustic events in the audio stream. The acoustic event recognition network model can also be trained in the same way as the scene recognition network model, and will not be described further here.

[0101] In some embodiments, after the electronic device performs step S304, it extracts features from the audio stream to obtain audio features, and calls a trained sound event recognition network model to process the audio features, thereby obtaining the sound events included in the audio stream, i.e., the sound event recognition result.

[0102] In some embodiments, feature extraction of an audio stream by an electronic device includes extracting short-term audio features from the audio stream. A scene recognition network model processes these short-term audio features to obtain the sound events included in the audio stream. Short-term audio features can be understood as extracting features of audio from a short time period or a specific moment in the audio stream, such as a segment of audio lasting a few seconds or a single second.

[0103] In one possible implementation, the electronic device may be configured with an audio event recognition network model. This network model can recognize more than one audio event; for example, it can identify whether the audio stream includes human voices, chewing sounds, applause, and fireworks sounds. The electronic device's ability to recognize different audio events through a single audio event recognition network model offers the advantage of a simple processing flow.

[0104] Electronic devices may also use other methods to perform sound event recognition on audio and obtain sound event recognition results, and this application embodiment does not limit this.

[0105] In other embodiments, the electronic device can perform sound event recognition on the audio separated from the video, obtain sound event recognition results, and determine the scene to which the audio belongs based on the sound event recognition results. One scenario is configured with key audio, i.e., specific audio; the electronic device can identify the key audio in the audio and determine the scene to which the audio belongs based on the scene corresponding to the key audio.

[0106] S307. Electronic devices combine scene recognition results to determine whether the sound event recognition results include the target sound event.

[0107] In some embodiments, the electronic device can be configured with target sound events corresponding to a scene. After the electronic device executes steps S305 and S306, it can use the scene recognition results to determine the scene to which the audio belongs and whether the sound event recognition results include the target sound event corresponding to the scene to which the audio belongs.

[0108] In some embodiments, a target sound event can be understood as a key sound or indicative sound in a scene, which can indicate the scene to which the audio belongs.

[0109] In some embodiments, human voice events are all target sound events in all scenarios. Different target sound events can be set for specific scenarios, such as chewing in a kitchen scene, applause in a conference room scene, and fireworks in a square scene.

[0110] If the electronic device recognizes the sound event and the recognition result includes the target sound event, then step S308 is executed. If the electronic device recognizes the sound event and the recognition result does not include the target sound event, then step S303 is executed and the electronic device acquires the next video.

[0111] In some embodiments, after the electronic device identifies that the sound event identification result includes the target sound event, it may also record the target sound event included in the sound event identification result. In one possible implementation, the electronic device records the target sound event included in the sound event identification result by setting a tag for the target sound event and recording the tag.

[0112] For example, such as Figure 5As shown in (a), the scenes identified by the scene recognition network model include: kitchen, conference room, square, and park. The sound events identified by the sound event recognition network model include: human voice events, chewing sound events, clapping sound events, fireworks sound events, and birdsong sound events. The target sound events corresponding to the kitchen scene include: human voice events and chewing sound events; the target sound events corresponding to the conference room scene include: human voice events and clapping sound events; the target sound events corresponding to the square scene include: human voice events and fireworks sound events; and the target sound events corresponding to the park scene include: human voice events and birdsong sound events.

[0113] like Figure 5 As shown in (b), the electronic device performs scene classification and sound event recognition on the input audio. Scene classification includes determining whether the audio belongs to a kitchen scene, a conference room scene, a square scene, or a park scene. Sound event recognition includes determining whether the audio contains human voice events, chewing sound events, clapping sound events, fireworks sound events, and bird call sound events. The electronic device combines the scene recognition results to determine whether the sound event recognition results include the target sound event.

[0114] S308. Electronic devices determine whether the target sound event includes only human voices.

[0115] After obtaining the target sound events included in the sound event recognition results, the electronic device can further identify whether the target sound events only include human voices and do not include other sound events.

[0116] In some embodiments, the electronic device identifies the tags of the recorded target sound events. If only the tags of human voice events are identified, it is determined that the target sound event includes only human voices. If the tags of human voice events and tags of other sound events are identified, it is determined that the target sound event does not include only human voices.

[0117] If the electronic device determines that the target sound event only includes human voice, then it executes step S309; ​​if the electronic device determines that the target sound event does not only include human voice, then it returns to execute step S310.

[0118] S309. The electronic device performs noise reduction processing on the audio separated from the video.

[0119] In some embodiments, the electronic device may use Wiener filtering or deep neural network models to perform noise reduction on the audio; the specific implementation process of noise reduction will not be described here. Of course, the electronic device may also use other noise reduction methods to perform noise reduction on the audio.

[0120] The audio only includes human voice events, meaning it only contains the sound of people speaking. To obtain clean human voices, electronic devices perform noise reduction processing on the audio. Generally, audio noise reduction processing refers to removing noise other than human voices from the audio. In scenarios where the target sound events only include human voice events, the electronic device performs noise reduction processing on the audio without affecting other sound events, thus allowing for a more complete preservation of target sounds other than human voices in one-click short videos.

[0121] In some embodiments, step S308 may not be performed. In scenarios where the event recognition result includes only human voice events, or may include other human voice events, the electronic device performs noise reduction processing on the audio.

[0122] In other embodiments, steps S308 and S309 may also be omitted.

[0123] S310, The electronic device marks the audio as audio to be merged.

[0124] The electronic device determines in step S307 that the audio separated from the video includes the target sound event, thereby determining that the audio needs to be retained in the one-click video. The electronic device can mark the audio as audio to be merged.

[0125] In some embodiments, the electronic device can mark the audio as the audio to be merged by configuring parameters for the audio, where the parameters can be the start and end time periods of the audio. These parameters can then be used in the merging action of step S311 described below. For example, if the configured parameters for an audio segment are 00:01:10 to 00:01:30, then these parameters indicate that the audio exists during the time period of 00:01:30 to 00:01:30, and can be merged with background music.

[0126] S311. The electronic device performs fusion processing on the audio to be fused and the background music in the acquired video at the corresponding time period to obtain the fused audio.

[0127] The electronic device filters out the audio to be merged from the audio separated from the user-selected video through steps S303 to S310. The electronic device then merges the audio to be merged with the background music configured in step S302. In some embodiments, the audio to be merged and the background music are audio streams configured with timestamps. Therefore, the audio to be merged and the background music are merged in corresponding time periods.

[0128] In some embodiments, the fusion of the audio to be fused and the background music within a corresponding time period can be understood as the background music being fused with the audio within the time period indicated by the parameters of the audio to be fused.

[0129] In some embodiments, the audio to be merged and the background music can be merged in a certain proportion. It is understood that merging the audio to be merged and the background music in a certain proportion means that the amplitude of the audio to be merged and the amplitude of the background music are merged in a certain proportion during a corresponding time period. This proportion can be any value.

[0130] In one possible implementation, the fusion ratio of the audio to be fused is greater than that of the background music, thereby highlighting the audio to be fused in the one-click short video. In some embodiments, the target sound events included in the audio to be fused include human voice events, and the degree to which the fusion ratio of the audio to be fused is greater than that of the background music is increased, which can significantly highlight the human voice in the one-click short video.

[0131] For example, if the target sound events in the audio to be blended include human voice events, the audio to be blended and the background music can be mixed in a 7:3 ratio; or if the target sound events in the audio to be blended include human voice events and other target sound events, the audio to be blended and the background music can be mixed in a 6:4 ratio. The above blending ratios are not specifically limited and can be other ratios.

[0132] In other embodiments, the audio to be merged and the background music are dynamically merged based on the amplitude (or energy) of the audio to be merged and the background music.

[0133] In one possible implementation, the principle for blending the audio to be blended and the background music is that the energy of the audio to be blended and the energy of the background music should be kept at a comparable level; that is, the energy of the audio to be blended and the energy of the background music are basically the same in the blended audio. The average energy of the audio to be blended is denoted as Pow(original sound), and the average energy of the background music is denoted as Pow(background). The blending coefficient of the audio to be blended can be denoted as Pow(background) / (Pow(background)+Pow(original sound)), and the blending coefficient of the background music can be denoted as Pow(original sound) / (Pow(background)+Pow(original sound)). The product of the amplitude of the audio to be blended and its blending coefficient is taken as the amplitude of the audio to be blended in the blended audio, and the product of the amplitude of the background music and its blending coefficient is taken as the amplitude of the background music in the blended audio.

[0134] It should be noted that for scenarios where there are multiple audio files to be merged, the merging ratio of each audio file and the background music in the corresponding time period can be found in the methods provided above.

[0135] In some embodiments, there are multiple audio files to be merged. The electronic device can combine the multiple audio files into a single audio file in chronological order and then merge it with the background music. If there are multiple background music files, they can also be combined into a single background music file beforehand.

[0136] In other embodiments, there are multiple audio files to be merged. The electronic device can also merge each audio file to be merged with the background music, and then combine the merged audio files in chronological order to form a whole segment of merged audio.

[0137] S312 The electronic device merges the images separated from the fused audio and video, as well as the acquired images, to obtain a short video with one click.

[0138] In some embodiments, the electronic device can first fuse the images separated from the fused audio and video to obtain a video. The electronic device then further fuses the video and the acquired images to obtain a short video that can be created with a single click. The electronic device can edit the acquired images after the video, and the electronic device can edit the images after the video in order of their storage time from newest to oldest.

[0139] In one possible implementation, during the process of merging the video and the acquired images, the electronic device can edit the video, retaining some video segments.

[0140] In some embodiments, the duration of the background music is longer than the duration of the video selected by the user in step S301, and the user selects multiple videos, with the duration of the background music exceeding the total duration of the selected videos. Thus, in the one-click short video generated by the electronic device, the images selected by the user also have background music.

[0141] In some embodiments, the electronic device is configured with style-matching effects for user-selected images and videos. These effects can also be added to images extracted from the video and to the acquired images to create special effects.

[0142] Another embodiment of this application also provides a method for generating one-click short videos, which is described below in conjunction with... Figure 6 This paper introduces the process of generating short videos with one click using video data processing methods on electronic devices.

[0143] like Figure 6 As shown in the embodiment of this application, a method for processing video data includes:

[0144] S601, Electronic devices acquire videos and images selected by the user.

[0145] The content of step S601 can be found in the aforementioned step S301, and will not be repeated here.

[0146] S602, Electronic devices match background music to images and videos.

[0147] The content of step S602 can be found in the aforementioned step S302, and will not be repeated here.

[0148] S603, Electronic device acquires a video.

[0149] The content of step S603 can be found in the aforementioned step S303, and will not be repeated here.

[0150] S604, The electronic device separates the image and audio of the video.

[0151] The content of step S604 can be found in the aforementioned step S304, and will not be repeated here.

[0152] S605: The electronic device performs speech recognition on the audio extracted from the video and obtains the speech recognition result.

[0153] Human voice, as a special sound in audio, usually needs to be preserved in short videos created with a single click. Therefore, for a video, after the electronic device separates the audio from the video, it performs speech recognition on the audio to determine whether the audio contains a speech signal, which is usually understood as human voice.

[0154] In some embodiments, the electronic device uses the cepstral method to determine the fundamental frequency. If the electronic device determines that the fundamental frequency exists, it means that the audio includes a speech signal, that is, the speech recognition result is that the audio includes a speech signal.

[0155] In the field of speech recognition, an audio segment can be feature-extracted to obtain audio features. These features are then used for speech or non-speech recognition to determine whether the audio segment contains a speech signal. The cepstral method for determining the fundamental frequency can be understood as a method of audio feature extraction. Of course, electronic devices can also use other methods for feature extraction.

[0156] In other embodiments, the electronic device employs a neural network model for speech recognition to obtain speech recognition results. One possible implementation involves the electronic device using a recurrent neural network model to perform speech endpoint detection (or speech activity detection) on the audio. The recurrent neural network model can use short-time zero-crossing rate and short-time energy as features for endpoint detection, and then use these features to identify whether the audio includes a speech signal, i.e., human voice.

[0157] S606. Electronic devices determine whether the speech recognition result includes human voice.

[0158] When an electronic device determines whether a speech recognition result includes human voice, it can be understood as whether the electronic device determines whether the speech recognition result indicates that the audio includes human voice.

[0159] If the electronic device determines that the speech recognition result includes human voice, it proceeds to step S607; if the electronic device determines that the speech recognition result does not include human voice, it returns to step S603 and the electronic device acquires the next video.

[0160] If the electronic device determines that the audio in the voice recognition result does not include human voice, it means that the audio can be retained in the one-click short video. Therefore, the electronic device obtains the next video selected by the user in step S603.

[0161] S607, The electronic device performs noise reduction processing on the audio separated from the video.

[0162] In some embodiments, the electronic device may use Wiener filtering or deep neural network models to perform noise reduction on the audio; the specific implementation process of noise reduction will not be described here. Of course, the electronic device may also use other noise reduction methods to perform noise reduction on the audio.

[0163] Speech recognition results indicate that the audio includes human voices. To obtain clean human voices, electronic devices perform noise reduction processing on the audio. Generally, audio noise reduction processing refers to removing noise other than human voices from the audio.

[0164] S608, The electronic device marks the audio as audio to be merged.

[0165] The content of step S608 can be found in the aforementioned step S310, and will not be repeated here.

[0166] S609. The electronic device performs fusion processing on the video to be merged and the background music in the acquired video at the corresponding time period to obtain the merged audio.

[0167] The content of step S609 can be found in the aforementioned step S311, and will not be repeated here.

[0168] The S610 electronic device merges the images separated from the fused audio and video, along with the acquired images, to obtain a short video with one click.

[0169] The content of step S610 can be found in the aforementioned step S312, and will not be repeated here.

[0170] Figure 1 This demonstrates a scenario where a user opens a gallery app and uses the videos and images stored in the app to generate a short video with a single click. In some embodiments, this one-click video generation function can also be applied to video shooting scenarios.

[0171] When a user opens the camera app and uses the one-click video creation function to shoot a short video, the electronic device can at least retain some of the audio captured by the microphone during the video shooting process, and can also highlight specific audio.

[0172] In some embodiments, a method for generating one-click short videos in a video shooting scenario includes:

[0173] The electronic device launches a camera application to capture video and obtain video data. The electronic device captures the image of the video data through a camera and the audio of the video data through a microphone.

[0174] The electronic device configures background music that matches the style of the video. In some embodiments, the electronic device can also configure special effects that match the video. In some embodiments, the electronic device can configure background music and / or special effects for the video in response to user-specified background music and / or special effects. The implementation of configuring background music that matches the style of the video by the electronic device can be found in step S302 of the foregoing embodiments, and will not be repeated here.

[0175] During video recording, the images and audio in the video can be understood as being captured and stored separately. Only after the user operates the electronic device screen to input and save the video are the images and audio merged into a single video. Therefore, the electronic device can perform scene recognition on the audio captured by the microphone to obtain scene recognition results, and also perform sound event recognition to obtain sound event recognition results. The implementation methods for scene recognition and sound event recognition can be found in steps S305 and S306 of the aforementioned embodiments, and will not be repeated here.

[0176] The electronic device, combining the scene recognition results, identifies sound events, including target sound events. It then fuses the audio captured by the microphone with the background music within the corresponding time period to obtain the fused audio. The specific implementation of this step can be found in step S311 of the aforementioned embodiment, and will not be repeated here.

[0177] The electronic device then merges the combined audio and the images captured by the camera to create a short video with a single click.

[0178] In some embodiments, after the electronic device identifies the sound event and the identification result includes the target sound event, it may further determine whether the target sound event only includes human voice. If it only includes human voice, then noise reduction processing is performed on the audio. The specific implementation process of noise reduction processing can be found in step S309 of the aforementioned embodiments, and will not be repeated here.

[0179] In other embodiments, the method for generating one-click short videos in a video shooting scenario includes:

[0180] The electronic device launches a camera application to capture video and obtain video data. The electronic device captures the image of the video data through a camera and the audio of the video data through a microphone.

[0181] The electronic device configures background music that matches the style of the video. In some embodiments, the electronic device can also configure special effects that match the video. In some embodiments, the electronic device can configure background music and / or special effects for the video in response to user-specified background music and / or special effects. The implementation of configuring background music that matches the style of the video by the electronic device can be found in step S302 of the foregoing embodiments, and will not be repeated here.

[0182] The electronic device performs speech recognition on the audio captured by the microphone to obtain a speech recognition result. The implementation method of speech recognition can be found in step S605 of the aforementioned embodiments, and will not be repeated here. In some embodiments, if the electronic device recognizes human voices in the speech recognition result, noise reduction processing is performed on the audio. The specific implementation process of noise reduction processing can be found in step S309 of the aforementioned embodiments, and will not be repeated here.

[0183] When the electronic device determines that the speech recognition result includes human voice, it fuses the audio captured by the microphone and the background music within the corresponding time period to obtain the fused audio. The specific implementation of this step can be found in step S311 of the aforementioned embodiment, and will not be repeated here.

[0184] The electronic device then merges the combined audio and the images captured by the camera to create a short video with a single click.

[0185] Another embodiment of this application provides a computer-readable storage medium storing instructions that, when executed on a computer or processor, cause the computer or processor to perform one or more steps of any of the above methods.

[0186] Computer-readable storage media can be non-transitory computer-readable storage media, such as read-only memory (ROM), random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage devices.

[0187] Another embodiment of this application provides a computer program product containing instructions. When the computer program product is run on a computer or processor, it causes the computer or processor to perform one or more steps of any of the methods described above.

Claims

1. A method for processing video data, characterized in that, include: The electronic device displays a first interface, which shows a thumbnail of the cover of a first video and a first button; The electronic device responds to the user's click on the first button by configuring background music, which matches the first video; Speech recognition is performed on the audio of the first video to obtain a speech recognition result. Based on the speech recognition result, it is determined that the audio of the first video includes human voices. Then, the audio of the first video and the background music are fused to obtain a fused audio, and the image of the first video and the fused audio are fused to obtain a second video. Alternatively, short-time audio features are extracted from the audio of the first video, and the short-time audio features are processed by a sound event recognition network model to obtain the sound events included in the audio of the first video. If it is determined that the sound events included in the audio of the first video include sound events of key audio, then the audio of the first video and the background music are fused to obtain a fused audio, and the image of the first video and the fused audio are fused to obtain a second video. The second interface is then displayed, which is the playback interface for the second video. The second video includes at least part of the audio and background music of the first video.

2. The video data processing method according to claim 1, characterized in that, Before displaying the second interface, the following is also included: The electronic device displays a third interface, which is a buffer interface for generating the second video.

3. The video data processing method according to claim 1, characterized in that, Before the electronic device fuses the audio of the first video and the background music to obtain the fused audio, it further includes: The electronic device performs noise reduction processing on the audio of the first video to obtain noise-reduced audio; The electronic device merges the audio of the first video and the background music to obtain merged audio, including: the electronic device merges the noise-reduced audio of the first video and the background music to obtain merged audio.

4. The video data processing method according to claim 1, characterized in that, In the process of the electronic device recognizing the audio of the first video, including key audio, the following further steps are included: The electronic device extracts features from the audio of the first video to obtain long-term audio features, and performs scene recognition on the long-term audio features to obtain scene recognition results. The scene recognition results indicate that the first video belongs to a first scene. The electronic device identifies key audio from the first video, including: The electronic device identifies that the audio of the first video includes key audio from the first scene.

5. The video data processing method according to claim 1, characterized in that, The electronic device is configured with multiple scenarios, and the electronic device identifies key audio from the first video, including: The electronic device identifies key audio belonging to the first scene from the audio of the first video.

6. The video data processing method according to claim 1, characterized in that, The key audio may include at least human voice, or the key audio may not include human voice.

7. The video data processing method according to claim 1, characterized in that, The process of fusing the audio of the first video and the background music to obtain the fused audio includes: The audio of the first video and the background music are merged according to the audio amplitude fusion ratio in the corresponding time period to obtain the merged audio.

8. The video data processing method according to claim 7, characterized in that, The audio amplitude fusion ratio is used to indicate that the amplitude of the first video audio in the fused audio is greater than the amplitude of the background music.

9. The video data processing method according to claim 1, characterized in that, The process of fusing the audio of the first video and the background music to obtain the fused audio includes: The audio of the first video and the background music are fused together according to the principle that the energy of the audio of the first video and the energy of the background music are basically the same to obtain the fused audio.

10. The video data processing method according to claim 9, characterized in that, The step of fusing the audio of the first video and the background music according to the principle that the energy of the audio of the first video and the energy of the background music are substantially the same to obtain the fused audio includes: The product of the audio energy of the first video and the first coefficient is taken as the audio energy of the first video in the fused audio. The product of the background music energy and the second coefficient is taken as the background music energy in the fused audio. The first coefficient is the proportion of the background music energy in the sum of the audio energy of the first video and the energy of the background music. The second coefficient is the proportion of the audio energy of the first video in the sum of the audio energy of the first video and the energy of the background music.

11. The video data processing method according to claim 6, characterized in that, The process of fusing the audio of the first video and the background music to obtain the fused audio includes: The key audio includes at least human voice. The audio of the first video and the background music are fused together in a first ratio of audio amplitude during the corresponding time period to obtain the fused audio. Alternatively, the key audio may not include human voices. The audio of the first video and the background music may be merged according to a second ratio of audio amplitude during the corresponding time period to obtain the merged audio. The first ratio and the second ratio of the audio amplitude are both used to indicate the proportional relationship between the audio of the first video and the background music; the first ratio of the audio amplitude is greater than the second ratio of the audio amplitude.

12. An electronic device, characterized in that, include: One or more processors, memory, and a display screen; The memory and the display screen are coupled to the one or more processors. The memory is used to store a computer program, the computer program including computer instructions, and when the one or more processors execute the computer instructions, the electronic device performs a video data processing method as described in any one of claims 1 to 11.

13. A computer-readable storage medium, characterized in that, Used to store a computer program, which, when executed, is specifically used to implement the video data processing method as described in any one of claims 1 to 11.

Citation Information

Patent Citations

  • Background music adding method and device and electronic device

    CN110740262A

  • Video processing method and electronic equipment

    CN115484389A