Music video generation method, equipment and medium
By switching to a video with human voices removed and recording audio during music video playback to synthesize a new music video, the problem of limited pre-made music video resources on the server is solved, storage and download costs are reduced, and user experience and enjoyment are enhanced.
Patent Information
- Application Number
- CN202511402201.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-28
- Publication Date
- 2025-12-05
AI Technical Summary
In existing technologies, the pre-built music video resources on servers are limited, which affects user experience. Storage and download bandwidth costs are high, and the experience is repetitive and monotonous when users sing the same song multiple times.
During music video playback, respond to the recording trigger operation, switch the music video to a video with vocals removed, record audio, and synthesize a new music video.
It eliminates the need to retrieve pre-made, voiceless music and video resources from servers, reducing storage and download bandwidth costs and enhancing the user's recording experience and enjoyment.
Smart Images

Figure CN121075293A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio processing technology, and in particular to a method, device and medium for generating music videos. Background Technology
[0002] Singing using an app is a common form of entertainment. Currently, existing technologies and products primarily involve users first viewing a music video (MV) with vocals and accompaniment. If a user wants to record, the app retrieves and downloads a corresponding MV with only accompaniment and no vocals. This MV is then used for recording and karaoke. These resources need to be pre-processed and stored on a server, resulting in high storage and download bandwidth costs. Furthermore, pre-made MV versions are often limited, leading to a repetitive and monotonous experience when users sing the same song multiple times, limiting the enjoyment of singing. Moreover, with the increasing number of user-created short video content, including many high-quality song videos, the server may not have the corresponding pre-made MV resources, preventing users from recording. In summary, during the development of this invention, the inventors discovered that existing technologies suffer from limited pre-made MV resources on servers, impacting user experience, and incurring high storage and download bandwidth costs. Summary of the Invention
[0003] In view of this, the purpose of this application is to provide a method, device, and medium for generating music videos, which can avoid the problem of limited pre-made MV resources on servers affecting user experience, and reduce storage costs and download bandwidth costs. The specific solution is as follows:
[0004] Firstly, this application provides a method for generating music videos, including:
[0005] During the playback of the first music video, in response to the recording trigger operation, the first music video is switched to the second music video; the second music video is the music video obtained by removing the first vocal audio from the first music video;
[0006] Audio recording is performed during the playback of the second music video;
[0007] After the audio recording is completed, a third music video generated based on the second human voice audio and the second music video is displayed.
[0008] Optionally, switching the playing first music video to a second music video includes:
[0009] Extract the audio file and video background file of the first music video being played;
[0010] Extract the accompaniment from the audio file to obtain the accompaniment file;
[0011] The video background file and the accompaniment file are played synchronously to switch the first music video to the second music video;
[0012] Correspondingly, after completing the audio recording, the following is also included:
[0013] The third music video is synthesized based on the recorded second human voice audio, the video background file, and the accompaniment file.
[0014] Optionally, the accompaniment is extracted from the audio file to obtain an accompaniment file, including:
[0015] Perform a sound accompaniment separation operation on the audio file to obtain an initial accompaniment file;
[0016] The initial accompaniment file is subjected to human voice detection. If human voice noise is detected, the human voice noise is removed from the initial accompaniment file to obtain the accompaniment file.
[0017] Optionally, after audio recording is complete, the following may also be included:
[0018] Determine the audio pickup device corresponding to the second human voice audio;
[0019] If the sound pickup device is an external speaker, then echo cancellation and noise cancellation operations are performed on the second human voice audio to obtain the second human voice audio after cancellation.
[0020] Accordingly, based on the recorded second human voice audio, the video background file, and the accompaniment file, a third music video is synthesized, including:
[0021] A third music video is obtained by synthesizing the second human voice audio after elimination, the video background file, and the accompaniment file.
[0022] Optionally, a third music video is obtained by synthesizing the recorded second human voice audio, the video background file, and the accompaniment file, including:
[0023] The second human voice audio and the accompaniment file are combined into a single audio file to obtain the target audio file;
[0024] The target audio file and the video background file are combined to obtain a third music video.
[0025] Optionally, the second vocal audio and the accompaniment file are combined into a single audio track to obtain the target audio file, including:
[0026] The auditory delay is determined based on the timestamp corresponding to the human voice in the second human voice audio and the timestamp of the lyrics corresponding to the accompaniment file.
[0027] The second human voice audio is cut based on the auditory delay to obtain the cut second human voice audio.
[0028] The cut second human voice audio and the accompaniment file are combined into a single audio file to obtain the target audio file.
[0029] Optionally, determining the auditory delay based on the timestamp corresponding to the human voice in the second human voice audio and the timestamp of the lyrics corresponding to the accompaniment file includes:
[0030] Determine the timestamp of the first word in the second human voice audio to obtain the first timestamp;
[0031] Determine the timestamp of the first character in the lyrics timestamp of the accompaniment file to obtain the second timestamp;
[0032] Hearing delay is determined based on the first timestamp and the second timestamp.
[0033] Optional, also includes:
[0034] During the synchronized playback of the video background file and the accompaniment file, if a playback jump operation is detected, the accompaniment file is played according to the timestamp corresponding to the playback jump operation, while the timestamp of the video background file is aligned with the accompaniment file for playback.
[0035] Optionally, after displaying the third music video generated based on the recorded second vocal audio and the second music video, the method further includes:
[0036] Generate prompt information and display it on the user interface, wherein the prompt information includes prompts for saving and / or publishing;
[0037] The response is used for save and / or publish operations, saving and / or publishing the third music video.
[0038] Secondly, this application provides a music video generation apparatus, comprising:
[0039] The video switching module is used to switch the first music video to a second music video in response to a recording trigger operation during the playback of the first music video; the second music video is the music video obtained by removing the first vocal audio from the first music video.
[0040] An audio recording module is used to record audio during the playback of the second music video;
[0041] The video generation module is used to display a third music video generated based on the second human voice audio and the second music video after the audio recording is completed.
[0042] Thirdly, this application provides an electronic device, including a memory and a processor, wherein:
[0043] The memory is used to store computer programs;
[0044] The processor is used to execute the computer program to implement the aforementioned music video generation method.
[0045] Fourthly, this application provides a computer-readable storage medium for storing a computer program, wherein the computer program, when executed by a processor, implements the aforementioned music video generation method.
[0046] Fifthly, this application provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the aforementioned music video generation method.
[0047] As can be seen from the above scheme, the present invention provides a music video generation method, including: during the playback of a first music video, in response to a recording trigger operation, switching the first music video to a second music video; the second music video is a music video obtained by removing the first human voice audio from the first music video; during the playback of the second music video, performing audio recording; and after the audio recording is completed, displaying a third music video generated based on the recorded second human voice audio and the second music video.
[0048] As can be seen, the beneficial effects of this application are as follows: In this application, during the playback of the first music video, in response to the recording trigger operation, the first music video being played is switched to a second music video obtained by removing the first vocal audio from the first music video. During the playback of the second music video, audio recording is performed, and further, a third music video generated based on the recorded second vocal audio and the second music video is displayed. In this way, there is no need to obtain pre-made music video resources without vocals from the server. The processing and playback are switched directly based on the currently playing music video, while simultaneously recording vocals for subsequent synthesis. This is not limited by the resources on the server, and the server does not need to store too many pre-made music video resources, thus avoiding the problem of limited pre-made MV resources on the server affecting the user experience, and reducing storage costs and download bandwidth costs.
[0049] Correspondingly, the music video generation apparatus, device, readable storage medium, and product provided in this application also have the above-mentioned technical effects. Attached Figure Description
[0050] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0051] Figure 1 A flowchart of a music video generation method provided in this application embodiment;
[0052] Figure 2 A flowchart for generating a music video is provided as an embodiment of this application;
[0053] Figure 3 This is a schematic diagram of the structure of a music video generation device provided in an embodiment of this application;
[0054] Figure 4 This is a structural diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0055] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0056] Singing on mobile apps is a common form of entertainment. Often, the singing process is accompanied by pre-made music videos or real-time vocal scoring, providing a richer singing experience. However, the number of pre-made music video versions for the same song is often limited, such as official music videos or a few music videos produced by content creators. This limited quantity means that the experience can become repetitive and monotonous when users sing the same song repeatedly, limiting the enjoyment of singing. Nowadays, user-generated short video content is increasing, including many high-quality song videos or music videos. For example, many short videos on video platforms include song clips or the full song, but cannot be recorded directly.
[0057] Current industry technologies and products primarily involve users first viewing the original audio with vocals or a music video (MV) with both vocals and instrumental music. When a user wants to sing, the app retrieves and downloads the corresponding instrumental music or an MV with only instrumental music. This allows the user to record and sing karaoke. These resources need to be pre-processed and stored on a server. After recording, the recorded vocals are combined with the corresponding instrumental music or MV to create a new MV. The process begins with instrumental music production, where professional sound engineers or MV producers create the music and provide it to a commercial company. The company then uploads the files to a server. Users who want to sing use the karaoke software client to retrieve the instrumental music or MV from the server and play it locally on their client. Simultaneously, the user records using a microphone. After recording, the vocals captured by the microphone are combined with the instrumental music / MV to create a new song. This requires readily available backing tracks or music videos, which must be processed to be voiceless to ensure users are not disturbed by other voices while recording. If the server does not have the corresponding backing track or music video for a particular song, the user will be unable to record. Users must download the corresponding backing track or music video before recording can begin, which involves waiting and is relatively cumbersome, reducing user motivation. Furthermore, the server needs to store the corresponding backing track or music video, while users need to download them, incurring storage and download bandwidth costs. Therefore, this application provides a music video generation solution that avoids the problem of limited pre-made music video resources on the server affecting user experience, and reduces storage and download bandwidth costs.
[0058] See Figure 1 As shown in the figure, this application discloses a method for generating music videos, including:
[0059] Step S11: During the playback of the first music video, in response to a recording / singing trigger operation, the first music video is switched to a second music video; the second music video is the music video obtained by removing the first vocal audio from the first music video. The recording / singing trigger operation can be a click operation on a preset button in the graphical user interface. During the playback of the first music video, in response to the recording / singing trigger operation, the first music video can be switched to the second music video. The second music video can be the music video obtained by removing the first vocal audio from the first music video.
[0060] In an optional implementation, switching the first music video being played to a second music video includes: extracting the audio file and video background file of the first music video being played; extracting the accompaniment from the audio file to obtain an accompaniment file; and playing the video background file and the accompaniment file synchronously to switch the first music video to the second music video.
[0061] This embodiment of the application can extract the audio file and video background file of the currently playing first music video in response to a recording trigger operation. In this embodiment, the music video can be played through a client installed on the terminal device. The recording trigger operation can be a user's click on a preset button, for example, in response to a user's click on the recording button in the client's user interface, the client extracts the music file and video background file of the currently playing first music video. The currently playing first music video can be a music video downloaded from a server to the local device by the client, or it can be a music video played online by the client. If it is an online music video, it is first downloaded to the local device, and then the audio file and video background file of the currently playing first music video are extracted. Alternatively, the audio file and video background file of the currently playing first music video can be extracted based on received streaming media data. The audio file is the sound in the music video, and the video background file is the background image in the music video.
[0062] In addition, in this embodiment, in response to the recording trigger operation for the currently playing music video, the method further includes: generating prompt information during the recording process and displaying it on the user interface, for example, displaying it as a pop-up window, so as to inform the user to wait.
[0063] In this embodiment, the accompaniment can be extracted from the audio file to obtain an accompaniment file. This embodiment does not limit the method for extracting the accompaniment from the audio file; it can include DSP algorithms, models, etc. For example, a voice-accompaniment separation neural network model can be used to separate the voice and accompaniment from the audio file to obtain the accompaniment file and the vocal file. This scheme uses voice-accompaniment separation primarily to obtain the accompaniment file; a vocal removal scheme can also be used to achieve the same effect. Furthermore, the accompaniment can be extracted from the audio file on the client side to obtain the accompaniment file.
[0064] Furthermore, in an optional implementation, extracting the accompaniment from the audio file to obtain an accompaniment file includes: performing a sound-accompaniment separation operation on the audio file to obtain an initial accompaniment file; performing human voice detection on the initial accompaniment file, and if human voice noise is detected, removing the human voice noise from the initial accompaniment file to obtain the accompaniment file.
[0065] In this embodiment, the audio file can be separated into its accompaniment to obtain an initial accompaniment file. Then, a human voice detection algorithm is used to scan the initial accompaniment file to determine if there is any additional human voice noise. If so, the scanned human voice part is removed, thus ensuring the purity of the accompaniment.
[0066] In an optional implementation, after extracting the complete audio file and video background file corresponding to the currently playing first music video, the audio file can be separated from its audio accompaniment. Alternatively, a streaming processing method can be used, that is, after each audio file is extracted, the audio accompaniment is separated for that audio file, thereby further accelerating the processing speed.
[0067] Step S12: Record audio while playing the second music video.
[0068] In this embodiment, the video background file and the accompaniment file are played synchronously, and the recording of the user's voice is started at the same time to obtain the user's voice audio, i.e., the second voice audio.
[0069] In this embodiment, timestamp alignment can be used to synchronize the playback of the video background file and the accompaniment file. Timestamp alignment means that the current playback time is the same. This means the video background file and the accompaniment file have the same length and the same current playback timestamp, allowing for synchronized playback and recording. Synchronized playback provides a user experience of synchronized audio and video. In an optional implementation, the accompaniment file and the video background file can also be combined into a second music video and played.
[0070] Furthermore, during the synchronized playback of the video background file and the accompaniment file, if a playback jump operation is detected, the accompaniment file is played according to the timestamp corresponding to the playback jump operation, while the timestamp of the video background file is aligned with the accompaniment file for playback.
[0071] In this embodiment, when a user performs a seek (i.e., jump playback) operation, the timestamp of the accompaniment file is adjusted first for jump playback based on the timestamp of the jump playback operation (i.e., the timestamp to which the user jumps), while the timestamp of the video file is aligned with the accompaniment file for playback, achieving audio-visual synchronization after the jump playback operation. Since human hearing is more sensitive, this embodiment adjusts the accompaniment file first to ensure a better user experience.
[0072] This embodiment can activate the player component in the client to play the video background file and the accompaniment file synchronously. Simultaneously, it can also activate the recorder component in the client, which starts recording via a microphone. This microphone can be a headset or the microphone of the terminal device; either the headset can capture the user's voice audio, or the terminal device's microphone can capture the user's voice audio when the audio is played aloud. In this way, the user can simultaneously record and sing while watching the video background file in the player and listening to the accompaniment file. The recorder records and saves the user's voice.
[0073] Furthermore, after completing the audio recording, the method also includes: determining the pickup device corresponding to the second human voice audio; if the pickup device is an external speaker, then performing echo cancellation and noise cancellation operations on the second human voice audio to obtain the canceled second human voice audio.
[0074] In other words, in this embodiment, after starting the recording of the user's voice and obtaining the user's voice audio, the method further includes: determining the sound pickup device corresponding to the user's voice audio; if the sound pickup device is an external speaker, then performing echo cancellation and noise cancellation operations on the user's voice audio to obtain the canceled user voice audio. If the sound pickup device is headphones, echo cancellation and noise cancellation operations are not required. It is understood that if the user records using headphones, the recorded voice is clean and can be directly used for the next step of processing. If the user records using an external speaker, the accompaniment sound will be picked up, and the echo cancellation and noise cancellation performed in this embodiment can ensure the purity of the voice audio.
[0075] Step S13: After completing the audio recording, display the third music video generated based on the second human voice audio and the second music video obtained from the recording.
[0076] The display can be understood as follows: after the audio recording is completed, the third music video can be played automatically, and the user can watch the new video with the second voice audio without any operation; or, a preview of the third music video can be displayed, and the user can click the play button to view the full video; or, a pop-up window or prompt box can be used to notify the user that the third music video has been generated and provide a "play" button; or, through a jump button or link, the user can choose to enter a new page or interface to watch the generated third music video, etc.
[0077] In this embodiment, after the audio recording is completed, the method further includes: synthesizing the second human voice audio obtained from the recording, the video background file, and the accompaniment file to obtain a third music video.
[0078] In this embodiment, if the sound pickup device is an external speaker, the third music video is obtained by synthesizing the second human voice audio after deletion, the video background file, and the accompaniment file. If the sound pickup device is an earphone, the third music video is obtained by synthesizing the recorded second human voice audio, the video background file, and the accompaniment file.
[0079] In an optional implementation, a third music video is obtained by synthesizing the recorded second human voice audio, the video background file, and the accompaniment file, including: synthesizing the second human voice audio and the accompaniment file into a single audio track to obtain a target audio file; and synthesizing the target audio file and the video background file to obtain the third music video.
[0080] In this embodiment, the second human voice audio and the accompaniment file can be first combined into a single audio file using timestamp alignment to obtain the target audio file. Then, the target audio file and the video background file can be combined using timestamp alignment to obtain the target music video.
[0081] Furthermore, in an optional implementation, the second vocal audio and the accompaniment file are combined into a single audio track to obtain the target audio file, including: determining an auditory delay based on the timestamp corresponding to the vocals in the second vocal audio and the timestamp corresponding to the lyrics in the accompaniment file; cutting the second vocal audio based on the auditory delay to obtain a cut second vocal audio; and combining the cut second vocal audio and the accompaniment file into a single audio track to obtain the target audio file.
[0082] Although the video background file and accompaniment file are played at the same time as the user's voice is recorded, different people have different hearing sensitivities, especially the elderly, who may have a significant hearing delay. That is, although the accompaniment has reached a certain point in time, the user hears it with a delay. This application embodiment can determine the user's hearing delay based on the timestamp corresponding to the voice in the second voice audio and the timestamp of the lyrics corresponding to the accompaniment file. Based on the hearing delay, the second voice audio and the accompaniment file are aligned, that is, the start time of the lyrics corresponding to the accompaniment is aligned with the time when the user starts singing, and the second voice audio and the accompaniment file are combined into a single audio file.
[0083] In an optional implementation, a segment corresponding to the auditory delay at the beginning of the second vocal audio can be cut to obtain the cut user vocal audio. The cut user vocal audio and the accompaniment file are then combined into a single audio track to obtain the target audio file. The cut user vocal audio can be padded at the end, i.e., padded to the same length as the accompaniment file, such as by adding zeros.
[0084] Furthermore, in an optional implementation, determining the auditory delay based on the timestamp corresponding to the human voice in the second human voice audio and the timestamp of the lyrics corresponding to the accompaniment file may include: determining the timestamp of the first word of the human voice in the second human voice audio to obtain a first timestamp; determining the timestamp of the first word in the timestamp of the lyrics corresponding to the accompaniment file to obtain a second timestamp; and determining the auditory delay based on the first timestamp and the second timestamp.
[0085] That is, in this embodiment, the timing of the first word of the human voice in the audio and the timing of the first word of the lyrics in the accompaniment file can be determined to obtain the first timestamp and the second timestamp. The auditory delay can be obtained by subtracting the second timestamp from the first timestamp.
[0086] In another optional implementation, determining the auditory delay based on the timestamps corresponding to the vocals in the second vocal audio and the timestamps of the lyrics in the accompaniment file can include: determining the timestamps of the first preset number of words in the second vocal audio to obtain a preset number of first timestamps; determining the timestamps of the first preset number of words in the timestamps of the lyrics in the accompaniment file to obtain a preset number of second timestamps; calculating the differences between the preset number of first timestamps and the preset number of second timestamps in sequence, and taking the average to obtain the auditory delay. Here, the preset number is multiple, i.e., greater than or equal to two. By averaging the timestamps of multiple words, the accuracy of the alignment between the user's vocal audio and the accompaniment file can be improved, thus enhancing the quality of the target music audio.
[0087] Furthermore, in an optional implementation, in response to the click of the recording end button, a third music video can be synthesized based on the recorded second vocal audio, the video background file, and the accompaniment file to obtain a third music video. Further, in an optional implementation, after displaying the third music video generated based on the recorded second vocal audio and the second music video, the method further includes: generating a prompt message and displaying it on the user interface, wherein the prompt message includes a prompt message for saving and / or publishing; and in response to the save and / or publish operation, saving and / or publishing the third music video.
[0088] Alternatively, if it has been set as the default for saving and publishing, it will save and publish directly, and generate a message indicating that it has been saved and published.
[0089] In other words, in this embodiment, the user can save or publish the generated third-party music video. It should be noted that the music video generated using this embodiment is a music video that has already obtained the appropriate authorization.
[0090] As can be seen, in this embodiment, during the playback of the first music video, in response to a recording trigger operation, the first music video is switched to a second music video obtained by removing the first vocal audio from the first music video. During the playback of the second music video, audio recording is performed. Furthermore, a third music video generated based on the recorded second vocal audio and the second music video is displayed. This eliminates the need to retrieve pre-made music video resources without vocals from the server. Processing and switching playback are performed directly based on the currently playing music video, while simultaneously recording vocals for subsequent synthesis. This avoids the resource limitations of the server, and the server does not need to store excessive pre-made music video resources, thus preventing the problem of limited pre-made music video resources on the server affecting user experience. It also reduces storage and download bandwidth costs.
[0091] Furthermore, taking a music app as an example, this application's music video generation solution is further illustrated. (See [link]) Figure 2 As shown, Figure 2 This application provides a flowchart for generating a music video. It mainly includes MV audio extraction, audio file accompaniment separation, synchronized playback of accompaniment and video files, user-recorded vocals, and multi-track file synthesis. Audio extraction refers to the technical process of separating audio data from video files or other multimedia carriers and saving it as an independent audio file. Accompaniment separation refers to the technical process of separating vocals and accompaniment from mixed audio using signal processing or artificial intelligence technology. This application designs an audio extraction process that separates the audio and video in a music video into separate audio and video files, i.e., the aforementioned video background file. An audio separation process is designed to separate the vocals and accompaniment from the extracted individual audio files, resulting in separate vocal and accompaniment files. A synchronized audio and video playback process is designed to play the accompaniment and video files simultaneously, allowing users to sing along while watching the video and listening to the accompaniment, thus obtaining the user's recorded vocals. A workflow was designed to synthesize two audio files and one video file, combining vocal files, backing track files, and video files to create a new music video. Through audio extraction and separation, the backing track and music video playback resources needed for karaoke can be obtained in real time. By synchronizing audio and video, users can sing along to an interesting music video in real time. Finally, by merging multiple files, a work identical to the original music video is generated, which users can save or publish. The following sections describe the process in detail:
[0092] MV Audio Extraction: When users watch MVs in the app, they can click the "Record" button for an MV they are interested in. The app will then display a pop-up message informing the user that "Recording is in progress." At this time, the app's background process will start the audio extraction algorithm to process the MV and extract the audio file audio.file and the video file video.file from the MV. The audio file will be used to play the sound, while the video file will be used to play the background of the MV video.
[0093] Audio file audio accompaniment separation: Using an audio accompaniment separation algorithm / model, the audio file audio.file extracted in the previous step is subjected to audio accompaniment separation, generating an accompaniment file accompaniment.file and a vocal file vocal.file, which are stored separately. The accompaniment file accompaniment.file is further scanned using a vocal detection algorithm to determine if there is any additional vocal noise. If so, the scanned vocal parts are removed to ensure the accompaniment is pure.
[0094] Synchronized playback of instrumental and video files: The instrumental file accompany.file and the video file video.file are sent to the player for playback. Using a timestamp alignment scheme, the instrumental file accompany.file and the video file video.file are played synchronously, providing a synchronized audio-visual experience for the user.
[0095] When a user manually performs a seek operation, the timestamp of the accompany.file is adjusted first for seek playback based on the seek timestamp, while the timestamp of the video.file is aligned with that of the accompany.file for playback, achieving audio-visual synchronization after seek.
[0096] User-generated vocal recording: When the backing track and video files are played synchronously, a sound recorder is simultaneously activated to capture the user's voice. The user watches the video file on the player while listening to the backing track, recording simultaneously. The app's sound recorder records the user's voice and saves it separately as personal_voice.file. If the user records using headphones, the recorded voice is clean and can be used directly for further processing. If the user records using speakers, the backing track will be picked up, requiring echo cancellation and noise reduction.
[0097] Multi-track file merging: After the user finishes recording, clicking the "Done" button initiates the MV merging process. The app merges the user's recorded vocals (personal_voice.file), video file (video.file), and instrumental file (accompany.file). Using timestamp alignment, the user's recorded vocals (personal_voice.file) and instrumental file (accompany.file) are combined into a single audio file (audio.file). Then, using timestamp alignment again, the merged audio file (audio.file) and video file (video.file) are combined to generate a new MV. Upon completion, the user is prompted to save and publish the work, or to set it to save or publish automatically.
[0098] This application achieves an efficient online karaoke method that allows users to directly record and sing from existing music videos (MVs) through operations such as audio extraction, audio-visual separation, audio-visual synchronization, sound recording, and multi-track synthesis. This simplifies the user experience, satisfies users' desire to sing from all MVs, and allows them to publish the same MV, thus improving user satisfaction. Furthermore, for song audio, in response to a recording trigger for the currently playing song audio, the application can extract the accompaniment from the song audio to obtain an accompaniment file; play the accompaniment file while simultaneously starting the recording of the user's vocals to obtain the user's vocal audio; and synthesize the user's vocal audio and the accompaniment file to obtain the target song audio. In other words, even song audio that is not a video can be recorded without downloading the accompaniment file from the server.
[0099] When a user watches a music video (MV), if they are interested in the content and want to sing along with the MV's visuals and accompaniment, they simply click the record button. The mobile app extracts the audio from the MV, separates it from the video, further separates the vocals and accompaniment based on the extracted audio, and then plays the separated accompaniment and video footage synchronously. The user can then sing. After singing, the separated accompaniment, the user's vocals, and the video footage are combined to create a new MV for saving or publishing. This embodiment can temporarily convert various music videos, such as high-quality short videos created by users, into karaoke MVs for users, greatly enriching the user's visual experience of recording on their mobile phone and bringing more emotional value and spiritual enjoyment to their singing. Furthermore, technically, it is also possible to combine relevant video materials and the user's recorded singing to recreate a personalized MV. This solves the problem of users being able to record and sing karaoke based on MVs of interest even when there are no readily available accompaniments or MV resources with only accompaniment and no vocals.
[0100] This lowers the barrier to entry for online karaoke, allowing users to sing any music video they're interested in anytime, satisfying their needs. All operations in this solution are performed on existing music videos; no additional backing track files or music video files with only backing tracks and no vocals are required. It can generate the backing track files or music video resources needed for user recording. Any music video that a user is interested in can be recorded. Furthermore, it reduces server storage costs and file download bandwidth costs. Since the solution doesn't require additional backing track files or music video files with only backing tracks and no vocals, these files don't need to be stored on the server, reducing storage costs. Users also don't need to download these files during recording, reducing bandwidth costs and time spent on file downloads, thus improving the overall user recording experience.
[0101] This solution can be applied to music software applications such as music playback and karaoke recording. When users watch music videos and suddenly feel like singing, they can use this solution to record themselves singing. It can record songs from any music video, increasing user motivation and improving the overall recording experience. The technical solutions regarding audio extraction, audio-visual separation, audio-visual synchronization, and multi-track synthesis can also be applied to other similar scenarios, such as offline karaoke rooms, singing on in-car screens, chat rooms, and online education software.
[0102] See Figure 3 As shown, this application embodiment provides a music video generation apparatus, including:
[0103] The video switching module 11 is used to switch the first music video to a second music video in response to a recording trigger operation during the playback of the first music video; the second music video is the music video obtained by removing the first human voice audio from the first music video.
[0104] The audio recording module 12 is used to record audio during the playback of the second music video;
[0105] The video generation module 13 is used to display a third music video generated based on the second human voice audio and the second music video after the audio recording is completed.
[0106] The video switching module 11 may specifically include:
[0107] The audio and video extraction submodule is used to extract the audio file and video background file of the first music video being played.
[0108] The accompaniment extraction submodule is used to extract the accompaniment from the audio file to obtain the accompaniment file;
[0109] The video switching submodule is used to synchronize the playback of the video background file and the accompaniment file to switch the first music video to the second music video;
[0110] Correspondingly, the video generation module 13 also includes a synthesis submodule, which is used to synthesize the second human voice audio obtained from the recording, the video background file, and the accompaniment file after the audio recording is completed to obtain a third music video.
[0111] The accompaniment extraction submodule is specifically used to perform a sound-accompaniment separation operation on the audio file to obtain an initial accompaniment file; to perform human voice detection on the initial accompaniment file, and if human voice noise is detected, to remove the human voice noise from the initial accompaniment file to obtain the accompaniment file.
[0112] The video generation module 13 may also include:
[0113] The elimination submodule is used to determine the sound pickup device corresponding to the second human voice audio; if the sound pickup device is an external speaker, then the second human voice audio is subjected to echo cancellation and noise cancellation operations to obtain the eliminated second human voice audio; correspondingly, the synthesis submodule is specifically used to synthesize based on the eliminated second human voice audio, the video background file and the accompaniment file to obtain a third music video.
[0114] Furthermore, the composition submodule may include:
[0115] An audio synthesis unit is used to synthesize the second human voice audio and the accompaniment file into a single audio file to obtain a target audio file;
[0116] The video synthesis unit is used to synthesize the target audio file and the video background file to obtain a third music video.
[0117] The audio synthesis unit can be specifically used to: determine an auditory delay based on the timestamp corresponding to the human voice in the second human voice audio and the timestamp of the lyrics corresponding to the accompaniment file; cut the second human voice audio based on the auditory delay to obtain a cut second human voice audio; and synthesize the cut second human voice audio and the accompaniment file into a single audio track to obtain a target audio file.
[0118] Furthermore, the audio synthesis unit can be specifically used to: determine the timestamp of the first word of the human voice in the second human voice audio to obtain a first timestamp; determine the timestamp of the first word in the lyrics timestamp corresponding to the accompaniment file to obtain a second timestamp; and determine the auditory delay based on the first timestamp and the second timestamp.
[0119] Furthermore, the device is also used to: during the synchronous playback of the video background file and the accompaniment file, if a jump playback operation is detected, jump playback of the accompaniment file according to the timestamp corresponding to the jump playback operation, and simultaneously align the timestamp of the video background file with the accompaniment file for playback.
[0120] Furthermore, the device may also include:
[0121] A prompt module is used to generate prompt information and display it on the user interface, wherein the prompt information includes prompts for saving and / or publishing;
[0122] A post-processing module is used to save and / or publish the third music video in response to save and / or publish operations.
[0123] As can be seen, in this embodiment, during the playback of the first music video, in response to a recording trigger operation, the first music video is switched to a second music video obtained by removing the first vocal audio from the first music video. During the playback of the second music video, audio recording is performed. Furthermore, a third music video generated based on the recorded second vocal audio and the second music video is displayed. This eliminates the need to retrieve pre-made music video resources without vocals from the server. Processing and switching playback are performed directly based on the currently playing music video, while simultaneously recording vocals for subsequent synthesis. This avoids the resource limitations of the server, and the server does not need to store excessive pre-made music video resources, thus preventing the problem of limited pre-made music video resources on the server affecting user experience. It also reduces storage and download bandwidth costs.
[0124] See Figure 4 As shown in the figure, this application discloses an electronic device 20, including a processor 21 and a memory 22; wherein, the memory 22 is used to store a computer program; the processor 21 is used to execute the computer program, the music video generation method disclosed in the foregoing embodiment.
[0125] For details regarding the specific process of the above-mentioned music video generation method, please refer to the relevant content disclosed in the foregoing embodiments, which will not be repeated here.
[0126] Furthermore, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk, or optical disk, and the storage method can be temporary storage or permanent storage.
[0127] In addition, the electronic device 20 also includes a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26; wherein, the power supply 23 is used to provide operating voltage for the various hardware devices on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.
[0128] Furthermore, embodiments of this application also disclose a computer-readable storage medium for storing a computer program, wherein the computer program, when executed by a processor, implements the music video generation method disclosed in the foregoing embodiments.
[0129] For details regarding the specific process of the above-mentioned music video generation method, please refer to the relevant content disclosed in the foregoing embodiments, which will not be repeated here.
[0130] Furthermore, this application also provides a computer program product, including a computer program / instructions, which, when executed by a processor, implements the music video generation method disclosed in the foregoing embodiments.
[0131] For details regarding the specific process of the above-mentioned music video generation method, please refer to the relevant content disclosed in the foregoing embodiments, which will not be repeated here.
[0132] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.
[0133] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0134] The above provides a detailed description of a music video generation method, device, and medium provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A method of generating a music video, characterized by, The method comprises the following steps: In the process of playing the first music video, in response to a recording trigger operation, switching the playing first music video to a second music video; The second music video is a music video obtained by removing the first vocal audio in the first music video; In the process of playing the second music video, audio recording is performed; After the audio recording is completed, a third music video generated based on the second vocal audio obtained by recording and the second music video is displayed.
2. The music video generation method of claim 1, wherein, Switching the playing first music video to a second music video comprises the following steps: Extracting an audio file and a video background file of the playing first music video; Extracting accompaniment in the audio file to obtain an accompaniment file; Synchronously playing the video background file and the accompaniment file to switch the first music video to the second music video; Correspondingly, after the audio recording is completed, the method further comprises the following steps: Synthesizing the second vocal audio obtained by recording, the video background file and the accompaniment file to obtain a third music video.
3. The music video generation method of claim 2, wherein, Extracting accompaniment in the audio file to obtain an accompaniment file comprises the following steps: Performing a vocal accompaniment separation operation on the audio file to obtain an initial accompaniment file; Performing vocal detection on the initial accompaniment file, and if vocal noise is detected, removing the vocal noise from the initial accompaniment file to obtain an accompaniment file.
4. The music video generation method of claim 2, wherein, After the audio recording is completed, the method further comprises the following steps: Determining a sound pickup device corresponding to the second vocal audio; If the sound pickup device is an external device, performing echo cancellation and noise cancellation operations on the second vocal audio to obtain an eliminated second vocal audio; Correspondingly, synthesizing the second vocal audio obtained by recording, the video background file and the accompaniment file to obtain a third music video comprises the following steps: Synthesizing the eliminated second vocal audio, the video background file and the accompaniment file to obtain a third music video.
5. The music video generation method of claim 2, wherein, Synthesizing the second vocal audio obtained by recording, the video background file and the accompaniment file to obtain a third music video comprises the following steps: Synthesizing the second vocal audio and the accompaniment file into a single-track audio file to obtain a target audio file; Synthesizing the target audio file and the video background file to obtain a third music video.
6. The music video generation method of claim 5, wherein, Synthesizing the second vocal audio and the accompaniment file into a single-track audio file to obtain a target audio file comprises the following steps: Determining an auditory delay based on a time stamp of vocal in the second vocal audio and a time stamp of lyrics corresponding to the accompaniment file; Cutting the second vocal audio based on the auditory delay to obtain a cut second vocal audio; Synthesizing the cut second vocal audio and the accompaniment file into a single-track audio file to obtain a target audio file.
7. The music video generation method of claim 6, wherein, Determining an auditory delay based on a time stamp of vocal in the second vocal audio and a time stamp of lyrics corresponding to the accompaniment file comprises the following steps: Determining a time stamp of a first character of vocal in the second vocal audio to obtain a first time stamp; Determining a time stamp of a first character of lyrics corresponding to the accompaniment file to obtain a second time stamp; determine an auditory delay based on the first timestamp and the second timestamp.
8. The music video generation method of claim 2, wherein, Further comprising: In the process of synchronously playing the video background file and the accompaniment file, if a skip playing operation is detected, the accompaniment file is played by skipping according to the timestamp corresponding to the skip playing operation, and the timestamp of the video background file is aligned to play the accompaniment file.
9. The music video generation method according to any one of claims 1 to 8, characterized in that, After displaying a third music video generated based on the recorded second person audio and the second music video, further comprising: generating prompt information and displaying the prompt information on a user interface, wherein the prompt information includes a prompt to save and / or publish; in response to a save and / or publish operation, saving and / or publishing the third music video.
10. An electronic device, comprising: comprising a memory and a processor, wherein: the memory is configured to store a computer program; the processor is configured to execute the computer program to implement the music video generation method according to any one of claims 1 to 9.
11. A computer readable storage medium, characterized in that, a computer program product, wherein the computer program product is configured to implement the music video generation method according to any one of claims 1 to 9 when executed by a processor.