Song editing processing method and device, equipment, medium and product
By marking the accompaniment section in the song and recording the narration audio data for synthesis, the complexity of traditional online music creation technology has been solved, enabling convenient song editing on mobile terminals and expanding the audience and user activity of online music services.
Patent Information
- Application Number
- CN202210345844.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-31
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2042-03-31
AI Technical Summary
Traditional online music creation technologies are difficult to implement for all users, mobile devices are complex to operate, making it difficult for users to create high-quality songs and thus failing to unlock the potential of music services.
By identifying the accompaniment segments without vocals in the song to be edited, visualizing them, acquiring the narration audio data, and synthesizing it into the song, a simple interactive method is provided for recording and synthesizing narration audio data.
The ability to easily and efficiently add user-generated narration to mobile devices enriches song content, lowers the barrier to user participation, expands the audience and user activity of online music services, and promotes the development of online music social activities.
Smart Images

Figure CN114863899B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of music information processing technology, and in particular to a song editing and processing method and its corresponding apparatus, computer equipment, computer-readable storage medium, and computer program product. Background Technology
[0002] Online music has enriched people's spiritual and cultural lives, and has therefore flourished. Music-assisted creation technology, by providing various operational conveniences, makes it easier for people to showcase their musical, literary, and performing talents, opening up greater opportunities for social workers to realize their value and expand social employment.
[0003] Traditional online music creation technology typically provides services for songs and lyrics. It offers users a user interface on a computer, providing tools for editing sheet music and lyrics to capture data corresponding to the user's creative ideas. This data is then processed in the background to produce the final song. However, due to the inherent professionalism and personal musicality required for music, the average user still cannot create high-quality songs independently. Therefore, the application of such music creation technology remains relatively niche, failing to truly achieve widespread participation in music creation and thus failing to unlock deeper social value.
[0004] On the other hand, terminal devices have shifted from personal computers to mobile devices such as mobile phones and tablets. These mobile devices generally have small screens and are inconvenient to operate. Therefore, music creation aids based on sheet music and lyrics are difficult to effectively utilize on mobile devices due to the complexity and professionalism required by music theory.
[0005] Based on the above aspects, it is difficult for online music creation technology to remain within the existing model to fully realize the potential of online music services. Therefore, the development of related technologies needs to explore new avenues. Summary of the Invention
[0006] The primary objective of this application is to solve at least one of the aforementioned problems by providing a song editing processing method and corresponding apparatus, computer equipment, computer-readable storage medium, and computer program product.
[0007] To achieve the various objectives of this application, the following technical solution is adopted:
[0008] A song editing method provided for one of the purposes of this application includes the following steps:
[0009] Identify the instrumental section corresponding to the vocal parts of the song to be edited;
[0010] The accompanying section of the song to be edited is visually marked;
[0011] Responding to the narration acquisition event, acquire the narration audio data corresponding to the accompaniment segment;
[0012] In response to the narration synthesis event, the narration audio data is synthesized into the audio data of the song to be edited.
[0013] In one of the more detailed embodiments, determining the accompaniment segment corresponding to the vocalless portion of the song to be edited includes the following steps:
[0014] Obtain the lyrics data and audio data of the song to be edited. The lyrics data includes the lyrics and the timestamps of the corresponding vocal parts in the audio data.
[0015] Based on the timestamps corresponding to the lyrics, the candidate accompaniment segments in the audio data are calculated, and the playback duration corresponding to each candidate accompaniment segment is determined.
[0016] Candidate accompaniment segments with a playback duration exceeding a preset threshold are selected as valid accompaniment segments.
[0017] In another embodiment of the refinement, the accompaniment segment corresponding to the vocalless portion of the song to be edited is determined, including the following steps:
[0018] Obtain the audio data of the song to be edited;
[0019] Human voice detection is performed on the audio data to identify multiple candidate accompaniment segments corresponding to the non-vocal singing content, and the playback duration of each candidate accompaniment segment is determined.
[0020] Candidate accompaniment segments with a playback duration exceeding a preset threshold are selected as valid accompaniment segments.
[0021] In one of the more detailed embodiments, the accompaniment section of the song to be edited is visualized, including the following steps:
[0022] The timeline corresponding to the song to be edited is displayed, and the timeline is used to show the playback duration information of the song to be edited;
[0023] The indexing position of the accompaniment segment on the time axis is determined based on the time range of the accompaniment segment;
[0024] A narration capture control corresponding to the accompaniment segment is displayed at the index position to trigger a narration capture event in response to user touch.
[0025] In a further embodiment, in response to a narration acquisition event, acquiring the narration audio data corresponding to the accompaniment segment includes the following steps:
[0026] In response to the narration acquisition event corresponding to the accompaniment segment, audio data recording begins;
[0027] The recorded audio data is stored as the corresponding narration audio data for the accompaniment segment;
[0028] The accompanying music segment is associated with a narration audio data visualization control.
[0029] In a specific embodiment, in the step of starting to record audio data in response to the narration acquisition event corresponding to the accompaniment segment, the recording duration is constrained to not exceed the playback duration of the accompaniment segment.
[0030] In some extended embodiments, after the step of visualizing the narration audio data associated with the accompaniment segment as a narration control, the following steps are included:
[0031] In response to a movement event applied to any narration indicator control, the offset of the narration audio data of the narration indicator control relative to its accompaniment segment is adjusted according to the corresponding movement amount.
[0032] In some extended embodiments, after the step of visualizing the narration audio data associated with the accompaniment segment as a narration control, the following steps are included:
[0033] In response to a song playback event, the audio data of the song to be edited and the narration audio data of the accompaniment segment are played synchronously according to the timing alignment relationship.
[0034] In a specific embodiment, after the step of starting audio data recording in response to the narration acquisition event corresponding to the accompaniment segment, the following steps are included:
[0035] Automatic speech recognition is performed on the recorded audio data to obtain the narration statements corresponding to the narration audio data. The narration statements are associated with timestamps that correspond to the playback time of the song to be edited.
[0036] In some extended embodiments, after the step of synthesizing the narration audio data into the audio data of the song to be edited in response to the narration synthesis event, the following steps are included:
[0037] The narration sentences generated by speech recognition of each narration audio data are associated with their corresponding timestamps and added to the lyrics data of the song to be edited to obtain the synthesized lyrics data.
[0038] In a further extended embodiment, after the step of synthesizing the narration audio data into the audio data of the song to be edited in response to the narration synthesis event, the following steps are included:
[0039] Responding to the playback event after synthesis, the synthesized audio data is played, and the narration and lyrics in the synthesized lyrics data are displayed synchronously according to the timestamp;
[0040] In response to the song release event, the synthesized audio data is published online.
[0041] A song editing and processing apparatus provided for one of the purposes of this application includes:
[0042] The accompaniment analysis module is used to determine the accompaniment section corresponding to the vocal-free parts of the song to be edited;
[0043] The accompaniment marking module is used to visually mark the accompaniment segments of the song to be edited;
[0044] The narration acquisition module is used to respond to the narration acquisition event and acquire the narration audio data corresponding to the accompaniment segment;
[0045] The song synthesis module is used to respond to the narration synthesis event and synthesize the narration audio data into the audio data of the song to be edited.
[0046] A computer device provided for one of the purposes of this application includes a central processing unit and a memory, wherein the central processing unit is used to invoke and run a computer program stored in the memory to perform the steps of the song editing processing method described in this application.
[0047] A computer-readable storage medium is provided for another purpose of this application, which stores, in the form of computer-readable instructions, a computer program implemented according to the described song editing processing method, which, when invoked by a computer, performs the steps included in the method.
[0048] A computer program product provided for another purpose of this application includes a computer program / instructions that, when executed by a processor, implement the steps of the song editing processing method described in any embodiment of this application.
[0049] Compared with existing technologies, this application has many technical advantages, including but not limited to the following aspects:
[0050] First, this application identifies the accompaniment section corresponding to the vocal-free parts of the song to be edited and visually marks it for user operation. Based on this, it obtains the narration audio data corresponding to the accompaniment section by responding to the narration acquisition event. Finally, it synthesizes the narration audio data into the audio data of the song to be edited by responding to the narration synthesis event. This opens up a music-assisted creation technology framework that is different from the traditional one, allowing users to easily and efficiently add their own created narration content to the song on the terminal device, thereby enriching the content of the music and increasing the information content of the music.
[0051] Secondly, after determining the accompaniment section of the song to be edited, this application can record narration audio data by responding to the narration acquisition event based on simple interaction, and update the song to be edited by responding to the narration synthesis event. The interaction is simple and convenient to implement on mobile terminal devices, which lowers the threshold for user participation and helps to improve the coverage of music-assisted creation.
[0052] Furthermore, this application only provides services by adding corresponding narration audio data to the accompaniment section of a song. The narration content is simpler than dictionary creation, which can allow more users with basic literary literacy to participate, thereby expanding the audience of online music services, activating user traffic on online music platforms, improving their daily active users and retention rates, and promoting the development of online music social activities, thus bringing many positive social benefits. Attached Figure Description
[0053] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:
[0054] Figure 1 A typical network deployment architecture diagram related to the implementation of the technical solution of this application;
[0055] Figure 2 This is a flowchart illustrating a typical embodiment of the song editing method of this application;
[0056] Figures 3 to 10 To implement the exemplary graphical user interface obtained in this application, wherein: Figure 3 Display song list, Figure 4 Show the music playback interface with narration editing controls. Figure 5 This shows an example of marking multiple accompaniment sections. Figure 6 This example shows how multiple backing tracks are stored in the sidebar. Figure 7 According to Figure 6 The effect of the sidebar after it expands. Figure 8 This example demonstrates how to pop up a half-window for audio recording for one of the backing tracks. Figure 9 This demonstrates the indicator effect after the accompaniment section has been recorded. Figure 10 The effect of displaying the narration synthesis control after all the accompaniment sections have been recorded is shown;
[0057] Figure 11 and Figure 12 This is a flowchart illustrating the process of obtaining a valid accompaniment segment under different embodiments of this application;
[0058] Figure 13 This is a flowchart illustrating the process of visualizing the accompaniment section in an embodiment of this application.
[0059] Figure 14 To implement the exemplary graphical user interface obtained in this application, the accompanying music segments corresponding to each narration are shown in blank rectangular bars parallel to the timeline of the song to be edited;
[0060] Figure 15 This is a flowchart illustrating the process of acquiring narration audio data in an embodiment of this application;
[0061] Figure 16 To implement the exemplary graphical user interface obtained in this application, an audio recording half-window pops up for an accompaniment segment;
[0062] Figure 17 This is a schematic block diagram of the song editing and processing device of this application;
[0063] Figure 18 This is a schematic diagram of the structure of a computer device used in this application. Detailed Implementation
[0064] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain this application, and should not be construed as limiting this application.
[0065] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this application means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we say an element is “connected” or “coupled” to another element, it can be directly connected or coupled to the other element, or there may be intermediate elements. Furthermore, “connected” or “coupled” as used herein can include wireless connections or wireless coupling. The term “and / or” as used herein includes all or any units and all combinations of one or more associated listed items.
[0066] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the same meaning as in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless specifically defined as herein.
[0067] Those skilled in the art will understand that the terms "client," "terminal," and "terminal device" as used herein include both devices that receive wireless signals, devices that only possess wireless signal receiver capabilities without transmission capabilities, and devices with receiving and transmitting hardware, devices that have receiving and transmitting hardware capable of bidirectional communication over a bidirectional communication link. Such devices may include: cellular or other communication devices such as personal computers or tablets, having single-line displays, multi-line displays, or cellular or other communication devices without multi-line displays; PCS (Personal Communications Service) that can combine voice, data processing, fax, and / or data communication capabilities; PDA (Personal Digital Assistant) that may include a radio frequency receiver, pager, internet / intranet access, web browser, notepad, calendar, and / or GPS (Global Positioning System) receiver; and conventional laptops and / or handheld computers or other devices that have and / or include radio frequency receivers. As used herein, "client," "terminal," and "terminal device" can be portable, transportable, installed in a means of transportation (air, sea, and / or land), or suitable and / or configured to operate locally and / or in a distributed manner, operating in any other location on Earth and / or in space. "Client," "terminal," and "terminal device" as used herein can also be a communication terminal, an internet access terminal, or a music / video playback terminal, such as a PDA, a MID (Mobile Internet Device), and / or a mobile phone with music / video playback capabilities, or a smart TV, set-top box, etc.
[0068] The hardware referred to by the names "server," "client," and "service node" in this application is essentially an electronic device with the equivalent capabilities of a personal computer. It is a hardware device with the necessary components revealed by the von Neumann architecture, such as a central processing unit (including an arithmetic logic unit and a control unit), memory, input devices, and output devices. The computer program is stored in its memory, and the central processing unit loads the program stored in the secondary storage into the main memory to run it, execute the instructions in the program, and interact with the input and output devices to complete specific functions.
[0069] It should be noted that the concept of "server" used in this application can also be extended to the case of server clusters. Based on the network deployment principles understood by those skilled in the art, the servers should be logically divided. Physically, these servers can be independent of each other but accessible through interfaces, or they can be integrated into a single physical computer or a computer cluster. Those skilled in the art should understand this flexibility and should not use it to constrain the implementation of the network deployment method in this application.
[0070] Please see Figure 1 The hardware infrastructure required for implementing the technical solutions of this application can be deployed according to the architecture shown in the figure. The server 70 mentioned in this application is deployed in the cloud and acts as a business server. It can further connect to relevant data servers and other servers providing related support, thereby forming a logically related service cluster to provide services to relevant terminal devices such as the smartphone 71 and personal computer 72 shown in the figure, or third-party servers (not shown). Both the smartphone and personal computer can access the Internet through known network access methods and establish a data communication link with the cloud server 70 to run terminal applications related to the services provided by the server.
[0071] For servers, the application is usually built as a service process, with corresponding program interfaces exposed for remote calls by applications running on various terminal devices. The relevant technical solutions in this application that are suitable for running on servers can be implemented in servers in this way.
[0072] The application mentioned refers to an application running on a server or terminal device. This application implements the relevant technical solutions of this application in a programmed manner. Its program code can be stored in a non-volatile storage medium that can be recognized by a computer in the form of computer-executable instructions, and is loaded into memory by the central processing unit for execution. The relevant device of this application is constructed by the operation of the application on the computer.
[0073] One or more of the technical features of this application, unless explicitly specified herein, can be deployed on a server and accessed by a client remotely calling the online service interface provided by the server, or can be directly deployed and run on a client for access.
[0074] Unless otherwise specified, the neural network models referenced or potentially referenced in this application may be deployed on a remote server and invoked remotely on the client, or deployed on a client with the capability to invoke directly. In some embodiments, when running on the client, the corresponding intelligence may be acquired through transfer learning in order to reduce the requirements on the client's hardware resources and avoid excessive consumption of the client's hardware resources.
[0075] Unless otherwise specified, all data involved in this application may be stored remotely on a server or on a local terminal device, as long as it is suitable for use by the technical solution of this application.
[0076] Those skilled in the art will understand that although the various methods in this application are described based on the same concept and thus present commonality among them, they can be performed independently unless otherwise specified. Similarly, the various embodiments disclosed in this application are all based on the same inventive concept; therefore, concepts expressed in the same way, as well as concepts that are appropriately changed for convenience but are expressed differently, should be understood equivalently.
[0077] Unless otherwise expressly stated, the various embodiments disclosed in this application can be combined in a cross-cutting manner to flexibly construct new embodiments, as long as such combination does not depart from the inventive spirit of this application and can meet the needs of the prior art or solve a certain deficiency in the prior art. Those skilled in the art should be aware of such modifications.
[0078] The song editing processing method of this application can be programmed into a computer program product, which is mainly deployed and run on a terminal device so that the method can be executed by human-computer interaction with the computer program product through a graphical user interface by accessing the interface opened after the computer program product is run.
[0079] Please see Figure 2 The song editing processing method of this application, in its typical embodiment, includes the following steps:
[0080] Step S1100: Determine the accompaniment section corresponding to the vocalless parts of the song to be edited.
[0081] Figure 3An exemplary graphical user interface is a music list interface displayed on a terminal device such as a mobile phone or tablet after an online music application is running. Through this music list interface, the user can select any song and touch the corresponding narration editing control to use that song as the song to be edited in this application, and then execute the various steps of this application.
[0082] Figure 4 The exemplary graphical user interface is based on Figure 3 The music playback interface displayed after selecting and playing a song can also provide narration editing controls for the user to manipulate, so that the playing song can be used as the song to be edited in this application, and the various steps of this application can be executed accordingly.
[0083] In this application, when the song to be edited is played, it generally outputs accompaniment music and vocal performance. These two parts may overlap or stagger in timing. Their corresponding audio data are combined and encapsulated into the audio data of the song to be edited. Therefore, observing the playback sequence, it can be seen that the song to be edited generally includes an accompaniment section without vocal performance and a main melody section with vocal performance. It should be noted that in certain special scenarios, such as when narration is added first and then the vocal performance, the main melody section of the song to be edited is allowed to be generated using MIDI music to replace the corresponding vocal performance. In this case, it should also be considered as the vocal performance section in this application.
[0084] Once a user has selected a song to edit, they can download the relevant data, including audio and lyrics, from the online music service server. This data is typically packaged as files and downloaded to local storage. If the audio and lyrics data are already stored on local storage or in a local cache, they can be accessed directly.
[0085] Furthermore, by identifying the song to be edited, one or more accompaniment segments corresponding to the vocalless parts of the song can be determined. The accompaniment segments can be determined based on either the lyrics data or the audio data of the song. Since the lyrics data generally contains timestamps corresponding to the main melody, these timestamps can be used to effectively identify the accompaniment segments. In the audio data, the timestamps corresponding to the appearance of vocals can be determined using various known vocal detection methods, thereby identifying the accompaniment segments. Therefore, based on the principles disclosed herein, those skilled in the art can flexibly implement methods to determine the accompaniment segments corresponding to the vocalless parts of the song to be edited.
[0086] In the same song, there may be multiple accompaniment sections, and the distribution of each accompaniment section in the song to be edited is generally discrete.
[0087] Step S1200: Visually mark the accompaniment section of the song to be edited:
[0088] Please see Figure 5 The example graphical user interface (GUI) demonstrates how, after identifying one or more backing tracks from a song to be edited, these tracks can be displayed in the GUI in the form of icons or text. They are generally arranged in order of their corresponding playback sequence within the song. If necessary, appropriate prompts can be provided, such as indicating "Intro," "Bridge," or "Ending." The labels for each backing track can be represented and loaded as controls; then, the corresponding icons or text can be loaded into these controls.
[0089] like Figure 6 The example graphical user interface demonstrates another form of indicating the accompaniment segments, where only a notification icon indicating that the accompaniment segment identification has been completed is displayed in the graphical user interface. When the user touches this notification icon, the controls corresponding to each accompaniment segment expand in response to the touch event, such as... Figure 7 As shown.
[0090] As can be seen, the accompaniment segment of the song to be edited can be displayed through the above methods. Further embodiments that enhance the user experience will be revealed later through other examples, which will not be discussed here. In addition, those skilled in the art can, of course, flexibly adapt the principles and examples disclosed herein to achieve on-demand display of the accompaniment segment.
[0091] When indicating the accompanying instrumental segments, the timestamps of the playback sequence of each segment corresponding to the song to be edited can also be marked. Specifically, this refers to the start and end timestamps of each instrumental segment in the playback sequence of the song to be edited. The format of the start and end timestamps can be as follows: Figure 5 The timestamps shown can be used directly, or they can be used to mark a partial segment of the playback sequence of the song to be edited. Those skilled in the art can implement this flexibly. It is easy to understand that the start and end timestamps corresponding to the accompaniment segment together define the time range of the accompaniment segment in the playback sequence of the song to be edited, that is, define the playback duration occupied by the entire accompaniment segment. These data have been obtained in advance in the previous step.
[0092] Step S1300: Respond to the narration acquisition event and acquire the narration audio data corresponding to the accompaniment segment:
[0093] In one embodiment, a single narration acquisition event can be triggered for multiple accompaniment segments, while the corresponding narration audio data for each accompaniment segment can be acquired at once. To this end, the user can trigger the narration acquisition event by touching any control representing the accompaniment segment or other dedicated controls.
[0094] In another embodiment, the narration audio data can be acquired separately for each accompaniment segment. Accordingly, controls used to represent each accompaniment segment can respond to user touch operations, triggering the corresponding narration acquisition event for the touched accompaniment segment.
[0095] In addition, in other embodiments, the narration acquisition event can also be triggered by pressing a button, shaking the terminal device, or any other triggering method. Those skilled in the art can make corresponding presets based on the principles disclosed in this application.
[0096] When the narration acquisition event is triggered, in response to the narration acquisition event, the process of acquiring the corresponding accompaniment segment is executed, such as... Figure 8 As an example, a recording interface can pop up in the graphical user interface of the terminal device, guiding the user to start recording audio or retrieve pre-stored audio data as narration audio data for the corresponding accompaniment segment of the event. For example, the user can touch... Figure 8 The recording interface shown uses a recording button to start recording, thus beginning the collection of user voice data until the user ends recording or the preset recording duration is reached. The recording duration can be determined based on the time range occupied by the corresponding accompaniment segment, and is generally set to an equal length. If the audio data recorded by the user without a set recording duration is relatively short compared to the time range occupied by the accompaniment segment, the recorded audio data can be discarded, and the user can be prompted to re-record. If the time range is relatively long, the recorded audio data can be truncated according to the time range of the corresponding accompaniment segment to obtain audio data of equal length as the narration audio data. Whether the recorded audio data is too short can be determined by those skilled in the art by setting a certain duration threshold or referring to a preset proportion of the duration defined by the time range of the accompaniment segment. If the duration of the recorded audio data is less than the duration threshold or preset proportion, it can be considered too short, discarded, and a re-recording is required.
[0097] The narration audio data refers to the recorded content, which is generally the speech of the user on the current terminal device. This content can constitute the narration part of the song to be edited. The content corresponding to the narration can be diverse. For example, it can be personal reflections based on the song lyrics, a story narration carefully crafted by the user with reference to the imagery of the lyrics, onomatopoeic voices to enrich the atmosphere of the song, or even environmental sounds without corresponding text generated by other animals or settings. All of these can be used as the narration part of the song to be edited, playing a positive role in enriching the song's content and enhancing its artistry. It should be noted that although this application technically implements a mechanism for acquiring and responding to narration audio data, if the user intentionally inputs disordered sounds into the terminal device during the process of acquiring narration audio data for the accompaniment segment in response to the narration acquisition event, this human intervention does not affect the technical implementation of this application, and therefore does not affect the inventiveness of this application.
[0098] In some embodiments, the narration audio data may be pre-recorded and archived, therefore... Figure 8 In the example interface, the corresponding voice data can be obtained as the narration audio data simply by using "local call" or "remote call". If necessary, an editing interface can also be provided for this voice data during the call process, so that users can extract a portion of the longer voice data that is the same length as the corresponding accompaniment segment as the corresponding narration audio data.
[0099] The narration audio data recorded by the user for each accompaniment segment can be associated with its corresponding accompaniment segment and stored in a cache for direct retrieval later. Furthermore, the graphical user interface can display status information indicating whether the narration audio data acquisition for the corresponding accompaniment segment has been completed. Figure 9 The example shown is displayed in the legend of the control corresponding to the accompaniment segment, or it may be presented in other ways.
[0100] Step S1400: In response to the narration synthesis event, synthesize the narration audio data into the audio data of the song to be edited.
[0101] To facilitate guiding users in synthesizing narration audio data with the song to be edited, in one embodiment, it can be done as follows: Figure 10As exemplified, after obtaining the narration audio data for any accompaniment segment, a narration synthesis control is made available to the user in the graphical user interface. This narration synthesis control is used to trigger a narration synthesis event to synthesize the various narration audio data with the song to be edited. In another embodiment, the user can be allowed to trigger the narration synthesis event first using any interaction method, such as shaking the phone or touching a button. Then, each accompaniment segment is detected. Only after confirming that at least one accompaniment segment has completed the acquisition of narration audio data is the process of synthesizing the various narration audio data into the song to be edited executed.
[0102] Since each accompaniment segment is associated with its start and end timestamps relative to the playback sequence of the song to be edited, the duration of its narration audio data is controlled within the time range of the accompaniment segment. Therefore, when merging the narration audio data into the song to be edited, the narration audio data can be time-aligned with the song to be edited according to the time range of the corresponding accompaniment segment. Typically, the narration audio data can be aligned with the playback sequence corresponding to the start timestamp of its respective accompaniment segment. Then, the narration audio data is merged with the song to be edited to obtain the merged audio data.
[0103] The technique for synthesizing the narration audio data with the song to be edited can be implemented using various traditional audio editing or synthesis algorithms, or by calling the open interface provided by audio editing software. It can be flexibly implemented by those skilled in the art and does not affect the inventive spirit of this application.
[0104] After the narration audio data is synthesized with the song to be edited, the synthesized audio data is obtained. Based on this, the user can further play the audio data to listen to the synthesized song. During the playback, the song not only includes the original vocal part of the song to be edited, but also allows the user to hear the playback effect of the narration audio data in the accompaniment section.
[0105] As can be seen from the typical embodiments and various variations of this application, compared with the prior art, this application has many technical advantages, including but not limited to the following aspects:
[0106] First, this application identifies the accompaniment section corresponding to the vocal-free parts of the song to be edited and visually marks it for user operation. Based on this, it obtains the narration audio data corresponding to the accompaniment section by responding to the narration acquisition event. Finally, it synthesizes the narration audio data into the audio data of the song to be edited by responding to the narration synthesis event. This opens up a music-assisted creation technology framework that is different from the traditional one, allowing users to easily and efficiently add their own created narration content to the song on the terminal device, thereby enriching the content of the music and increasing the information content of the music.
[0107] Secondly, after determining the accompaniment section of the song to be edited, this application can record narration audio data by responding to the narration acquisition event based on simple interaction, and update the song to be edited by responding to the narration synthesis event. The interaction is simple and convenient to implement on mobile terminal devices, which lowers the threshold for user participation and helps to improve the coverage of music-assisted creation.
[0108] Furthermore, this application only provides services by adding corresponding narration audio data to the accompaniment section of a song. The narration content is simpler than dictionary creation, which can allow more users with basic literary literacy to participate, thereby expanding the audience of online music services, activating user traffic on online music platforms, improving their daily active users and retention rates, and promoting the development of online music social activities, thus bringing many positive social benefits.
[0109] Please see Figure 11 In a further embodiment, step S1100, determining the accompaniment segment corresponding to the vocalless portion of the song to be edited, includes the following steps:
[0110] Step S1111: Obtain the lyrics data and audio data of the song to be edited. The lyrics data includes the lyrics and the timestamps of the corresponding vocal parts in the audio data.
[0111] When the user is Figure 3 When selecting a song to be edited in the example music list interface to begin the narration audio editing process of this application, the application background process can first detect whether the lyrics data and audio data of the song to be edited have been cached in the local cache. If they already exist, they can be directly called; otherwise, the corresponding lyrics data or audio data can be downloaded from the server providing the online music service.
[0112] The lyrics data can be in various common lyrics format files. These lyrics format files generally contain lyrics of the main melody section corresponding to the vocal parts of the song to be edited. Based on the time correspondence between the vocal parts of each lyrics statement and the playback sequence of the song to be edited, the lyrics format file is associated with the corresponding timestamp of each lyrics statement in the audio data of the song to be edited.
[0113] The audio data can be used to determine the total playback duration of the song to be edited, and can be used to establish the correspondence between the subsequently detected accompaniment segment and the song to be edited in terms of playback sequence.
[0114] Step S1112: Calculate the candidate accompaniment segments in the audio data based on the timestamps corresponding to the lyrics, and determine the playback duration corresponding to each candidate accompaniment segment;
[0115] Since the lyrics in the lyrics data are all marked with corresponding timestamps, the lyrics data of the song to be edited can be used to determine the corresponding time periods of the vocal parts, that is, the main melody sections, and then the accompaniment sections can be further determined.
[0116] It should be noted that in the lyrics format file, the lyrics of each main melody section are generally relatively continuous when sung. Although there may be brief breath transitions or pauses in between, these transitions or pauses are negligible. However, the time interval between different main melody sections may be relatively large. Therefore, for the accompaniment section in the middle of a song with a defined bridge-like nature:
[0117] First, based on the time distance between the timestamps of two adjacent lyric phrases in the lyric format file, it can be roughly estimated whether the two phrases belong to the same main melody segment, so as to determine the accompaniment segment between the two phrases based on the two adjacent main melody segments.
[0118] An example excerpt of lyrics is as follows:
[0119] "...
[0120] [00:05:10] The sun shines brightly.
[0121] [00:08:30]The flowers are smiling at me
[0122] [00:09:50] The little bird said, "Good morning, good morning!"
[0123] [00:12:30]Why are you carrying a small backpack?
[0124] ...
[0125] [00:30:10] I'm going to school
[0126] ..."
[0127] As shown in the excerpt above, each lyric phrase is marked with a corresponding timestamp. The first phrase, "The sun shines brightly," begins at 5:10 in the song to be edited, while the second phrase, "The flowers smile at me," begins at 8:30. The breath transition between these two phrases is relatively subtle. Therefore, a preset time threshold can be used to filter out such situations. That is, for two adjacent lyric phrases in the playback sequence, if the difference between their timestamps is less than the preset time threshold (e.g., 5 seconds), the two lyric phrases are considered to be part of the same main melody, and no accompaniment is added between them. If the difference between their timestamps is greater than or equal to the preset time threshold, then an accompaniment is considered to exist between them. In the example lyrics, the time difference between the timestamps [00:12:30] and [00:30:10] exceeds 18 seconds. Logically, the lyrics corresponding to the two timestamps should belong to different main melody sections. There may be an accompaniment section between the two main melody sections. Theoretically, the timestamps of the preceding and following lyrics can be used as the start and end timestamps of the accompaniment section to determine the time range of the accompaniment section.
[0128] However, since there is also a lyric "Why are you carrying a small schoolbag" at the corresponding location [00:12:30], directly setting the timestamp [00:12:30] as the start timestamp of the accompaniment segment would cause the accompaniment segment and the lyric to overlap in playback sequence. Sometimes this situation is undesirable, therefore:
[0129] Secondly, to make this estimation more accurate, the time range of the accompaniment segment, which is a transitional section, can be adjusted by relating it to the number of words in the lyrics. Specifically, the unit time occupied by each word can be estimated by dividing the total number of words in a main melody segment by the total duration of the entire main melody segment. For example, in the illustrative lyrics, the total number of words in the last lyric of the main melody segment is then multiplied by the unit time occupied by each word. An allowance of, for example, 1 second is added to obtain the estimated duration of the last lyric. This estimated duration is then added to the timestamp of the lyric, resulting in an estimated timestamp indicating when the lyric ends. This timestamp is then used as the start timestamp of the accompaniment segment following the lyric, while the end timestamp of the accompaniment segment is based on the timestamp of the first lyric of the next main melody segment. This allows for a relatively accurate determination of the time range of the transitional accompaniment segment.
[0130] For the instrumental section at the beginning of the song to be edited, its starting timestamp can be determined as [00:00:00], which is the beginning of the entire song, while its ending timestamp can be based on the timestamp of the first lyric line of the song.
[0131] For the accompaniment section at the end of the song to be edited, its end timestamp can be based on the timestamp corresponding to the end of the song's playback sequence, while its start timestamp can be determined by using the method of determining the start timestamp of the accompaniment section between the two main melody sections mentioned above.
[0132] Following the above method, the various accompaniment sections in the song to be edited can be determined based on the lyrics data. Usually, there are multiple accompaniment sections in each song, and the playback duration of each accompaniment section varies depending on its time range. Therefore, the accompaniment sections determined according to the above process can be used as candidate accompaniment sections for further selection in the future.
[0133] It should be noted that there are various flexible implementations for determining multiple candidate accompaniment segments based on the timestamps of the lyrics. Those skilled in the art can implement these methods flexibly based on the principles disclosed herein, such as adapting and adjusting various time thresholds in the above example process.
[0134] Step S1113: Select candidate accompaniment segments with a playback duration higher than a preset threshold as valid accompaniment segments:
[0135] It's easy to understand that if the accompaniment segment is short, for example, less than 5 seconds, it might prevent the user from effectively expressing their creative message. In this case, another preset threshold can be used to filter the candidate accompaniment segments identified in the previous step, selecting those with a playback duration higher than the preset threshold as valid accompaniment segments and discarding the others. Accordingly, it's easy to understand that the preset threshold set here is essentially the same in nature as the time threshold used to determine whether two adjacent lyric phrases belong to the same main melody segment. Therefore, in alternative embodiments, if the time threshold was applied in the previous step, segments with a time difference between two adjacent lyric phrases higher than the time threshold can be identified as valid accompaniment segments.
[0136] Therefore, it can be understood that the start and end timestamps of each valid accompaniment segment have been determined accordingly, and the playback duration of the valid accompaniment segment is also determined by the difference between its end and start timestamps.
[0137] In this embodiment, the accompaniment segment can be determined and optimized using the lyrics data of the song to be edited, ultimately identifying one or more valid accompaniment segments for adding corresponding narration audio data. The lyrics file is generally associated with the song's audio file, making it easy to access. Furthermore, calculating the accompaniment segment from the lyrics file requires minimal computation and is highly efficient and fast. Therefore, the technical solution of this application is easier to implement efficiently, resulting in a better user experience.
[0138] Please see Figure 12 In another embodiment of the refinement, step S1100, determining the accompaniment segment corresponding to the vocalless portion of the song to be edited, includes the following steps:
[0139] Step S1121: Obtain the audio data of the song to be edited:
[0140] In this embodiment, only the audio data of the song to be edited needs to be obtained. The audio data can be an audio file that has been downloaded locally, or it can be obtained by downloading the corresponding audio file from a server that provides online music services.
[0141] Step S1122: Perform human voice detection on the audio data to identify multiple candidate accompaniment segments corresponding to the non-vocal singing content, and determine the corresponding playback duration for each candidate accompaniment segment:
[0142] In one implementation, the audio data can first be separated into audio tracks to separate the accompaniment audio and the vocal audio. Then, by performing voice activity detection (VAD) on the vocal audio, the start and end timestamps corresponding to the parts of the vocal audio without vocal singing content can be determined, thereby identifying multiple candidate accompaniment segments.
[0143] In another implementation, a neural network model pre-trained to convergence can be used to detect human voices in the audio data. This neural network model is trained to obtain the ability to detect the start and end timestamps corresponding to the human voice singing content in the audio data, thereby determining multiple candidate accompaniment segments corresponding to the human voice singing part.
[0144] Based on the principles disclosed herein, those skilled in the art can use any conventional voice detection technique to determine the multiple candidate accompaniment segments from the audio data of the song to be edited.
[0145] Step S1123: Select candidate accompaniment segments with a playback duration higher than a preset threshold as valid accompaniment segments:
[0146] Similarly, since the playback duration of the accompaniment segments varies, effective accompaniment segments that meet the preset target can be selected from multiple candidate accompaniment segments. Therefore, a preset threshold can be used to filter each candidate accompaniment segment determined in the previous step, and the candidate accompaniment segments with a playback duration higher than the preset threshold can be selected as effective accompaniment segments, while other candidate accompaniment segments are discarded.
[0147] Therefore, it can be understood that the start and end timestamps of each valid accompaniment segment have been determined accordingly, and the playback duration of the valid accompaniment segment is also determined by the difference between its end and start timestamps.
[0148] In this embodiment, by using traditional human voice detection technology, the effective accompaniment segment is automatically determined from the audio data of the song to be edited, without relying on the lyrics data of the song to be edited. Therefore, it can be compatible with a wider range of scenarios. As long as the audio data exists, the accompaniment segment of the part without human vocals can be detected.
[0149] Please see Figure 13 In a further embodiment, step S1200, visually marking the accompaniment segment of the song to be edited, includes the following steps:
[0150] Step S1210: Display the timeline corresponding to the song to be edited. The timeline is used to display the playback duration information of the song to be edited.
[0151] In order to achieve effective identification of the accompaniment segment, in this embodiment, as follows: Figure 14 As exemplified, a playback component for the song to be edited is displayed in the graphical user interface of a terminal device, preparing the song to be edited for playback control via various controls in the playback component. The playback component provides a timeline constrained to the full playback duration of the song to be edited and uses a cursor to indicate the playback progress of the song in the playback sequence, thereby providing various playback duration information.
[0152] Step S1220: Determine the index position of the accompaniment segment on the time axis based on the time range of the accompaniment segment:
[0153] As mentioned above, after each accompaniment segment is determined in the various embodiments of this application, which mainly refer to the effective accompaniment segments, the time range of each accompaniment segment is determined. Therefore, its start timestamp and end timestamp have been determined. Accordingly, the corresponding positions of these timestamps on the timeline can be determined. The distance between the start timestamp and end timestamp of the same accompaniment segment on the timeline can be regarded as the indexing position of the corresponding accompaniment segment. The accompaniment segment can be prompted at this indexing position later.
[0154] Step S1230: Display the narration acquisition control corresponding to the accompaniment segment at the index position, which is used to trigger the narration acquisition event in response to user touch.
[0155] When indexing each accompaniment segment on the timeline of the song to be edited, a narration acquisition control can be added at the corresponding indexing position of each accompaniment segment, specifically at a position parallel to the timeline and corresponding to the indexing position. Essentially, it can be a button used to trigger a narration acquisition event in response to user touch operation.
[0156] Therefore, as Figure 14 As shown, each accompaniment segment has a corresponding narration capture control. By touching any of the narration capture controls, the narration capture event of the corresponding accompaniment segment can be triggered.
[0157] This embodiment displays narration acquisition controls for each accompaniment segment in a graphical user interface, integrated with the timeline of the song to be edited. Each narration acquisition control can trigger a corresponding narration acquisition event for the accompaniment segment, making it convenient for users to initiate the acquisition of narration content for each accompaniment segment. This implementation simplifies the user's operation and allows users to easily start the narration content acquisition process. It has the technical advantages of being highly efficient and intuitive, especially for mobile devices with smaller screens.
[0158] Please see Figure 15 In a further embodiment, step S1300, responding to a narration acquisition event and acquiring the narration audio data corresponding to the accompaniment segment, includes the following steps:
[0159] Step S1310: Respond to the narration acquisition event corresponding to the accompaniment segment and start recording audio data:
[0160] When the narration acquisition event corresponding to any accompaniment segment is triggered, such as Figure 16 As shown, an audio recording half-window will pop up in the graphical user interface, displaying a record button for the user to use to start or stop recording. In one embodiment, the user presses and holds the record button to start recording, and releases the button to stop recording. In another embodiment, the user touches the record button to start recording, and touches it again to stop recording. In a further embodiment, when a narration acquisition event for an accompaniment segment is triggered, the corresponding duration of automatically recorded audio data can be pre-constrained based on the playback duration of the accompaniment segment. The user only needs to click the record button to start recording, and the recording will automatically stop when the corresponding playback duration of the accompaniment segment is reached, ensuring that the recording duration of the recorded audio data does not exceed the time range defined by the start and end timestamps of the accompaniment segment. When recording ends, the corresponding recorded audio data is generated.
[0161] Step S1320: Store the recorded audio data as the corresponding narration audio data for the accompaniment segment:
[0162] After the audio data recording is completed, the background process establishes a corresponding mapping relationship between the audio data and the narration audio data of the corresponding accompaniment segment, so that it can be called in the subsequent song synthesis.
[0163] Step S1330: Associate the accompaniment segment and visualize its narration audio data as a narration indicator control:
[0164] In the graphical user interface, the narration audio data can be visually displayed for each recorded accompaniment segment. This narration audio data can be displayed using a narration indicator control. This narration indicator control can be a separate control suitable for independent manipulation, or it can share a narration acquisition control pre-programmed to trigger a narration acquisition event in response to user touch. The key is that this control changes sequentially between recording and acquiring the narration audio data to indicate that the accompaniment segment has been recorded successfully.
[0165] Once the user has completed the recording of the narration audio data for each accompaniment segment, the background process will visualize and display the narration audio data in the graphical user interface, thus completing the process of adding narration for all accompaniment segments and completing the user's narration creation.
[0166] This embodiment provides a more convenient human-computer interaction process for acquiring narration audio data, making the acquisition of narration audio data more convenient and efficient. Furthermore, by visually displaying each piece of narration audio data, users can easily perform desired operations such as shifting on the recorded narration audio data.
[0167] In a further extended embodiment, following step S1330, which involves associating the accompaniment segment with its narration audio data and visualizing it as a narration control, the following steps are included:
[0168] Step S1340: In response to a movement event acting on any narration indicator control, adjust the offset of the narration audio data of the narration indicator control relative to its accompaniment segment according to the corresponding movement amount:
[0169] Users can drag and drop any of the narration indicator controls along the timeline of the song to be edited. When the user releases the drag and drop operation, a corresponding movement event is triggered. In response to this movement event, the background process calculates the movement amount corresponding to the movement event, which is usually the movement amount corresponding to the release position of the drag and drop operation. Then, based on this movement amount, the narration indicator control is controlled to follow the movement and display. In addition, the offset of the narration audio data corresponding to the narration indicator control to the start timestamp of its accompaniment segment is also corrected in the background. That is, the start timestamp of the narration audio data relative to the playback sequence of the song to be edited is re-determined, thereby realizing the adjustment of the playback sequence of the narration audio data.
[0170] This embodiment further realizes the adjustability of the playback timing of the recorded narration audio data, allowing users to listen to the relative effect between their recorded narration audio data and the original song to be edited, and fine-tune the narration audio data as needed, thus enriching the functions achieved by the technical solution of this application.
[0171] In some extended embodiments, after step S1330, which involves associating the accompaniment segment with its narration audio data and visualizing it as a narration control, the following steps are included:
[0172] Step S1350: Respond to the song playback event and, according to the timing alignment relationship, synchronously play the audio data of the song to be edited and the narration audio data of the accompaniment segment.
[0173] Since the narration audio data establishes a temporal correspondence with the playback of the song to be edited through its corresponding accompaniment segment, the user can trigger a song playback event by touching the music playback control provided for the song to be edited. In response to this song playback event, the background program process synchronously plays the audio data of the song to be edited and each narration audio data along the same time axis according to the temporal alignment relationship between the narration audio data and the audio data of the song to be edited. Thus, although the narration audio data and the audio data of the song to be edited have not yet been synthesized, they can be played to the user in the form of two audio tracks playing synchronously. Therefore, a preview effect can be achieved. The user can decide whether to re-record the narration audio data of a certain accompaniment segment or make a fine-tuning of the temporal sequence by listening to the preview. It can be seen that this embodiment can facilitate the user to debug their narration audio data without having to process it after the song is synthesized, thus improving the efficiency of the user in creating narration content.
[0174] In a specific embodiment, after step S1310, which involves responding to the narration acquisition event corresponding to the accompaniment segment and starting to record audio data, the following steps are included:
[0175] Step S1360: Perform automatic speech recognition on the recorded audio data to obtain the narration statements corresponding to the narration audio data. The narration statements are associated with and labeled with the timestamps corresponding to the playback sequence of the song to be edited.
[0176] Once the user has recorded the narration audio for a backing track, the background process automatically calls a speech recognition interface to perform speech recognition on the recorded audio data, thereby obtaining the recognized text as the corresponding narration. This narration is suitable for storage in the lyrics data of the song to be edited and / or displayed in the graphical user interface for user correction. To facilitate the user's understanding of timing details, a timestamp corresponding to the playback sequence of the song to be edited can be added to the narration, completing the annotation of the narration. This allows the narration corresponding to the timestamp to be displayed at the appropriate playback time when the song to be edited is played.
[0177] This embodiment further enriches the assisted creation function, so that users do not need to manually enter the text content corresponding to their narration audio data, but can obtain the narration sentences corresponding to the narration audio data through automatic speech recognition, thereby making narration creation more convenient and efficient.
[0178] In some embodiments that are extended based on the previous embodiment, after step S1400, which involves responding to a narration synthesis event and synthesizing the narration audio data into the audio data of the song to be edited, the following steps are included:
[0179] Step S1500: Associate the narration sentences generated from each narration audio data through speech recognition with their corresponding timestamps and add them to the lyrics data of the song to be edited to obtain the synthesized lyrics data:
[0180] In the previous embodiment, each narration audio data is processed by speech recognition to obtain its corresponding narration statement, and then associated with its playback sequence on the timeline of the song to be edited. This results in a format identical to the lyrics data. Thus, the lyrics data can be inserted into the corresponding position of the lyrics data of the song to be edited based on its timestamp, forming part of the lyrics data, thereby obtaining the synthesized lyrics data. Subsequently, when the song to be edited is played, it can be called and displayed on the music playback interface based on the timestamp.
[0181] In this embodiment, the narration audio data is synthesized with the song to be edited, and the narration statements of the narration audio data are also synthesized with the lyrics data of the song to be edited. Thus, not only is the audio data corresponding to the synthesized song obtained through narration creation, but also the lyrics data corresponding to the song containing the narration statements is obtained, making the synthesized song more complete. No user intervention is required in the process, which can realize efficient and fast narration-assisted creation.
[0182] In a further extended embodiment, after step S1400, which involves responding to a narration synthesis event and synthesizing the narration audio data into the audio data of the song to be edited, the following steps are included:
[0183] Step S1600: Respond to the playback event after synthesis, play the synthesized audio data, and synchronously display the narration and lyrics in the synthesized lyrics data according to the timestamp:
[0184] After the user completes the synthesis of the song to be edited, the narration audio data, and the narration statements, the synthesized song is obtained. The user can then access the song's playback controls. Touching these controls triggers a playback event, initiating the playback of the synthesized audio data. During this process, according to the inherent playback logic of the music playback software, the lyrics data is also invoked while the audio data is playing. Based on the individual statements in the lyrics data, including the narration and lyrics statements, and according to the timestamps of these statements, the corresponding narration or lyrics statement is displayed in the graphical user interface before the corresponding playback time arrives. Throughout the song's playback, both the narration and lyrics statements can be displayed in the graphical user interface for easy viewing by the user.
[0185] Step S1700: In response to the song release event, publish the synthesized audio data online.
[0186] In the music playback interface of the synthesized song, a sharing control can be further provided. When the user touches the sharing control, a song publishing event can be triggered. After the user adds the corresponding text information, the synthesized song can be published to the music publishing area opened by the online music service, realizing the online sharing of the user's re-created work.
[0187] This embodiment presents the process of displaying and sharing songs containing narration content according to the present application. According to this process, it can be seen that users can achieve social sharing based on their own re-creation of the work. For individual users, this can expand their online social boundaries. For online music platforms, such user sharing activities can increase their platform user traffic and stimulate economic growth.
[0188] Please see Figure 17The aforementioned graphical user interface, along with a song editing processing apparatus provided for one of the purposes of this application, is a functionally deployed solution based on the song editing processing method. It includes: an accompaniment analysis module 1100, an accompaniment marking module 1200, a narration acquisition module 1300, and a song synthesis module 1400. The accompaniment analysis module 1100 is used to determine the accompaniment segment corresponding to the vocalless portion of the song to be edited; the accompaniment marking module 1200 is used to visually mark the accompaniment segment of the song to be edited; the narration acquisition module 1300 is used to acquire the narration audio data corresponding to the accompaniment segment in response to a narration acquisition event; and the song synthesis module 1400 is used to synthesize the narration audio data into the audio data of the song to be edited in response to a narration synthesis event.
[0189] In a further embodiment, the accompaniment analysis module 1100 includes: a first data retrieval unit, used to acquire lyrics data and audio data of the song to be edited, wherein the lyrics data includes lyrics statements and timestamps of the corresponding vocal parts in the audio data; a first candidate determination unit, used to calculate candidate accompaniment segments existing in the audio data based on the timestamps corresponding to the lyrics statements, and determine the playback duration corresponding to each candidate accompaniment segment; and a first filtering and selection unit, used to select candidate accompaniment segments with playback durations higher than a preset threshold as valid accompaniment segments.
[0190] In another embodiment of the deepening, the accompaniment analysis module 1100 includes: a second data retrieval unit for acquiring audio data of the song to be edited; a second candidate determination unit for performing human voice detection on the audio data, determining multiple candidate accompaniment segments corresponding to the non-vocal singing content, and determining the playback duration corresponding to each candidate accompaniment segment; and a second filtering and optimization unit for selecting candidate accompaniment segments with playback durations higher than a preset threshold as valid accompaniment segments.
[0191] In a further embodiment, the accompaniment marking module 1200 includes: an interface display unit for displaying a timeline corresponding to the song to be edited, the timeline being used to display the playback duration information of the song to be edited; an accompaniment indexing unit for determining the indexing position of the accompaniment segment on the timeline based on the time range of the accompaniment segment; and a control display unit for displaying a narration acquisition control corresponding to the accompaniment segment at the indexing position, for triggering a narration acquisition event in response to user touch.
[0192] In a further embodiment, the narration acquisition module 1300 includes: a recording start unit, used to respond to a narration acquisition event corresponding to the accompaniment segment and start recording audio data; an associated storage unit, used to store the recorded audio data as narration audio data corresponding to the accompaniment segment; and a visual indicator unit, used to associate the accompaniment segment and visualize its narration audio data as a narration indicator control.
[0193] In some specific embodiments, the recording start unit constrains the recording duration to not exceed the playback duration of the accompaniment segment.
[0194] In some extended embodiments, the narration acquisition module 1300 further includes an offset adjustment unit that operates after the visual indicator unit, for responding to a movement event acting on any narration indicator control, and adjusting the offset of the narration audio data of the narration indicator control relative to its accompaniment segment according to the corresponding movement amount.
[0195] In some extended embodiments, the narration acquisition module 1300 further includes a playback debugging unit that runs after the visual indication unit, used to respond to song playback events and synchronously play the audio data of the song to be edited and the narration audio data of the accompaniment segment according to the timing alignment relationship.
[0196] In some specific embodiments, the narration acquisition module 1300 further includes a speech recognition unit that operates after the visual indication unit, used to perform automatic speech recognition on the recorded audio data to obtain narration statements corresponding to the narration audio data, wherein the narration statements are associated with and marked with timestamps corresponding to the playback sequence of the song to be edited.
[0197] In some extended embodiments, the song editing processing device of this application further includes a lyrics synthesis module that runs after the song synthesis module 1400, used to associate the narration sentences generated by speech recognition of each narration audio data with their corresponding timestamps and add them to the lyrics data of the song to be edited, so as to obtain the synthesized lyrics data.
[0198] In a further extended embodiment, the song editing processing device of this application also includes, after the song synthesis module 1400, a synthesis playback module, used to respond to a post-synthesis playback event, play the synthesized audio data, and synchronously display the narration and lyrics in the synthesized lyrics data according to the timestamp; and a song publishing module, used to respond to a song publishing event, and publish the synthesized audio data online.
[0199] To address the aforementioned technical problems, embodiments of this application also provide computer equipment. For example... Figure 18The diagram shows the internal structure of a computer device. The computer device includes a processor, a computer-readable storage medium, a memory, and a network interface connected via a system bus. The computer-readable storage medium stores an operating system, a database, and computer-readable instructions. The database may store a sequence of control information. When the computer-readable instructions are executed by the processor, they enable the processor to implement a song editing processing method. The processor of the computer device provides computing and control capabilities, supporting the operation of the entire computer device. The memory of the computer device may store computer-readable instructions. When these computer-readable instructions are executed by the processor, they enable the processor to execute the song editing processing method of this application. The network interface of the computer device is used for communication with a terminal. Those skilled in the art will understand that… Figure 18 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0200] In this embodiment, the processor is used to execute... Figure 17 The specific functions of each module and its sub-modules are defined within the device. The memory stores the program code and various data required to execute these modules or sub-modules. The network interface is used for data transmission between the user terminal and the server. In this embodiment, the memory stores the program code and data required to execute all modules / sub-modules in the song editing processing device of this application. The server can call the server's program code and data to execute the functions of all sub-modules.
[0201] This application also provides a storage medium storing computer-readable instructions, which, when executed by one or more processors, cause the one or more processors to perform the steps of the song editing processing method of any embodiment of this application.
[0202] This application also provides a computer program product, including a computer program / instructions that, when executed by one or more processors, implement the steps of the song editing processing method described in any embodiment of this application.
[0203] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. This computer program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. The aforementioned storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.
[0204] In summary, this application provides a solution for users to create efficient and convenient narration content for songs to be edited on terminal devices.
[0205] Those skilled in the art will understand that the steps, measures, and solutions in the various operations, methods, and processes discussed in this application can be alternated, modified, combined, or deleted. Furthermore, other steps, measures, and solutions in the various operations, methods, and processes discussed in this application can also be alternated, modified, rearranged, decomposed, combined, or deleted. Furthermore, steps, measures, and solutions in the prior art that are similar to those disclosed in this application can also be alternated, modified, rearranged, decomposed, combined, or deleted.
[0206] The above description is only a partial embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A song editing and processing method, characterized in that, Includes the following steps: Identify the instrumental section corresponding to the vocal parts of the song to be edited; The accompanying section of the song to be edited is visually marked; In response to a narration acquisition event, the narration audio data corresponding to the accompaniment segment is acquired, including: in response to a narration acquisition event corresponding to the accompaniment segment, starting to record audio data; storing the recorded audio data as narration audio data corresponding to the accompaniment segment; and associating the accompaniment segment with its narration audio data and visually displaying it as a narration indicator control. In response to the narration synthesis event, the narration audio data is synthesized into the audio data of the song to be edited.
2. The song editing and processing method according to claim 1, characterized in that, Identify the instrumental section corresponding to the vocal parts of the song to be edited, including the following steps: Obtain the lyrics data and audio data of the song to be edited. The lyrics data includes the lyrics and the timestamps of the corresponding vocal parts in the audio data. Based on the timestamps corresponding to the lyrics, the candidate accompaniment segments in the audio data are calculated, and the playback duration corresponding to each candidate accompaniment segment is determined. Candidate accompaniment segments with a playback duration exceeding a preset threshold are selected as valid accompaniment segments.
3. The song editing and processing method according to claim 1, characterized in that, Identify the instrumental section corresponding to the vocal parts of the song to be edited, including the following steps: Obtain the audio data of the song to be edited; Human voice detection is performed on the audio data to identify multiple candidate accompaniment segments corresponding to the non-vocal singing content, and the playback duration of each candidate accompaniment segment is determined. Candidate accompaniment segments with a playback duration exceeding a preset threshold are selected as valid accompaniment segments.
4. The song editing and processing method according to claim 1, characterized in that, Visualizing the accompaniment section of the song to be edited includes the following steps: The timeline corresponding to the song to be edited is displayed, and the timeline is used to show the playback duration information of the song to be edited; The indexing position of the accompaniment segment on the time axis is determined based on the time range of the accompaniment segment; A narration capture control corresponding to the accompaniment segment is displayed at the index position to trigger a narration capture event in response to user touch.
5. The song editing method according to claim 1, characterized in that, In the step of starting audio data recording in response to the narration acquisition event corresponding to the accompaniment segment, the recording duration is constrained to not exceed the playback duration of the accompaniment segment.
6. The song editing method according to claim 1, characterized in that, After the step of associating the accompaniment segment with its narration audio data and visualizing it as a narration control, the following steps are included: In response to a movement event applied to any narration indicator control, the offset of the narration audio data of the narration indicator control relative to its accompaniment segment is adjusted according to the corresponding movement amount.
7. The song editing method according to claim 1, characterized in that, After the step of associating the accompaniment segment with its narration audio data and visualizing it as a narration control, the following steps are included: In response to a song playback event, the audio data of the song to be edited and the narration audio data of the accompaniment segment are played synchronously according to the timing alignment relationship.
8. The song editing method according to claim 1, characterized in that, After responding to the narration acquisition event corresponding to the accompaniment segment and starting the recording of audio data, the following steps are included: Automatic speech recognition is performed on the recorded audio data to obtain the narration statements corresponding to the narration audio data. The narration statements are associated with timestamps that correspond to the playback time of the song to be edited.
9. The song editing method according to claim 1, characterized in that, Following the step of synthesizing the narration audio data into the audio data of the song to be edited in response to the narration synthesis event, the following steps are included: The narration sentences generated by speech recognition of each narration audio data are associated with their corresponding timestamps and added to the lyrics data of the song to be edited to obtain the synthesized lyrics data.
10. The song editing method according to claim 9, characterized in that, Following the step of synthesizing the narration audio data into the audio data of the song to be edited in response to the narration synthesis event, the following steps are included: Responding to the playback event after synthesis, the synthesized audio data is played, and the narration and lyrics in the synthesized lyrics data are displayed synchronously according to the timestamp; In response to the song release event, the synthesized audio data is published online.
11. A song editing and processing device, characterized in that, It includes: The accompaniment analysis module is used to determine the accompaniment section corresponding to the vocal-free parts of the song to be edited; The accompaniment marking module is used to visually mark the accompaniment segments of the song to be edited; The narration acquisition module is used to respond to the narration acquisition event and acquire the narration audio data corresponding to the accompaniment segment, including: responding to the narration acquisition event corresponding to the accompaniment segment and starting to record audio data; storing the recorded audio data as the narration audio data corresponding to the accompaniment segment; associating the accompaniment segment with its narration audio data and visually displaying it as a narration indicator control; The song synthesis module is used to respond to the narration synthesis event and synthesize the narration audio data into the audio data of the song to be edited.
12. A computer device comprising a central processing unit and a memory, characterized in that, The central processing unit is used to invoke and run a computer program stored in the memory to perform the steps of the method as described in any one of claims 1 to 10.
13. A computer-readable storage medium, characterized in that, It stores, in the form of computer-readable instructions, a computer program implemented according to any one of claims 1 to 10, which, when invoked by a computer, executes the steps included in the corresponding method.
14. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method described in any one of claims 1 to 10.
Citation Information
Patent Citations
Music sharing method and device and computer readable storage medium
CN111583973A