Song editing method, device and equipment
By extracting the audio features and lyrics files of the song, and using neural network and QRC lyrics files to correct time nodes, the problem of incomplete semantic structure in song editing is solved, and an efficient song editing method is realized, and multilingual song editing is supported.
Patent Information
- Application Number
- CN202210276635.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-21
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2042-03-21
AI Technical Summary
In the prior art, the song editing method cannot ensure the integrity and coherence of the semantic structure of the edited song segments, and the manual editing efficiency is low. The automatic editing method cannot recognize the main chorus passage, resulting in incomplete editing.
By extracting the audio feature information and lyrics files of the song, using neural network and QRC lyrics files to correct time nodes, automatically align and calibrate song clips, ensuring the integrity and coherence of semantic structure.
It realizes the semantic structure integrity and coherence of song editing, improves editing efficiency, reduces labor costs, and supports multilingual song editing.
Smart Images

Figure CN114639367B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of multimedia technology, and in particular to a song editing method, device and equipment. Background Art
[0002] With the development of the multimedia industry, especially the rise of short videos, the fragmented consumer era has led to a shift in user demand for music content. This has led to demands for music ringtones, song previews, and chorus (or climax) soundtracks, prompting the development of song editing technology. The verse-chorus section is one of the most important and commonly used musical forms. Each verse-chorus section can be divided into smaller sections with similar semantic structures and repetitive A and B characters. Currently, manual editing or traditional automatic editing methods are commonly used to generate song highlights. However, due to differences in people's understanding of song structure, editing standards are inconsistent, and manual editing is very inefficient. While traditional automatic editing methods can improve efficiency, they cannot ensure the integrity of the semantic structure of song clips. Summary of the Invention
[0003] In response to the above technical problems, the present application provides a song editing method, device and equipment, which can not only retain the semantic structure integrity and coherence of the edited song clips, but also improve editing efficiency and reduce overhead.
[0004] In a first aspect, an embodiment of the present application provides a song editing method. The method can be executed by a computer device (such as a terminal or a server), and the specific method includes:
[0005] Processing the audio file of the song to be edited and extracting the audio feature information of the song to be edited;
[0006] Determining first audio information of the song to be edited based on audio feature information of the song to be edited;
[0007] Determine a first structure according to first audio information of the song to be edited and first text information of the song to be edited;
[0008] The second structure is processed for time node correction according to the specified lyrics file of the song to be edited and the first structure to obtain a third structure. The second structure includes part or all of the content of the preset song to be edited, and the third structure includes part or all of the content of the preset song to be edited after the time node correction.
[0009] This method leverages the semantic information of structured song segments to automatically align and time-correct the music clips to be edited. This not only preserves the semantic structure and coherence of the edited song clips, but also improves editing efficiency and reduces overhead.
[0010] Among them, the computer device obtains the first audio information of the song to be edited (for example, the audio feature information of the chorus of the song to be edited), and can edit the song clip more accurately based on the characteristics of the chorus providing variability to the song melody and having strong memorability.
[0011] In one possible implementation, the computer device inputs audio feature information of the song to be edited into a neural network to obtain a first audio information probability set of the song to be edited, where the audio feature information includes CQT local features and MIDI vocal melody features of the song to be edited;
[0012] According to the first audio information probability set, first audio information of the song to be edited is determined, where the first audio information includes the chorus audio information of the song to be edited.
[0013] It can be seen that since the CQT local features are consistent with the perception frequency of the human ear, the MIDI vocal melody features can more accurately hit the vocal time points at the boundary of the first audio information. Therefore, by inputting the CQT local features and vocal melody features of the song to be edited into the neural network to determine the first audio information, the accuracy of the obtained first audio information can be improved.
[0014] In one possible implementation, the computer device calculates the edit distance between any two lyrics in the song to be edited based on a specified lyrics file of the song to be edited;
[0015] According to the edit distance, first text information of the song to be edited is obtained, where the first text information includes a text similarity matrix of the song to be edited.
[0016] It can be seen that this method uses a specified lyrics file (a special lyrics file provided in this application, for example, it can be called a QRC lyrics file) to obtain the first text information, so that based on the timestamp characteristics of the QRC lyrics file, the lyrics text corresponding to each time node can be obtained more accurately, which is conducive to improving the accuracy of editing.
[0017] In one possible implementation, the computer device divides the song to be edited into sections based on the first text information of the song to be edited, and obtains first time information, where the first time information includes time nodes corresponding to different sections.
[0018] Performing fuzzy matching on the first time information and the time nodes corresponding to the first audio information of the song to be edited, determining the lyrics text information corresponding to the first audio information and the lyrics text information corresponding to the second audio information, where the second audio information includes the main song audio information of the song to be edited;
[0019] According to the word overlap of the lyrics text and the similarity of the lyrics composition structure in the song to be edited, the lyrics text information corresponding to the first audio information and the lyrics text information corresponding to the second audio information are structurally segmented to determine a first structure, which is the structured segmentation result of the song to be edited.
[0020] As can be seen, by fuzzy matching the time nodes corresponding to the first audio of the song to be edited with the time nodes corresponding to different paragraphs in the song to be edited, the computer device can assign semantic information to different paragraph structures, thereby dividing the verse and chorus information of the song to be edited. Based on the word overlap of each lyric text and the structural similarity of each lyric line in the song to be edited, the verse and chorus information can be divided into structurally symmetrical small paragraphs, thereby obtaining the semantic structure information of the entire song.
[0021] In a possible implementation, the computer device obtains a preset starting time point and a preset duration of the second structure;
[0022] Calibrate the preset start time point and end time point of the second structure according to the specified lyrics file of the song to be edited, and obtain the start time point and end time point of the second structure after the calibration time node;
[0023] According to the first structure, a time node correction process is performed on the start time point and the end time point of the second structure after the time node is calibrated to obtain a third structure.
[0024] In an embodiment of the present application, the computer device performs time point correction processing on the start time point and end time point of the calibrated second structure based on the first structure, which can improve the structural connection and listening continuity of the second structure, thereby enhancing the user's listening experience.
[0025] In one possible implementation, the computer device updates the first text information according to the specified lyrics file, where the updated first text information includes lyrics text information corresponding to the first audio information and lyrics text information corresponding to the second audio information;
[0026] According to the updated first text information and the time information corresponding to the updated first text information, the preset starting time point and ending time point of the second structure are calibrated to obtain the starting time point and ending time point of the second structure after the calibrated time node.
[0027] In an embodiment of the present application, the computer device calibrates the preset starting time point and ending time point of the second structure through the QRC lyrics file, and can obtain the precise time of the starting time point and ending time point of the second structure, thereby improving the accuracy of the editing.
[0028] In a possible implementation, the computer device obtains, from the first structure, a first time difference corresponding to a starting time point of a second structure after the calibration time node;
[0029] Obtaining, from the first structure, a second time difference corresponding to an end time point of a second structure after the calibration time node;
[0030] performing a time node correction process on the starting time point of the second structure body after the calibration time node according to the starting time point of the second structure body after the calibration time node and the first time difference;
[0031] performing a time node correction process on the end time point of the second structure body after the calibration time node according to the end time point of the second structure body after the calibration time node and the second time difference;
[0032] The third structure includes the start time point and the end time point of the second structure after the time node correction processing.
[0033] In an embodiment of the present application, since the first structure includes the semantic structure information of the entire song, the computer device performs time node correction processing on the start time point and end time point of the second structure after the calibrated time node based on the first structure, which can improve the semantic structure integrity and listening coherence of the obtained third structure.
[0034] In a possible implementation, the computer device determines the corresponding duration according to the start time point and the end time point of the second structure after the calibration time node;
[0035] According to the duration, different paragraphs in the first structure are connected end to end to obtain a third structure.
[0036] In an embodiment of the present application, the computer device can connect the different paragraphs in the first structure end to end according to the editing duration, and can improve the connection and coherence of the song while freely splicing the paragraphs.
[0037] In a second aspect, an embodiment of the present application provides a song editing device, the device comprising:
[0038] A pre-processing module is used to process the audio file of the song to be edited and extract the audio feature information of the song to be edited;
[0039] A determination module, configured to determine first audio information of the song to be edited based on audio feature information of the song to be edited;
[0040] The determining module is further configured to determine a first structure according to the first audio information of the song to be edited and the first text information of the song to be edited;
[0041] The processing module is used to perform time node correction processing on the second structure according to the QRC lyrics file of the song to be edited and the first structure to obtain a third structure, wherein the second structure includes part or all of the content of the preset song to be edited, and the third structure includes part or all of the content of the preset song to be edited after the time node correction.
[0042] In a third aspect, an embodiment of the present application further provides a computer device, comprising: a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, any of the above methods is implemented.
[0043] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, any of the above methods is implemented.
[0044] In a fifth aspect, embodiments of the present application further provide a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the method provided in embodiments of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0046] Figure 1 This is a flowchart of a song editing method provided by an embodiment of the present application;
[0047] Figure 2 Schematic diagram of a CQT local feature provided in an embodiment of the present application;
[0048] Figure 3 is a schematic diagram of a MIDI vocal melody feature provided by an embodiment of the present application;
[0049] Figure 4 A short-term climax probability curve diagram provided in an embodiment of the present application;
[0050] Figure 5 This is a schematic diagram of filtering a short-term climax probability curve provided by an embodiment of the present application;
[0051] Figure 61 is a flow chart of another song editing method provided in an embodiment of the present application;
[0052] Figure 7 Schematic diagram of a lyric text similarity matrix shown in an embodiment of the present application;
[0053] Figure 8 is a schematic diagram of a first structure provided in an embodiment of the present application;
[0054] Figure 9 is a schematic diagram of a continuous editing provided by an embodiment of the present application;
[0055] Figure 10 is a schematic diagram of a splicing and editing method provided in an embodiment of the present application;
[0056] Figure 11 This is a framework diagram of a song editing method provided in an embodiment of the present application;
[0057] Figure 12 is a schematic diagram of a song editing device provided in an embodiment of the present application;
[0058] Figure 13 This is a schematic diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0059] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0060] To facilitate understanding of the embodiments disclosed in this application, some concepts involved in the embodiments of this application are first explained. The explanation of these concepts includes but is not limited to the following.
[0061] 1. QRC lyrics file
[0062] A lyrics file in Extensible Markup Language (XML) format that can precisely control the timing of each word in the lyrics.
[0063] 2. CQT local features
[0064] The Constant-Q Transform (CQT) local feature is a nonlinear frequency domain feature obtained by filtering the time domain audio signal with a set of constant-Q filters. This feature is more consistent with music theory.
[0065] 3. MIDI vocal melody characteristics
[0066] The vocal melody characteristics of the digital music format (Musical Instrument Digital Interface, MIDI) are a set of features that describe the pitch, rhythm, and strength of the human voice, representing the ups and downs of the melody.
[0067] 4. Main Song
[0068] The verse is the backbone of a song. Its function is to slowly push the melody to a climax while clearly expressing the background story of the song, and it has a strong narrative quality.
[0069] 5. Chorus
[0070] A refrain (also known as a chorus or climax) is a repeated verse or line in a song, typically appearing between verses. The refrain contrasts with the verses in length, melody, rhythm, and emotion, adding variety to the song's melody and making it highly memorable.
[0071] 6. Edit distance
[0072] Edit distance is a quantitative measure of the difference between two strings. It refers to the minimum number of edit operations required to transform one string into the other. Generally speaking, the smaller the edit distance, the greater the similarity between the two strings.
[0073] 7. Structured segmentation
[0074] Structured segmentation, which contains rich semantic information, is one of the primary forms of musical expression. From a song content perspective, similar, repetitive lyrics are often grouped together or grouped into a single paragraph. The structure of popular music can generally be divided into alternating verse and chorus sections.
[0075] Currently, song editing primarily relies on manual and automatic editing. Manual editing suffers from inconsistent editing standards due to discrepancies in understanding song structure, and manual editing is also inefficient. Existing automatic editing methods, on the other hand, primarily rely on duration or simple audio signal processing. These methods fail to identify the main chorus sections of a song, resulting in an abrupt and incomplete ending to the edited song, and fail to ensure the semantic structural integrity of the edited song segment.
[0076] Based on this, embodiments of the present application provide a song editing method, apparatus, and device. This method utilizes the semantic information of a song's structured segments to automatically align and time-correct the music clips to be edited. This method not only preserves the semantic structure and coherence of the edited song clips, but also improves editing efficiency and reduces overhead.
[0077] It should be noted that: in a specific implementation, the above scheme can be executed by a computer device, which can be a terminal or a server; the terminals mentioned here can include but are not limited to: smart phones, tablet computers, laptops, desktop computers, smart watches, smart TVs, smart car terminals, etc.; the server mentioned here can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDN), as well as big data and artificial intelligence platforms and other basic cloud computing services, etc., which are not limited here.
[0078] In order to facilitate understanding of the embodiments of the present application, the specific implementation methods of the song editing method will be described in detail below by taking a computer device executing the song editing method as an example.
[0079] Figure 1 This is a flow chart of a song editing method provided by an embodiment of the present application. Figure 1 The method can be executed by a computer device, and includes the following steps S101-S102.
[0080] S101: Process the audio file of the song to be edited and extract audio feature information of the song to be edited.
[0081] The song to be edited refers to a user-specified song (including an audio file and a lyrics file). For example, the song to be edited may include multiple song segments (e.g., a verse segment and a chorus segment), along with the corresponding lyrics for each segment. Optionally, the lyrics file may also include non-lyric information (e.g., the title, artist, production, mix, etc.).
[0082] Optionally, the file format of the audio file of the song to be edited includes but is not limited to MP3, MP4, Wave Audio File Format (WAV), etc.
[0083] In an optional embodiment, the computer device extracts audio feature information of the song to be edited, including extracting CQT local features and MIDI vocal melody features of the song to be edited. Figure 2 This is a schematic diagram of a CQT local feature provided by an embodiment of the present application. Figure 2 As shown in FIG, the CQT local feature is a nonlinear time-frequency spectrum, which is transformed based on log base 2 (ie, log2). Figure 2 In the figure, the horizontal axis represents time and the vertical axis represents frequency. When the pitch is distributed on a logarithmic span with log2 as the base, it is consistent with the human ear's perception frequency (the higher the frequency, the lower the sensitivity). Figure 3 is a schematic diagram of a MIDI vocal melody feature provided by an embodiment of the present application, Figure 3 In the figure, the horizontal axis represents time, and the vertical axis represents the height (or pitch) of the vocals. This feature has a stronger representation of the vocals and can more accurately hit the vocal timing points at the boundaries of the chorus segment.
[0084] S102: Determine first audio information of the song to be edited based on audio feature information of the song to be edited.
[0085] The first audio information is a chorus segment of the song to be edited. For example, if the song to be edited includes multiple chorus segments, the first audio information may include audio information of each of the multiple chorus segments.
[0086] In an optional embodiment, the computer device determines the first audio information of the song to be edited, which may include the following steps: inputting the audio feature information of the song to be edited into a neural network to obtain a first audio information probability set of the song to be edited; and determining the first audio information of the song to be edited based on the first audio information probability set, wherein the first audio information includes the chorus audio information of the song to be edited.
[0087] Specifically, for example, the probability set includes {P1, P2, ..., P n}, n represents the nth time point, and each probability represents the possibility that the time point is the climax part. The computer device can obtain the first audio information probability curve based on the first audio information probability set. For example, Figure 4 A climax probability curve diagram provided in an embodiment of the present application is shown. The climax probability curve diagram is the first audio information probability curve described above, wherein the horizontal axis represents the index of the time point and the vertical axis represents the probability.
[0088] The computer device can determine the first audio information of the song to be edited based on the first audio information probability curve. Figure 4 , determine the climax segment (i.e. chorus segment) of the song to be edited. Figure 4As shown in the figure, the probability between point a and point b is the highest, so we can preliminarily assume that the section between point a and point b is the climax of the song to be edited. Similarly, based on the probability value, we can also assume that the section between point c and point d is the climax of the song to be edited.
[0089] In the embodiment of determining the first audio information of the song to be edited, the computer device may further: input the audio feature information of the song to be edited into the neural network to obtain a probability curve of the first audio information of the song to be edited; and determine the first audio information of the song to be edited based on the first audio information probability curve, where the first audio information includes the audio information of the chorus of the song to be edited. For example, the computer device may be based on Figure 4 , determine the climax segment (i.e. chorus segment) of the song to be edited.
[0090] It can be understood that, in this embodiment, the probability curve of the first audio information can be directly obtained without being based on a probability set.
[0091] In an implementation method of determining the first audio information of a song to be edited based on a first audio information probability curve, a computer device may obtain a first audio filter curve by filtering the first audio information probability curve; and determine the first audio information of the song to be edited based on the first audio filter curve.
[0092] In an embodiment of determining the first audio information of a song to be edited based on a first audio filter curve, the computer device may determine the time information and confidence of the first audio information of the song to be edited based on the maximum point in the filter curve and the minimum point after the maximum point; the time information of the first audio includes a starting time point and an ending time point, and the confidence of the first audio information includes a starting confidence level and an ending confidence level; based on the starting confidence level and the ending confidence level of the first audio information, the average confidence level of the first audio information in the song to be edited is calculated, and the first audio information of the song to be edited is determined. The horizontal and vertical coordinates of the maximum point in the filter curve are the starting time point and the starting confidence level of the first audio information, respectively, and the horizontal and vertical coordinates of the minimum point after the maximum point are the ending time point and the ending confidence level of the first audio information, respectively.
[0093] In this embodiment, the computer device calculates the average confidence of the first audio information in the song to be edited, and determines the first audio information of the song to be edited based on the average confidence, thereby ensuring the reliability of the first audio information.
[0094] Alternatively, the computer device may determine the first audio information by setting a threshold value [a, b] of the average confidence. For example, if the average confidence of the first audio information is c, c∈[a, b], then the first audio information with the average confidence c may be used as the first audio information of the song to be edited. For another example, if the average confidence of the first audio information is d, Then the first audio information with the average confidence level d cannot be used as the first audio information of the song to be edited.
[0095] Figure 5 This is a schematic diagram of filtering a short-term climax probability curve provided by an embodiment of the present application. Figure 5 As shown, the computer device uses the Harr filter to analyze the short-term climax probability curve v p Perform filtering to obtain the filtering curve v f ; According to the filtering curve v f The first audio information of the song to be edited can be obtained. First, the computer device can input the CQT local features and MIDI life melody features of the song to be edited into the neural network, and predict the short-term climax probability every 500ms as a frame to obtain the short-term climax probability set of the song to be edited; according to the short-term climax probability set, the climax probability curve v is obtained. p The computer device uses a Haar filter with a width of 20s to slide through the entire climax probability curve v p , get the filtering curve v f . Determine the filter curve v f The maximum point (i.e. Figure 5 The time of point A in the figure is the climax starting time point t c_start , t c_start The corresponding value is the initial confidence s c_start , and determine the minimum point after the maximum point (i.e. Figure 5 The time of point B in the figure is the climax end time t c_end , t c_end The corresponding value is the end confidence s c_end Using the initial confidence s c_start and end confidence s c_end , calculate the average confidence of the climax segment
[0096] S103: Determine a first structure according to the first audio information of the song to be edited and the first text information of the song to be edited.
[0097] The first text information is the lyrics text corresponding to the song to be edited. For example, the first text information includes the lyrics text of the verse segment and the lyrics text of the chorus segment in the song to be edited. Optionally, the first text information may also include a similarity matrix of the lyrics text of the song to be edited.
[0098] The first structure is the structured segmentation result of the song to be edited, including the semantic structure information of the song to be edited. For example, the first structure includes the verse segment, chorus segment, and character A and B segments of the song to be edited.
[0099] S104 , performing time node correction processing on the second structure according to the designated lyrics file of the song to be edited and the first structure to obtain a third structure.
[0100] The second structure includes part or all of the content of the preset song to be edited, and the second structure can also be called a segment to be edited. For example, the second structure can be one or more segments in the preset song to be edited. The second structure includes a preset starting point and a preset duration. According to the preset starting time point and the preset duration, the end time point of the second structure can be obtained. For example, the preset starting time point can be 00:30.21 in the song to be edited, the preset duration can be 20s, and the end time point is 00:50.21.
[0101] The third structure includes part or all of the content of the preset song to be edited after the time node is corrected, that is, the second structure after the time node is corrected, which can also be called the edited audio. The third structure can also include the starting time point and the ending time point of the second structure after the time node is corrected.
[0102] In an optional embodiment, the computer device performs time node correction processing on the second structure according to the specified lyrics file of the song to be edited (for example, the specified lyrics file is a QRC lyrics file) and the first structure to obtain a third structure, including: obtaining a preset starting time point and a preset duration of the second structure; calibrating the preset starting time point and ending time point of the second structure according to the specified lyrics file of the song to be edited to obtain the starting time point and ending time point of the second structure after the calibration time node; performing time node correction processing on the starting time point and ending time point of the second structure after the calibration time node according to the first structure to obtain the third structure.
[0103] By using the semantic information of the song's structured segments, the embodiments of this application automatically align and time-correct the music clips to be edited. This not only preserves the semantic structure and coherence of the edited song clips, but also improves editing efficiency and reduces overhead. Furthermore, because the lyrics text information in the QRC lyrics file supports multiple languages, the embodiments of this application can support song editing in multiple languages and minority languages.
[0104] Figure 6 This is a flow chart of another song editing method provided by an embodiment of the present application. Figure 6As shown, the method described in the embodiment of the present application includes steps S601a-S604. It should be noted that step a (such as S601a and S602a) in the process of this method represents the operation performed on the audio information of the song to be edited, and step b (such as S601b and S602b) represents the operation performed on the second structure.
[0105] S601a: Process the audio file of the song to be edited and extract audio feature information of the song to be edited.
[0106] S602a: Determine first audio information of the song to be edited based on audio feature information of the song to be edited.
[0107] The specific processes of steps S601a and S402a can be found in the descriptions of S101 and S102 above, and will not be repeated here.
[0108] S603: Determine a first structure according to the first audio information of the song to be edited and the first text information of the song to be edited.
[0109] In an optional embodiment, when the computer device determines the first structure based on the first audio information of the song to be edited and the first text information of the song to be edited, it may include a preprocessing part, a multimodal fusion part and a post-processing part.
[0110] First, the preprocessing portion includes obtaining first text information of the song to be edited based on the QRC lyrics file of the song to be edited. Therefore, before determining the first structure based on the first audio information and the first text information of the song to be edited, the computer device also includes: obtaining the QRC lyrics file of the song to be edited; calculating the edit distance between any two lyrics in the song to be edited based on the QRC lyrics file; and obtaining the first text information of the song to be edited based on the edit distance, wherein the first text information includes a text similarity matrix of the song to be edited.
[0111] Optionally, the computer device may limit the file format of the QRC lyrics file to the qrc format.
[0112] Figure 7 is a schematic diagram of a lyrics text similarity matrix shown in an embodiment of the present application. Figure 7 In the graph, both the horizontal and vertical axes represent the index of the lyrics. For example, (10,35) represents the similarity between the 10th and 35th lines of lyrics.
[0113] Optionally, the preprocessing part further includes: dividing the song to be edited into sections according to the first text information of the song to be edited, and obtaining first time information, where the first time information includes time nodes corresponding to different sections;
[0114] Secondly, the multimodal fusion process involves fuzzy matching the first time information with the time nodes corresponding to the first audio information of the song to be edited, and determining the lyric text information corresponding to the first audio information and the lyric text information corresponding to the second audio information, where the second audio information includes the verse audio information of the song to be edited. Multimodality refers to multiple sources, media, or forms of information, such as text, audio, and images. Multimodal fusion involves fusing multiple types of information.
[0115] In the computer device's fuzzy matching of the first time information and the time nodes corresponding to the first audio information of the song to be edited, the first time information includes the time nodes corresponding to different sections of the song to be edited, including the starting and ending time points of the verse segment, the starting and ending time points of the chorus segment, and the like. The time nodes corresponding to the first audio information include the starting and ending time points of the chorus segment. Matching the time nodes corresponding to different sections of the song to be edited with the starting and ending time points of the chorus segment is called fuzzy matching.
[0116] Finally, the post-processing part includes: structurally segmenting the lyrics text information corresponding to the first audio information and the lyrics text information corresponding to the second audio information based on the word overlap of each sentence of lyrics and the structural similarity of each sentence of lyrics in the song to be edited, and determining the first structure, which is the structured segmentation result of the song to be edited.
[0117] The specific implementation process of the above-mentioned preprocessing part, multimodal fusion part, and post-processing part is illustrated below. Assuming that the first audio information is a chorus segment of the song to be edited, when determining the first structure, the computer device can obtain the QRC lyrics file of the song to be edited, and based on the QRC lyrics file, calculate the edit distance between any two lyrics in the song to be edited to obtain the text similarity matrix of the song to be edited; based on the text similarity matrix, the optimal path search algorithm is used to combine the lyrics segments with similarity greater than a preset threshold to obtain the song to be edited after segmentation (i.e., the verse segment and the chorus segment), as well as the time nodes corresponding to the different segments. Fuzzy matching is performed on the time nodes corresponding to the different segments in the song to be edited and the start and end time points of the chorus segment determined in step S402a to determine the lyrics text information corresponding to the chorus segment and the lyrics text information corresponding to the verse segment. Based on the word overlap of each lyric text in the song to be edited and the structural similarity of each lyric composition, the lyrics text information corresponding to the chorus segment and the lyrics text information corresponding to the verse segment are divided into structurally symmetrical small segments AB, thereby determining the first structure of the song to be edited. The first structure includes the semantic structure information of the song to be edited, as well as the main song section, chorus section, and character AB section. The first structure (or the semantic structure of the song to be edited) can be recorded as
[0118] Sec=[V1,A1,B1,C1,A2,B2,C2,…,V n ,C n ,A n ,B n ], n∈[1,N], where V represents the main song section, C represents the chorus section, A represents the small section of character A, B represents the small section of character B, and N represents the total number of sentences in the lyrics of the song to be edited. Figure 8 is a schematic diagram of a first structure provided in an embodiment of the present application, such as Figure 8 As shown, the song to be edited includes two main song sections V1 and V2, two chorus sections C1 and C2, four role A sections A1, A2, A3 and A4, and three role B sections B1, B2 and B4.
[0119] S601b: Obtain a preset starting time point and a preset duration of the second structure.
[0120] For example, the computer device may obtain the preset starting time point t of the second structure from the user side. u_start and preset duration dur .
[0121] S602b: Calibrate the preset start time point and end time point of the second structure according to the designated lyrics file of the song to be edited, and obtain the start time point and end time point of the second structure after the calibration time node.
[0122] In an optional embodiment, the computer device calibrates the preset starting time point and ending time point of the second structure according to the specified lyrics file of the song to be edited, and obtains the starting time point and ending time point of the second structure after the calibrated time node, including: updating the first text information according to the specified lyrics file, the updated first text information includes the lyrics text information corresponding to the first audio information and the lyrics text information corresponding to the second audio information; calibrating the preset starting time point and ending time point of the second structure according to the updated first text information and the time information corresponding to the updated first text information, and obtaining the starting time point and ending time point of the second structure after the calibrated time node.
[0123] Optionally, assuming that the designated lyrics file is a QRC lyrics file, in an implementation method of updating the first text information according to the designated lyrics file, the computer device removes the prelude part of the song to be edited and the song information such as the title, singer, production, and remix in the lyrics text that does not contain lyrics based on the QRC lyrics file and the filtering module.
[0124] In an implementation method in which the preset starting time point and ending time point of the second structure are calibrated according to the updated first text information and the time information corresponding to the updated first text information to obtain the starting time point and ending time point of the second structure after the calibrated time node: the preset starting time point of the second structure is processed using the updated first text information to obtain the first starting time point of the second structure, and the first ending time point of the second structure is obtained based on the first starting time point and the preset duration; the first starting time point and the first ending time point are calibrated using the time information corresponding to the updated first text information to obtain the second starting time point and the second ending time point, the second starting time point and the second ending time point being the starting time point and the ending time point of the second structure after the calibrated time node.
[0125] For example, assuming that the preset starting time point of the second structure obtained by the computer device from the user side is t u_start , the preset duration is l dur The computer device can use the lyrics text content of the QRC lyrics file and the filtering module to remove the prelude of the song to be edited and the song information such as the title, singer, production, and mixing in the lyrics text, and obtain the updated lyrics text information. u_startProcessing is performed to determine the first starting time point t' of the second structure u_start ; For the preset duration l dur and t′ u_start Perform addition operation to obtain the estimated end time point t' of the second structure u_end Then, use the precise time information of the QRC lyrics file to align the first starting time point t′ u_start and end time t′ u_end , and get t″ u_start and t″ u_end . t″ u_start and t″ u_end That is, the starting time point and the ending time point of the second structure after the calibration time node.
[0126] Optionally, the computer device may also directly calibrate the preset start time point and end time point of the second structure based on the updated first text information and the time information corresponding to the updated first text information to obtain the start time point and end time point of the second structure after the calibration time node.
[0127] For example, assuming that the preset starting time point of the second structure obtained by the computer device from the user side is T u_start , the preset duration is L dur The computer device can first remove the prelude of the song to be edited and the song information such as the title, singer, production, and mixing in the lyrics text according to the lyrics text content of the QRC lyrics file and the filtering module to obtain the updated lyrics text information; then, according to the updated lyrics text information and the time information corresponding to the updated lyrics text information, the preset starting time point T of the second structure is set. u_start Process and determine the starting time point T' of the second structure u_start ; Finally, for the preset duration L dur and T′ u_start Perform addition operation to determine the estimated end time point T' of the second structure u_end Among them, T′ u_start and T′ u_end That is, the starting time point and the ending time point of the second structure after the calibration time node.
[0128] S604 : Based on the first structure, perform time node correction processing on the start time point and the end time point of the second structure after the time node is calibrated to obtain a third structure.
[0129] In an optional embodiment, the computer device performs time node correction processing on the start time point and end time point of the second structure after the calibration time node based on the first structure to obtain a third structure, including: obtaining the first time difference corresponding to the start time point of the second structure after the calibration time node from the first structure; obtaining the second time difference corresponding to the end time point of the second structure after the calibration time node from the first structure. According to the start time point of the second structure after the calibration time node and the first time difference, the start time point of the second structure after the calibration time node is subjected to time node correction processing; according to the end time point of the second structure after the calibration time node and the second time difference, the end time point of the second structure after the calibration time node is subjected to time node correction processing. The third structure includes the start time point and end time point of the second structure after the time node correction processing. It can be understood that the method of obtaining the third structure in this embodiment can also be called a continuous editing method, that is, an editing method that maintains the continuity from the start to the end of the segment.
[0130] Among them, the first time difference includes the difference between the boundary times of the two paragraphs adjacent to the paragraph where the start time point of the second structure after the calibration time node is located; the second time difference includes the difference between the boundary times of the two paragraphs adjacent to the paragraph where the end time point of the second structure after the calibration time node is located.
[0131] Figure 9 This is a schematic diagram of a continuous clip provided by an embodiment of the present application. Figure 9 As shown, assuming that the first structure is Sec described in step S403, the starting time point and the ending time point of the second structure are t″ u_start and t″ u_end The computer device performs t″ according to the first structure. u_start and t″ u_end When the time node correction processing is performed and the third structure is obtained, the time nodes corresponding to t″ can be obtained from the first structure Sec. u_start and t″ u_end The difference between the boundary time of the two adjacent paragraphs is Δup and Δdown. The boundary point of the paragraph with the minimum difference is taken as the corrected start time. The calculation formula is as follows: Formula (1). In Formula (1), t″′ u_star It indicates the starting time point of the second structure after the time node is corrected. For example, Figure 9 As shown, t″ u_start The value is 00:51.68, which is consistent with t″ u_start The boundary time of the previous paragraph adjacent to the current paragraph is 00:42.30, which is the same as t″ u_startThe boundary time of the next paragraph adjacent to the current paragraph is 01:02.59, so Δup and Δdown can be calculated: Δup = 00:51.68-00:42.30 = 00:09.38, Δdown = 01:02.59-00:51.68 = 00:10.91. Since Δup < Δdown, the boundary time point of the paragraph where Δup is located is used to compare t″ u_start After correction, according to formula (1), t″′ u_star 00:51.68-00:09.38=00:42.30. Similarly, we can get the end time t″ u_end The corresponding corrected end time point t″′ u_end Finally, the actual duration of the second structure after correction based on the first structure is l″′ in the following formula (2): dur , wherein the actual duration of the second structure is approximately equal to the preset duration.
[0132] t″′ u_start =t″ u_start ±min(Δup,Δdown) (1)
[0133] l″′ dur =t″′ u_end -t″′ u_start ≈l dur (2)
[0134] In another optional embodiment, the computer device performs time node correction processing on the start time point and end time point of the second structure after the calibration time node based on the first structure to obtain a third structure, including: determining the target duration between the start time point and the end time point of the second structure after the calibration time node; and according to the target duration, connecting the different paragraphs in the first structure end to end to obtain the third structure. It is understandable that the method of obtaining the third structure in this embodiment can also be called a splicing editing method, that is, an editing method that freely splices together different main and chorus paragraphs or character AB paragraphs according to user needs and under the condition that the editing segment duration is met.
[0135] Figure 10 Schematic diagram of splicing and editing provided by the embodiment of the present application. Figure 10As shown, assuming the first structure is Sec described in step S403, then, subject to the clip duration requirement, the computer device can freely splice together the different chorus sections or character A and B sections in the first structure. That is, the V1, C1, and C2 sections in Sec are spliced end to end to obtain a third structure. At this point, the computer device can perform a fade-in and fade-out of the sound of As at the splicing point to further reduce the abruptness of the splicing, where A can be 0.5 seconds. This editing method is called splicing editing.
[0136] Optionally, after obtaining the third structure, the computer device may use an audio trimming tool to cut the third structure from the audio file of the song to be edited based on the start time point and end time point of the third structure (i.e., the start time point and end time point of the second structure after the time nodes are corrected), and output the third structure. The output third structure is the edited song segment.
[0137] By adopting the implementation method of the present application, the semantic information of the structured segmentation of the song is utilized to automatically align, calibrate and time-correct the music clips to be edited. This not only preserves the semantic structural integrity and auditory coherence of the edited song clips, but also improves editing efficiency and reduces overhead.
[0138] Furthermore, since the embodiments of this application support editing of songs at any specified starting point and duration (any continuous time region), and also support editing by freely splicing semantic segments based on duration, this application can automatically and flexibly edit music clips with complete structure and coherent sound according to different scene requirements. For example, this application can be applied to short video soundtracks, music games, ringtones, choruses, and other scenes, without limitation.
[0139] Figure 11This is a framework diagram of a song editing method provided in an embodiment of the present application, corresponding to steps S601a to S604 above. The computer device first extracts the CQT audio local features and MIDI vocal melody features of the song to be edited, and then inputs the CQT audio local features and MIDI vocal melody features into a neural network to determine the chorus segment. Secondly, the text similarity matrix of the song to be edited is calculated based on the QRC lyrics file, and the song to be edited is divided into paragraphs based on the lyrics similarity; combined with the determined chorus segment, a structured segmentation result of the song to be edited is obtained using multimodal fusion technology. Then, based on the preset start time and preset duration of the segment to be edited, the preset start time point and end time point of the segment to be edited are aligned and calibrated based on the QRC lyrics file to obtain the start time point and end time point of the segment to be edited after the aligned and calibrated time node. The structured segmentation result of the song to be edited is then used to correct the start time point and end time point of the segment to be edited after the aligned and calibrated time point, so that it is adaptively corrected to the boundary of the nearest neighboring paragraph of the paragraph where the start time point and end time point are located respectively. Finally, based on the corrected time point of the clip to be edited, the song clip is cut by the song editing tool to obtain the edited song clip. It can be seen that the embodiment of the present application uses the semantic information of the song structured segmentation and the QRC lyrics file to automatically align and correct the clip to be edited, which not only preserves the semantic structure integrity and coherence of the clip
[0140] In addition, compared with manual editing methods, this method saves time and effort, can improve editing efficiency, and thus reduce editing costs; compared with existing automatic editing methods, it can improve the success rate and coverage of song editing.
[0141] Figure 12 Schematic diagram of a song editing device provided in an embodiment of the present application. The song editing device described in this embodiment may include the following parts:
[0142] The pre-processing module 1201 is used to process the audio file of the song to be edited and extract the audio feature information of the song to be edited;
[0143] The determining module 1202 is configured to determine first audio information of the song to be edited based on audio feature information of the song to be edited;
[0144] The determining module 1202 is further configured to determine a first structure according to the first audio information of the song to be edited and the first text information of the song to be edited;
[0145] Processing module 1203 is used to perform time node correction processing on the second structure based on the specified lyrics file of the song to be edited and the first structure to obtain a third structure, wherein the second structure includes part or all of the content of the preset song to be edited, and the third structure includes part or all of the content of the preset song to be edited after the time node correction.
[0146] In an optional implementation, when the determining module 1202 is used to determine the first audio information of the song to be edited based on the audio feature information of the song to be edited, it is specifically used to:
[0147] Inputting audio feature information of the song to be edited into a neural network to obtain a first audio information probability set of the song to be edited, the audio feature information including CQT local features and MIDI vocal melody features of the song to be edited;
[0148] According to the first audio information probability set, first audio information of the song to be edited is determined, where the first audio information includes the chorus audio information of the song to be edited.
[0149] In an optional embodiment, the processing module 1203 is further configured to calculate the edit distance between any two lyrics in the song to be edited based on a specified lyrics file of the song to be edited;
[0150] According to the edit distance, first text information of the song to be edited is obtained, where the first text information includes a text similarity matrix of the song to be edited.
[0151] In an optional implementation, when the determining module 1202 is used to determine the first structure according to the first audio information of the song to be edited and the first text information of the song to be edited, it is specifically used to:
[0152] Divide the song to be edited into sections according to the first text information of the song to be edited, and obtain first time information, where the first time information includes time nodes corresponding to different sections;
[0153] Performing fuzzy matching on the first time information and the time nodes corresponding to the first audio information of the song to be edited, determining the lyrics text information corresponding to the first audio information and the lyrics text information corresponding to the second audio information, where the second audio information includes the main song audio information of the song to be edited;
[0154] According to the word overlap of each lyric text in the song to be edited and the structural similarity of each lyric text, the lyric text information corresponding to the first audio information and the lyric text information corresponding to the second audio information are structured segmented to determine a first structure, which is the structured segmentation result of the song to be edited.
[0155] In an optional embodiment, when the processing module 1203 is used to perform time node correction processing on the second structure according to the specified lyrics file of the song to be edited and the first structure to obtain the third structure, it is specifically used to:
[0156] Obtaining a preset start time point and a preset duration of the second structure; calibrating the preset start time point and end time point of the second structure according to the specified lyrics file of the song to be edited, and obtaining the start time point and end time point of the second structure after the calibration time node;
[0157] According to the first structure, a time node correction process is performed on the start time point and the end time point of the second structure after the time node is calibrated to obtain a third structure.
[0158] In an optional embodiment, the processing module 1203 is used to calibrate the preset start time point and end time point of the second structure according to the specified lyrics file of the song to be edited, and obtain the start time point and end time point of the second structure after the calibration time node, specifically for:
[0159] updating the first text information according to the specified lyrics file, wherein the updated first text information includes lyrics text information corresponding to the first audio information and lyrics text information corresponding to the second audio information;
[0160] According to the updated first text information and the time information corresponding to the updated first text information, the preset starting time point and ending time point of the second structure are calibrated to obtain the starting time point and ending time point of the second structure after the calibrated time node.
[0161] In an optional embodiment, when the processing module 1203 is used to perform time node correction processing on the start time point and the end time point of the second structure after the calibration time node according to the first structure to obtain the third structure, it is specifically used to:
[0162] Obtaining, from the first structure, a first time difference corresponding to a starting time point of the second structure after the calibration time node;
[0163] Obtaining, from the first structure, a second time difference corresponding to an end time point of the second structure after the calibration time node;
[0164] Performing time node correction processing on the starting time point of the second structure body after the calibration time node according to the starting time point of the second structure body after the calibration time node and the first time difference;
[0165] Performing a time node correction process on the end time point of the second structure body after the calibration time node according to the end time point of the second structure body after the calibration time node and the second time difference;
[0166] The third structure includes the start time point and the end time point of the second structure after the time node correction processing.
[0167] In an optional embodiment, when the processing module 1203 is used to perform time node correction processing on the start time point and the end time point of the second structure after the calibration time node according to the first structure to obtain the third structure, it is specifically used to:
[0168] Determine the target duration between the start time point and the end time point of the second structure after the calibration time node;
[0169] According to the target duration, different paragraphs in the first structure are connected end to end to obtain a third structure.
[0170] It can be understood that the specific implementation of each module in the song editing device described in the embodiment of the present application and the beneficial effects that can be achieved can be referred to the description of the aforementioned related embodiments, and will not be repeated here.
[0171] Figure 13 1304 is a schematic diagram of the structure of a computer device according to an embodiment of the present application. The computer device described in the embodiment of the present application includes: a processor 1301, a user interface 1302, a communication interface 1303, and a memory 1304. The processor 1301, user interface 1302, communication interface 1303, and memory 1304 may be connected via a bus or other means. The embodiment of the present application uses a bus connection as an example.
[0172] The processor 1301 (also known as the Central Processing Unit (CPU)) is the computing and control core of the computer device. It can interpret various instructions within the computer device and process various data from the computer device. For example, the CPU can interpret power on / off commands sent by the user to the computer device and control the computer device to perform power on / off operations. Another example is that the CPU can transmit various interactive data between the internal components of the computer device, etc. The user interface 1302 is the medium for user interaction and information exchange between the computer device and the user. Its specific embodiment may include a display for output and a keyboard for input, etc. It should be noted that the keyboard here can be a physical keyboard, a touchscreen virtual keyboard, or a combination of physical and touchscreen virtual keyboards. The communication interface 1303 can optionally include a standard wired interface or a wireless interface (such as Wi-Fi, mobile communication interface, etc.), and is controlled by the processor 1301 to send and receive data. The memory 1304 (Memory) is a storage device in the computer device used to store programs and data. It is understood that the memory 1304 herein may include both the built-in memory of the computer device and, of course, the extended memory supported by the computer device. The memory 1304 provides storage space for storing the operating system of the computer device, which may include but is not limited to: Android system, iOS system, Windows Phone system, etc., and this application does not limit this.
[0173] In the embodiment of the present application, the processor 1301 performs the following operations by running the executable program code in the memory 1304:
[0174] Processing the audio file of the song to be edited and extracting the audio feature information of the song to be edited;
[0175] Determining first audio information of the song to be edited based on audio feature information of the song to be edited;
[0176] Determine a first structure according to first audio information of the song to be edited and first text information of the song to be edited;
[0177] The second structure is processed for time node correction according to the specified lyrics file of the song to be edited and the first structure to obtain a third structure. The second structure includes part or all of the content of the preset song to be edited, and the third structure includes part or all of the content of the preset song to be edited after the time node correction.
[0178] In an optional implementation, when the processor 1301 is configured to determine the first audio information of the song to be edited based on the audio feature information of the song to be edited, it is specifically configured to:
[0179] Inputting audio feature information of the song to be edited into a neural network to obtain a first audio information probability set of the song to be edited, the audio feature information including CQT local features and MIDI vocal melody features of the song to be edited;
[0180] According to the first audio information probability set, first audio information of the song to be edited is determined, where the first audio information includes the chorus audio information of the song to be edited.
[0181] In an optional embodiment, the processor 1301 is further configured to calculate the edit distance between any two lyrics in the song to be edited based on a specified lyrics file of the song to be edited;
[0182] According to the edit distance, first text information of the song to be edited is obtained, where the first text information includes a text similarity matrix of the song to be edited.
[0183] In an optional implementation manner, when the processor 1301 is used to determine the first structure according to the first audio information of the song to be edited and the first text information of the song to be edited, it is specifically used to:
[0184] Divide the song to be edited into sections according to the first text information of the song to be edited, and obtain first time information, where the first time information includes time nodes corresponding to different sections;
[0185] Performing fuzzy matching on the first time information and the time nodes corresponding to the first audio information of the song to be edited, determining the lyrics text information corresponding to the first audio information and the lyrics text information corresponding to the second audio information, where the second audio information includes the main song audio information of the song to be edited;
[0186] According to the word overlap of each lyric text in the song to be edited and the structural similarity of each lyric text, the lyric text information corresponding to the first audio information and the lyric text information corresponding to the second audio information are structured segmented to determine a first structure, which is the structured segmentation result of the song to be edited.
[0187] In an optional embodiment, when the processor 1301 is used to perform time node correction processing on the second structure according to the specified lyrics file of the song to be edited and the first structure to obtain the third structure, it is specifically used to:
[0188] Obtaining a preset start time point and a preset duration of the second structure; calibrating the preset start time point and end time point of the second structure according to the specified lyrics file of the song to be edited, and obtaining the start time point and end time point of the second structure after the calibration time node;
[0189] According to the first structure, a time node correction process is performed on the start time point and the end time point of the second structure after the time node is calibrated to obtain a third structure.
[0190] In an optional embodiment, when the processor 1301 is used to calibrate the preset start time point and end time point of the second structure according to the specified lyrics file of the song to be edited, and obtain the start time point and end time point of the second structure after the calibration time node, it is specifically used to:
[0191] updating the first text information according to the specified lyrics file, wherein the updated first text information includes lyrics text information corresponding to the first audio information and lyrics text information corresponding to the second audio information;
[0192] According to the updated first text information and the time information corresponding to the updated first text information, the preset starting time point and ending time point of the second structure are calibrated to obtain the starting time point and ending time point of the second structure after the calibrated time node.
[0193] In an optional embodiment, when the processor 1301 is configured to perform time node correction processing on the start time point and the end time point of the second structure after the calibration time node according to the first structure to obtain the third structure, it is specifically configured to:
[0194] Obtaining, from the first structure, a first time difference corresponding to a starting time point of the second structure after the calibration time node;
[0195] Obtaining, from the first structure, a second time difference corresponding to an end time point of the second structure after the calibration time node;
[0196] Performing time node correction processing on the starting time point of the second structure body after the calibration time node according to the starting time point of the second structure body after the calibration time node and the first time difference;
[0197] Performing a time node correction process on the end time point of the second structure body after the calibration time node according to the end time point of the second structure body after the calibration time node and the second time difference;
[0198] The third structure includes the start time point and the end time point of the second structure after the time node correction processing.
[0199] In an optional embodiment, when the processor 1301 is configured to perform time node correction processing on the start time point and the end time point of the second structure after the calibration time node according to the first structure to obtain the third structure, it is specifically configured to:
[0200] Determine the target duration between the start time point and the end time point of the second structure after the calibration time node;
[0201] According to the target duration, different paragraphs in the first structure are connected end to end to obtain a third structure.
[0202] In a specific implementation, the processor 1301, user interface 1302, communication interface 1303 and memory 1304 described in the embodiments of the present application can execute the implementation method of the computer device described in the song editing method provided in the embodiments of the present application, and can also execute the implementation method described in the song editing device provided in the embodiments of the present application, which will not be repeated here.
[0203] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. The computer program includes program instructions. When the program instructions are executed by a processor, the song editing method provided by the embodiment of the present application is implemented. For details, please refer to the implementation methods provided in the above steps, which will not be repeated here.
[0204] The present application also provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the method described in the present application. The specific implementation method is described above and will not be repeated here.
[0205] It should be noted that for the aforementioned various method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.
[0206] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0207] The above disclosure is only part of the embodiments of the present application, and certainly cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application are still within the scope covered by the present application.
Claims
1. A song editing method, characterized in that: The method comprises: Processing the audio file of the song to be edited and extracting audio feature information of the song to be edited; Determining first audio information of the song to be edited based on the audio feature information of the song to be edited, where the first audio information includes audio information of the chorus of the song to be edited; Obtaining first text information of the song to be edited, the first text information including lyrics text information corresponding to the verse audio information of the song to be edited, lyrics text information corresponding to the chorus audio information, and a text similarity matrix of the song to be edited, the text similarity matrix being used to indicate the similarity between any two lyrics in the song to be edited; According to the first text information of the song to be edited, based on the text similarity matrix, the song to be edited is divided into paragraphs to obtain first time information, where the first time information includes time nodes corresponding to different paragraphs; Performing fuzzy matching on the first time information and the time nodes corresponding to the first audio information of the song to be edited, determining lyrics text information corresponding to the first audio information and lyrics text information corresponding to the second audio information, wherein the second audio information includes the audio information of the main song of the song to be edited; Based on the word overlap and structural similarity of the lyrics in the song to be edited, the lyrics text information corresponding to the first audio information and the lyrics text information corresponding to the second audio information are structurally segmented to determine a first structure, wherein the first structure is a result of the structured segmentation of the song to be edited; obtaining a preset starting time point and a preset duration of a second structure, wherein the second structure includes part or all of the content of the preset song to be edited, wherein the part or all of the content of the preset song to be edited is a preset segment to be edited; Calibrate the preset start time point and end time point of the second structure according to the specified lyrics file of the song to be edited, and obtain the start time point and end time point of the second structure after the calibration time node; According to the first structure, the start time point and the end time point of the second structure after the calibration time node are subjected to time node correction processing to obtain a third structure, wherein the third structure includes part or all of the content of the preset song to be edited after the time node correction, and the start time point and the end time point of the third structure are located at the boundary of the nearest adjacent paragraph of the paragraph where the start time point and the end time point of the second structure after the calibration time node are respectively located.
2. The method according to claim 1, characterized in that The determining, based on the audio feature information of the song to be edited, the first audio information of the song to be edited comprises: Inputting audio feature information of the song to be edited into a neural network to obtain a first audio information probability set of the song to be edited, the audio feature information including CQT local features and MIDI vocal melody features of the song to be edited; The first audio information of the song to be edited is determined based on the first audio information probability set.
3. The method according to claim 1, characterized in that The method further comprises: Determining, based on a specified lyrics file of a song to be edited, an edit distance between any two lyrics in the song to be edited; According to the edit distance, first text information of the song to be edited is obtained.
4. The method according to claim 1, wherein The step of calibrating the preset start time point and end time point of the second structure according to the designated lyrics file of the song to be edited, and obtaining the start time point and end time point of the second structure after the calibration time node, includes: updating the first text information according to the designated lyrics file, wherein the updated first text information includes lyrics text information corresponding to the first audio information and lyrics text information corresponding to the second audio information; According to the updated first text information and the time information corresponding to the updated first text information, the preset starting time point and ending time point of the second structure are calibrated to obtain the starting time point and ending time point of the second structure after the calibrated time node.
5. The method according to claim 1 or 4, characterized in that According to the first structure, a time node correction process is performed on the start time point and the end time point of the second structure after the calibration time node to obtain a third structure, including: Obtaining, from the first structure, a first time difference corresponding to a starting time point of a second structure after the calibration time node, wherein the first time difference comprises a difference between boundary times of two adjacent paragraphs to which the starting time point of the second structure after the calibration time node belongs; Obtaining, from the first structure, a second time difference corresponding to an end time point of a second structure after the calibration time node, the second time difference comprising a difference between boundary times of two adjacent paragraphs to which the end time point of the second structure after the calibration time node belongs; performing a time node correction process on the starting time point of the second structure body after the calibration time node according to the starting time point of the second structure body after the calibration time node and the first time difference; performing a time node correction process on the end time point of the second structure body after the calibration time node according to the end time point of the second structure body after the calibration time node and the second time difference; The third structure includes the start time point and the end time point of the second structure after the time node correction processing.
6. The method according to claim 1 or 4, characterized in that According to the first structure, a time node correction process is performed on the start time point and the end time point of the second structure after the calibration time node to obtain a third structure, including: Determine a target duration between a start time point and an end time point of a second structure after the calibration time node; According to the target duration, different paragraphs in the first structure are connected end to end to obtain a third structure.
7. A computer device, characterized in that: It includes a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the steps of the song editing method as described in any one of claims 1 to 6 are implemented.
8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the song editing method according to any one of claims 1 to 6.
9. A computer program product comprising computer instructions, characterized in that The computer instructions are stored in a computer-readable storage medium, and when read and executed by a processor of a computer device, the computer device executes the steps of the song editing method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Music structure determination method and device, equipment and medium
CN112037764A