A song processing method, device, equipment, medium and program product
Patent Information
- Application Number
- CN202610866016.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-15
- Publication Date
- 2026-09-01
AI Technical Summary
[0003]在此提供一种歌曲处理方法、装置、设备、介质以及程序产品,可以解决歌曲拆分不合理的问题
[0009]上述方法通过在第一页面展示第一控件,响应于针对第一控件的触发操作,获得第一控件对应的歌曲数据;响应于针对歌曲数据的确认操作,获得歌曲数据对应的歌曲片段集合,歌曲片段集合用于生成音乐视频,其中,歌曲片段集合是基于音乐结构、歌词信息、节奏信息和时长规则中至少一项对应的切点,切分歌曲数据得到的。通过上述方法实现结合音乐结构、歌词信息、节奏信息和时长规则,生成一组考虑歌曲表达完整性、镜头完整性及整体视听节奏的切点,实现基于切点切分歌曲数据得到的歌曲片段用于生成效果符合预期的音乐视频,解决了歌曲拆分方式不合理而影响视频呈现效果的问题,提升了歌曲拆分的合理性,提升了音乐视频中镜头切换的平滑性,增强了镜头变化与音乐节奏的适配程度,优化了视频呈现效果。
Smart Images

Figure CN122672879A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer processing technology, and in particular to a song processing method, apparatus, device, medium, and program product. Background Technology
[0002] With the development of computer technology, content generation tools are becoming increasingly diverse in their capabilities. For example, users can upload songs to content generation tools to create music videos (MVs). However, current methods of splitting songs into their constituent parts are often flawed, which negatively impacts the final presentation of the MV. Summary of the Invention
[0003] This invention provides a song processing method, apparatus, device, medium, and program product that can solve the problem of unreasonable song splitting.
[0004] In one scenario, this paper provides a song processing method, which includes: Display the first control on the first page; In response to a trigger operation on the first control, obtain the song data corresponding to the first control; In response to the confirmation operation on the song data, a set of song segments corresponding to the song data is obtained. The set of song segments is used to generate a music video. The set of song segments is obtained by segmenting the song data based on a cut point corresponding to at least one of the following: music structure, lyrics information, rhythm information, and duration rules.
[0005] In one instance, this document also provides a song processing apparatus, which includes: The display module is used to display the first control on the first page; The first obtaining module is used to obtain the song data corresponding to the first control in response to a trigger operation on the first control; The second obtaining module is used to obtain a set of song segments corresponding to the song data in response to a confirmation operation on the song data. The set of song segments is used to generate a music video. The set of song segments is obtained by segmenting the song data based on a cut point corresponding to at least one of the following: music structure, lyrics information, rhythm information, and duration rules.
[0006] In one instance, this document also provides an electronic device comprising: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the song processing method as described herein.
[0007] In one instance, this document also provides a storage medium containing computer-executable instructions that, when executed by a computer processor, are used to perform song processing methods as described herein.
[0008] In another scenario, this document also provides a computer program product, including a computer program that, when executed by a processor, implements the song processing method as described herein.
[0009] The above method displays a first control on the first page. In response to a trigger operation on the first control, it obtains the song data corresponding to the first control. In response to a confirmation operation on the song data, it obtains a set of song fragments corresponding to the song data. This set of song fragments is used to generate a music video. The song fragment set is obtained by segmenting the song data based on cut points corresponding to at least one of the following: music structure, lyrics, rhythm, and duration rules. This method combines music structure, lyrics, rhythm, and duration rules to generate a set of cut points that consider the completeness of song expression, shot integrity, and overall audiovisual rhythm. This allows the song fragments obtained by segmenting the song data based on these cut points to generate a music video with the expected effect. It solves the problem of unreasonable song segmentation affecting video presentation, improves the rationality of song segmentation, enhances the smoothness of shot transitions in the music video, strengthens the adaptation of shot changes to music rhythm, and optimizes the video presentation effect. Attached Figure Description
[0010] The above and other features, advantages, and aspects of the embodiments described herein will become more apparent when taken in conjunction with the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0011] Figure 1 This is a schematic diagram of a song processing system under one specific scenario. Figure 2 This is a flowchart illustrating a song processing method under one specific scenario. Figure 3 This is a schematic diagram of the first page in one scenario; Figure 4 This is a flowchart illustrating a song processing method for another scenario. Figure 5 This is a flowchart illustrating the song data segmentation process in one scenario of song processing. Figure 6This is a flowchart illustrating another method for song processing. Figure 7 This is a flowchart illustrating the duration rule verification method in a song processing scenario. Figure 8 This is a schematic diagram of a song processing device in one scenario. Figure 9 This is a schematic diagram of an electronic device used to implement a song processing method in one scenario. Detailed Implementation
[0012] The embodiments will now be described in more detail with reference to the accompanying drawings. While some embodiments are shown in the drawings, it should be understood that the technical solutions can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the technical solutions herein. It should be understood that the illustrated drawings and embodiments are for illustrative purposes only and are not intended to limit the scope of protection of the technical solutions.
[0013] It should be understood that the steps described in the method implementation may be performed in different orders and / or in parallel. Furthermore, the method implementation may include additional steps and / or omit the steps shown. The scope of this document is not limited in this respect.
[0014] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one situation" means "at least one situation"; the term "another situation" means "at least one additional situation"; the term "some situations" means "at least some situations". Definitions of other terms will be given in the following description.
[0015] It should be noted that the concepts of "first" and "second" mentioned are only used to distinguish different devices, modules or units, and are not used to limit the order of the functions performed by these devices, modules or units or their interdependencies.
[0016] It should be noted that the terms "one" and "more" used in this document are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0017] The names of messages or information exchanged between multiple devices in this document are for illustrative purposes only and are not intended to limit the scope of these messages or information.
[0018] It is understood that before using the technical solutions disclosed in the various embodiments of this document, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this document in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0019] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as electronic devices, applications, servers, or storage media, that perform the operations described herein, based on the prompt message.
[0020] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0021] It is understood that the above notification and user authorization process is merely illustrative and does not limit the implementation method described in this article. Other methods that comply with relevant laws and regulations may also be applied to the implementation method described in this article.
[0022] It is understood that the data involved in the technical solutions in this article (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0023] In some cases, the provided solution can be applied to Figure 1 The illustrated song processing system may include a client 101 and a server 102. The client 101 may include, but is not limited to, web applications such as browsers, applications (Apps), HyperText Markup Language (HTML) applications, lightweight applications (also known as mini-programs), or cloud applications. The client 101 may be deployed on an electronic device and relies on the operation of that device or certain applications within the device to implement its functions. The electronic device may be, for example, a device with a display screen that supports information browsing, such as a smartphone, tablet, personal computer, or other client terminal. For ease of understanding, Figure 1The client is primarily represented by a device. Other applications, such as content generation applications, can also be configured on the electronic device. Server 102 can be one or more servers providing various services. That is, it can be implemented as a distributed server cluster composed of multiple servers, or as a single server; furthermore, it can be a server for a distributed system, a server integrating blockchain technology, a cloud server, or an intelligent cloud computing server or intelligent cloud host deployed with machine learning models, etc.
[0024] The song processing method described in this paper allows interaction between client 101 and server 102, such as receiving or sending messages. For example, in this paper, server 102 can receive song fragments sent by client 101 based on an information carrier, analyze the song fragments, obtain music structure, lyrics information, and rhythm information, and segment the song data to obtain a set of song fragments based on the cut points corresponding to at least one of the music structure, lyrics information, rhythm information, and duration rules. The set of song fragments is then sent to client 101 for display on the display interface.
[0025] It should be noted that the song processing method can be executed by client 101, or by client 101 and server 102, with different functional parts of the corresponding song processing device deployed on client 101 and server 102 respectively; wherein, the front-end interaction module and local data acquisition module of the device are deployed on client 101, and the back-end data processing module and data storage module are deployed on server 102, and client 101 and server 102 achieve data interaction and functional collaboration through network communication. It should be understood that... Figure 1 The number of clients and servers shown is for illustrative purposes only. Any number of clients and servers can be configured to meet specific implementation requirements.
[0026] Figure 2 This is a flowchart illustrating a song processing method for one scenario. This method is applicable to song processing scenarios, especially music creation scenarios, such as generating music videos. This song processing method can be executed by a song processing device, which can be implemented in software and / or hardware, optionally through an electronic device, such as a mobile terminal, PC, or server. Figure 2 As shown, the song processing method may specifically include: S210. Display the first control on the first page.
[0027] The first page represents the interactive page of a content generation application. This application can be a music creation application, or similar. Optionally, a music creation application can be a tool that generates and publishes songs based on user-input prompts. These prompts include at least one of the following: song creation requirements, style prompts, timbre prompts, segmentation instructions, and genre prompts.
[0028] The first control is used to select the song data to be processed within the entire song. Song data refers to the processing range targeted by this processing task within the song. The first control includes at least one of a playback progress bar and tabs. The playback progress bar includes a selection window that corresponds to a time segment of the song, and the tabs represent at least a portion of the song data. Specifically, the position and length of the selection window correspond to the time segment of the song. For example, sliding the selection window to select song data from second 30 to second 55. Alternatively, selecting the tab corresponding to second 35 to second 75 of the song indicates that the content from second 35 to second 75 of the song will be used as song data.
[0029] Figure 3 This is a diagram illustrating the first page in one scenario. For example... Figure 3 As shown, a first page 310 is displayed, and a playback progress bar 320 is displayed on the first page 310. A selection window 330 is displayed at the position corresponding to the playback progress bar 320. Optionally, moving the selection window 330 can switch the selected song data. Dragging the boundary of the selection window 330 can modify the time segment corresponding to the selected song data. Optionally, multiple label items 340 are displayed below the playback progress bar 320. The multiple label items 340 are used to represent the entire song or the highlight segments of the song, etc.
[0030] S220. In response to a trigger operation on the first control, obtain the song data corresponding to the first control.
[0031] In one scenario, in response to an interactive operation on the selection window, the song time period corresponding to the selection window is adjusted; and song data corresponding to the song time period is obtained.
[0032] The selection window can be configured as a rectangle. Optionally, the playback progress bar can be configured as a waveform playback progress bar. A waveform playback progress bar is a user interface component that combines an audio waveform with a progress bar. See also Figure 3 As shown, in response to the interactive operation of the selection window 330, the song time period covered by the selection window 330 is adjusted. Optionally, time information 350 is displayed at the corresponding position of the selection window 330, wherein the time information represents the song time period covered by the selection window.
[0033] In another scenario, in response to a trigger operation on the tag item, song data corresponding to the triggered tag item is obtained, wherein the tag item is used to represent the song or a highlight segment in the song.
[0034] In this context, the highlight segment in a song refers to the most impactful combination of melody and / or lyrics. Optionally, a multimodal model can be used to process the song's Mel spectrogram, lyrics, and audio data to extract audio features in the time and frequency domains, semantic features of the lyrics, and music theory features (including beat and musical structure). A multi-channel feature matrix is constructed based on the audio, semantic, and music theory features. An encoder is used to encode the audio, semantic, and music theory features in the multi-channel feature matrix, and the encoding results are input into a cross-modal attention layer. Cross-modal attention is calculated through the cross-modal attention layer, outputting modal enhancement features that enhance each modal feature with information from other modalities. The modal enhancement features are fused into a decision vector through gating fusion. This vector is input into a lightweight classification head to output the highlight probability for each frame of the song. In this paper, multimodal features are updated through a cross-modal attention layer. For example, for a melody of moderate energy, if the attention detects that the lyrics at the same moment have extremely high intensity, the audio features will shift towards the highlight direction. If attention is focused on a repeated lyric and a simultaneous shift in chord from dominant to tonic, the confidence level of the lyric's key phrase is strengthened. Similarly, a prediction of paragraph boundaries increases the probability of a valid paragraph start if attention is focused on a simultaneous improvement in audio quality.
[0035] S230. In response to the confirmation operation on the song data, a set of song segments corresponding to the song data is obtained. The set of song segments is used to generate a music video. The set of song segments is obtained by segmenting the song data based on a cut point corresponding to at least one of the following: music structure, lyrics information, rhythm information, and duration rules.
[0036] The song fragment set is a collection of song fragments corresponding to the song data. Song fragments are obtained by splitting the song data based on cut points. Cut points are gradually determined and corrected based on the musical structure, lyrics, and rhythm information of the song data, and long or short song fragments are determined based on duration rules. Long fragments are further split based on musical structure, lyrics, and rhythm information. Short fragments are merged based on their duration. Optionally, musical structure is correlated with song segments. For example, musical structure includes verses, pre-chorus, chorus, interludes, and bridges. Duration rules are duration-related splitting strategies. Duration rules are related to the generation duration range of a single video generated by the generative model. For example, duration rules include minimum segment duration (also called lower duration limit), maximum segment duration (also called upper duration limit), recommended segment duration intervals, recommended single segment duration, and maximum duration of vocal segments. The minimum segment duration refers to the shortest duration of a song fragment; if a song fragment is too short, it usually lacks sufficient information and easily disrupts the overall rhythm of the shot. The maximum duration of a single segment refers to the longest possible length of a song fragment. This maximum duration corresponds to the generation duration range of the generative model, ensuring that individual music video fragments fall within a relatively stable generation duration range. The ideal segmentation duration range indicates the priority to find suitable cut points within this range. The recommended single segment duration prioritizes the use of this default duration for average segmentation when no corresponding cut points exist for musical structure, lyrics, or rhythm information. The maximum duration of a vocal segment prioritizes the use of this maximum duration for segmentation when the song data includes vocal performance content, thereby improving the expressive fit.
[0037] In one scenario, a confirmation control is displayed on the first page. Responding to a trigger on this control leads to a confirmation action on the song data. In another scenario, a voice command triggers the confirmation action. Alternatively, a gesture can be used to trigger the confirmation action.
[0038] In one scenario, in response to a confirmation operation on the song data, a model with music analysis capabilities is invoked. Based on the song data, the music structure, lyric boundaries, and rhythm information are determined. Cutting points are determined hierarchically based on the priority order of the music structure, lyric boundaries, and rhythm information, and the song data is segmented based on these cutting points to obtain a set of song segments. The lyric boundaries identify the start and end positions of each lyric line. If the cutting points corresponding to the music structure, lyric boundaries, and rhythm information do not satisfy the duration rule, the song data is split based on the cutting points corresponding to the duration rule to obtain a set of song segments. Optionally, if no cutting points corresponding to the music structure, lyric boundaries, and rhythm information exist, the song data is split based on the cutting points corresponding to the duration rule to obtain a set of song segments.
[0039] In another scenario, in response to the song data not containing lyrics, cut points are determined based on musical structure and rhythm information, and the song data is segmented based on these cut points to obtain a set of song segments. If the cut points corresponding to the musical structure and rhythm information cannot satisfy the duration rules, the song data is split based on the cut points corresponding to the duration rules to obtain a set of song segments. Optionally, if no cut points corresponding to the musical structure and rhythm information exist, the song data is split based on the cut points corresponding to the duration rules to obtain a set of song segments.
[0040] Determining cut points using the above methods avoids abruptly severing lyrics or musical expression, improving the integrity of the expression; it also makes shot transitions more closely match changes in musical structure, enhancing the naturalness of the rhythm; and it controls the duration of each segment within a range more suitable for generative models, improving generation stability. This paper addresses the continuity and stability issues associated with determining cut points based on fixed durations. Dividing song data into segments of fixed duration can easily lead to incomplete expression by directly splitting a single lyric, and can also forcibly cut during transitions between verses, choruses, or interludes, disrupting the continuity of musical expression. Furthermore, while segments divided by fixed durations may appear uniform in length, their content density is uneven, failing to truly meet the requirements for generation stability. This paper also addresses the integrity issues associated with determining cut points based on lyrics. Segmenting song data according to lyrics can easily lead to fragmented visuals. Furthermore, the duration of different sentences in the lyrics varies greatly, which can result in segments that are too short or too long. In addition, interludes or bridges in the music are not entirely defined by the lyrics. Simply following the lyrics cannot cover the complete musical structure, and segmenting song data according to the lyrics makes it difficult to take into account the changes in beat and the overall audiovisual rhythm.
[0041] This paper presents a first control on the first page. In response to a trigger operation on the first control, it obtains the song data corresponding to the first control. In response to a confirmation operation on the song data, it obtains a set of song fragments corresponding to the song data. This set of song fragments is used to generate a music video. The song fragment set is obtained by segmenting the song data based on cut points corresponding to at least one of the following: music structure, lyrics, rhythm, and duration rules. This method combines music structure, lyrics, rhythm, and duration rules to generate a set of cut points that consider the completeness of song expression, shot integrity, and overall audiovisual rhythm. This allows the song fragments obtained by segmenting the song data based on these cut points to generate a music video with the expected effect. This solves the problem of unreasonable song segmentation affecting video presentation, improves the rationality of song segmentation, enhances the smoothness of shot transitions in the music video, strengthens the adaptation of shot changes to music rhythm, and optimizes the video presentation effect.
[0042] Figure 4This is a flowchart illustrating a song processing method under another scenario. The technical solution in this scenario can be combined with implementation methods in other scenarios. For identical or related parts, descriptions of other scenarios can be used, and will not be repeated here. Figure 4 As shown, the method in this case may specifically include: S410, Display the first control on the first page.
[0043] S420. In response to a trigger operation on the first control, obtain the song data corresponding to the first control.
[0044] S430. In response to the confirmation operation on the song data, based on the priority corresponding to the music structure, lyrics boundary and rhythm information, obtain the cut point corresponding to the music structure, lyrics boundary and rhythm information, wherein the lyrics boundary includes the start position and end position of the entire lyric.
[0045] Here, priority refers to the order in which musical structure, lyric boundaries, and rhythmic information are used to determine cut points. Optionally, the priority is positively correlated with the order in which cut points are determined. For example, the higher the priority, the earlier the cut points are determined; the lower the priority, the later the cut points are determined. Cut points can be gradually determined and corrected based on the priority of musical structure, lyric boundaries, and rhythmic information to form a set of song segments suitable for subsequent music video generation.
[0046] In one scenario, a first type of cut point is obtained based on the musical structure; a second type of cut point is obtained by correcting the first type of cut point located inside the entire lyric based on the lyric boundary; and a third type of cut point is determined based on the rhythm information of the first song segment, wherein the first song segment is a song segment whose duration is greater than the upper limit of duration corresponding to the first type of cut point or the second type of cut point.
[0047] Musical structure is associated with the sections of a song, and the division of sections usually corresponds to changes in the song's energy expression. For example, in a song, the verse serves to build energy, the chorus to release energy, and the interlude serves as a transition between the verse and chorus. Using musical structure as the basis for identifying cut points is more consistent with the song's expressive changes. Optionally, priority should be given to finding cut points near the boundaries of the musical structure (the beginning and end of sections). These boundaries usually better align with the user's intuitive understanding of the musical content and make it easier to create natural transitions. Lyric boundaries determine the completeness of the content expression. If the first type of cut point falls within a line of lyrics, it is necessary to refine the cut point based on the beginning and end of that line to obtain a second type of cut point, in order to reduce the probability of semantic truncation.
[0048] In one scenario, based on the first type of cut points corrected from the lyrics boundary, multiple second type of cut points are obtained, including: based on the deviation between the first type of cut points within the entire lyric line and the lyrics boundary, the first type of cut points are corrected to the start or end position of the lyric line to obtain the second type of cut points.
[0049] The deviation can be the time difference between the first type of cut point and the start or end position of the lyrics. Alternatively, the deviation can be the word count difference between the position of the first type of cut point and the start or end position of the lyrics. Optionally, if the deviation between the first type of cut point within the entire lyric line and the start position of the lyrics is less than the deviation between the first type of cut point and the end position of the lyrics, then the first type of cut point is corrected to the start position, and the corrected cut point is used as the second type of cut point. If the deviation between the first type of cut point within the entire lyric line and the end position of the lyrics is less than the deviation between the first type of cut point and the start position of the lyrics, then the first type of cut point is corrected to the end position, and the corrected cut point is used as the second type of cut point. Through the above method, the first type of cut point located within the entire lyric line is corrected to the lyric boundary closest to it, thus ensuring the integrity of the lyrics after segmentation.
[0050] In another scenario, if the duration of the song segment corresponding to either the first or second type of cut point exceeds the maximum duration of a single segment, rhythmic information is used to correct the first or second type of cut point to further break down the long segment. Cut points are determined based on rhythmic information to ensure that changes in camera movement align with the perceived rhythm.
[0051] Optionally, if a first type of cut point is not obtained based on the music structure, a fourth type of cut point is determined based on the lyrics boundary; a fifth type of cut point is determined based on the rhythm information of the second song segment, wherein the second song segment is a song segment whose duration is greater than the upper limit of duration corresponding to the fourth type of cut point.
[0052] If there are no clear musical boundaries in the song data, the musical structure of the song data cannot be identified, and therefore the first type of cut point cannot be determined. In this case, the fourth type of cut point is determined based on the lyric boundaries included in the song data. If the duration of the song segment corresponding to the fourth type of cut point exceeds the maximum duration of a single segment, the fourth type of cut point is corrected based on rhythm information, that is, the fifth type of cut point within the song segment is determined based on the beat points.
[0053] Optionally, in response to the absence or failure to meet the requirements of the cut points corresponding to the music structure, lyric boundaries, and rhythm information, a sixth type of cut point is obtained based on the duration rule.
[0054] In one scenario, not meeting the requirements means not meeting the duration rules. For example, based on the cut points corresponding to the musical structure, the lyrics boundaries, and the rhythm information, the duration of the song segments obtained by splitting the song data does not conform to the duration rules. Obtaining the sixth type of cut point based on the duration rules can include determining the cut point based on the recommended single-segment duration in the duration rules to evenly distribute the song data.
[0055] In another scenario, for song segments segmented based on cut points corresponding to at least one of the following: musical structure, lyric boundaries, and rhythmic information, if at least one song segment's duration exceeds the maximum single-segment duration, then the existence of a musical structure boundary is prioritized. If such a boundary exists, the song segment is segmented based on that boundary. If no musical structure boundary exists, the existence of a lyric boundary is determined. If such a boundary exists, the song segment is segmented based on that boundary. If no lyric boundary exists, the existence of a cut point is determined based on the song segment's rhythm. If such a cut point exists, the song segment is segmented based on that rhythm. If no rhythmic cut point exists, the song segment is segmented based on the recommended single-segment duration in the duration rules.
[0056] S440. Based on the cut point, the song data is segmented to obtain the first type of song fragments.
[0057] S450. The first type of song fragments whose duration does not meet the duration rules are split or merged to obtain the second type of song fragments.
[0058] In one scenario, based on the duration and duration rules of the first type of song segments, excessively long segments and excessively short segments are identified. Excessively long segments are split, and excessively short segments are merged.
[0059] S460. Based on the first type of song fragments and the second type of song fragments whose durations conform to the duration rules, the set of song fragments is obtained.
[0060] Figure 5 This is a flowchart illustrating the song data segmentation process in one scenario of song processing. For example... Figure 5 As shown, the segmentation process includes: S510, Determine song data.
[0061] S520. Determine if the song data has a music structure boundary. If not, execute S530; otherwise, execute S540.
[0062] S530: Segment the song data based on lyric phrases.
[0063] If the song data lacks musical structure boundaries, the cutoff point is determined based on the lyric boundaries. For example, the start position of each line of lyrics can be used as the cutoff point. Alternatively, the end position of each line of lyrics can be used as the cutoff point. Another option is to determine the cutoff point based on the semantic coherence of the lyrics in the song data. For example, the end position of the second line in two semantically consecutive lines of lyrics can be used as the cutoff point.
[0064] S540: Segment the song data based on the boundaries of the music structure.
[0065] If the song data has musical structure boundaries, these boundaries are prioritized as cut points. For example, cut points are determined based on the song's verse, chorus, interlude, and bridging sections.
[0066] S550. Determine whether the cut point is located inside the entire lyric. If yes, execute S560; otherwise, execute S570.
[0067] S560, Correct the cutting point based on the lyrics boundary of the cut lyrics.
[0068] S570, Preserve the current cutoff point.
[0069] S580. Determine whether the song segment exceeds the maximum duration of a single segment. If so, execute S590; otherwise, execute S5100.
[0070] S590. Based on rhythm information, further segment the song fragments that exceed the maximum duration of a single segment.
[0071] For song segments segmented based on musical structure boundaries or lyric boundaries, if their duration exceeds the maximum duration of a single segment, the song segment is further segmented based on rhythm information. This segmentation may result in a song segment duration shorter than the maximum duration of a single segment, or the segmented song segment may still be longer than the maximum duration of a single segment. Therefore, duration verification is required using duration rules to ensure the stability of the model generation.
[0072] S5100 performs duration verification based on duration rules.
[0073] S5110: Determine whether the cut-off point does not conform to the duration rule. If it does, execute S5120; otherwise, execute S5130.
[0074] If the cut point corresponding to the music structure, the cut point corresponding to the lyrics boundary, and the cut point corresponding to the rhythm boundary all fail to meet the requirements of the duration rule, or if there is no cut point corresponding to the music structure, the cut point corresponding to the lyrics boundary, and the cut point corresponding to the rhythm boundary, then execute S5120.
[0075] S5120, segmenting song data based on duration rules.
[0076] S5130, Output a collection of song fragments.
[0077] This paper first determines the first type of cut point based on the musical structure. For the first type of cut point located inside the entire lyric, it is then corrected based on the lyric boundary to obtain the second type of cut point. If there is a song segment whose duration exceeds the maximum duration of a single segment, the song segment is then segmented based on rhythm information. This achieves the step-by-step determination of cut points based on musical structure, lyric boundary, and rhythm information, thereby reducing the abruptness of camera cuts and avoiding the situation where the entire lyric is cut off, thus improving the consistency between camera changes and musical rhythm.
[0078] Figure 6 This is a flowchart illustrating another song processing method. The technical solution in this scenario can be combined with implementation methods in other scenarios. For identical or related parts, descriptions of other scenarios can be used, and will not be repeated here. Figure 6 As shown, the method in this case may specifically include: S610: Display the first control on the first page.
[0079] S620. In response to a trigger operation on the first control, obtain the song data corresponding to the first control.
[0080] S630. In response to the confirmation operation on the song data, based on the priority corresponding to the music structure, lyric boundaries and rhythm information, obtain the cut point corresponding to the music structure, lyric boundaries and rhythm information, wherein the lyric boundaries include the start position and end position of the entire lyric.
[0081] S640. Based on the cut point, the song data is segmented to obtain the first type of song fragments.
[0082] S650. For the first type of song segment whose duration exceeds the upper limit, the first type of song segment is split based on the cutting points corresponding to the music structure, lyric boundaries, or rhythm information to obtain the second type of song segment.
[0083] In one scenario, the duration of the first type of song segment exceeds the maximum duration of a single segment. The first type of song segment is divided based on the musical structure, lyric boundaries, and rhythmic information. The specific implementation method has been described in the example above and will not be repeated here.
[0084] S660. For the first type of song segment whose duration is less than the minimum duration limit, merge the first type of song segment with the adjacent song segment.
[0085] In one scenario, if the duration of a song segment of type I is less than the minimum duration of a single segment, the song segment of type I will be merged with the adjacent song segments.
[0086] S670. In response to the merging duration being less than or equal to the upper limit of duration, the merging result of the first type of song segment and the adjacent song segment is taken as the second type of song segment.
[0087] The merged duration refers to the duration of the segment after merging the first type of song segment with adjacent song segments. If the merged duration is less than or equal to the maximum duration of a single segment, the merged segment will be classified as a second type of song segment.
[0088] S680, In response to the merging duration being greater than the upper limit of duration, the merging of the first type of song segment with adjacent song segments is abandoned.
[0089] In one scenario, if the combined length of a first-class song segment and an adjacent song segment exceeds the maximum length of a single segment, then the merging of the first-class song segment and the adjacent song segment is abandoned.
[0090] S690. Based on the first type of song fragments and the second type of song fragments whose durations conform to the duration rules, the set of song fragments is obtained.
[0091] In one scenario, song segments of type I and type II with durations between the minimum and maximum durations of a single segment are added to the song segment set.
[0092] Figure 7 This is a flowchart illustrating the duration rule validation method in a song processing scenario. Figure 7 As shown, the verification method includes: S710. Identify the first type of song fragment.
[0093] S720. Determine whether the duration of the first type of song segment conforms to the duration range. If yes, execute S780. If it is greater than the maximum duration of a single segment, execute S730. If it is less than the minimum duration of a single segment, execute S750.
[0094] S730. Based on the priority of music structure, lyric boundaries and rhythm information, the cutting points are determined layer by layer. Based on the cutting points, the first type of song segments are segmented to obtain the second type of song segments.
[0095] For example, prioritize finding cut points near the boundaries of the musical structure. If no suitable cut point is found, determine the cut point based on the boundaries of the lyrics. If no suitable cut point is found, determine the cut point based on rhythmic information.
[0096] S740. Determine whether the duration of the second type of song segment conforms to the duration range. If yes, execute S780. If it is greater than the maximum duration of a single segment, execute S730. If it is less than the minimum duration of a single segment, execute S750.
[0097] For example, if a tangent point exists near the boundary of the musical structure, the song segment is divided into two second-type song segments based on this tangent point. If the second-type song segment is still longer than the maximum duration of a single segment, it is first determined whether there is a musical structure boundary within the second-type song segment. If the second-type song segment is already a segment in the verse, it is considered that there is no musical structure boundary, and the tangent point is determined based on the lyric boundary. If the second-type song segment is divided into two new second-type song segments based on the tangent point corresponding to the lyric boundary, and if the new second-type song segment is still longer than the maximum duration of a single segment, it is first determined whether there is a musical structure boundary within the new second-type song segment. If the result is that there is no musical structure boundary, it is further determined whether the new second-type song segment has a lyric boundary. Assuming that the new second-type song segment is a segment corresponding to a complete line of lyrics, it is considered that the new second-type song segment does not have a lyric boundary. New second-type song segments can be further divided based on rhythm information.
[0098] S750, merge the first type of song fragment with the adjacent song fragment.
[0099] In this context, adjacent song segments can be either the preceding or following song segment of the first type of song segment. Optionally, adjacent song segments can be either the first type of song segment or the second type of song segment.
[0100] S760: Determine if the merging time exceeds the maximum duration of a single segment. If so, execute S770; otherwise, execute S780.
[0101] S770, Abandon merging the first type of song fragment with adjacent song fragments.
[0102] S780. Add song fragments to the song fragment collection.
[0103] This paper proposes a method to split song segments with a duration greater than the upper limit and merge song segments with a duration less than the lower limit. This approach avoids both excessively long individual song segments that could affect generation stability and excessively short segments that could lead to fragmented segments. This results in segment division that conforms to the content expression patterns of song data and improves the rationality of song data splitting.
[0104] Figure 8 This is a schematic diagram of the structure of a song processing device in one scenario, such as... Figure 8 As shown, the device includes: Display module 810 is used to display the first control on the first page; The first acquisition module 820 is used to obtain the song data corresponding to the first control in response to a trigger operation on the first control; The second obtaining module 830 is used to obtain a set of song segments corresponding to the song data in response to a confirmation operation on the song data. The set of song segments is used to generate a music video. The set of song segments is obtained by segmenting the song data based on a cut point corresponding to at least one of the following: music structure, lyrics information, rhythm information, and duration rules.
[0105] The aforementioned device displays a first control on a first page. In response to a trigger operation on the first control, it obtains the song data corresponding to the first control. In response to a confirmation operation on the song data, it obtains a set of song fragments corresponding to the song data. This set of song fragments is used to generate a music video. The set of song fragments is obtained by segmenting the song data based on cut points corresponding to at least one of the following: musical structure, lyrics, rhythm, and duration rules. This method combines musical structure, lyrics, rhythm, and duration rules to generate a set of cut points that consider the completeness of song expression, shot integrity, and overall audiovisual rhythm. This allows the song fragments obtained by segmenting the song data based on these cut points to generate a music video with the expected effect. This solves the problem of unreasonable song segmentation affecting video presentation, improves the rationality of song segmentation, enhances the smoothness of shot transitions in the music video, strengthens the adaptation of shot changes to musical rhythm, and optimizes the video presentation effect.
[0106] In one scenario, the first control includes at least one of a playback progress bar and tab items, wherein the playback progress bar includes a selection window corresponding to a song time segment, and the tab items are used to represent at least a portion of the song data.
[0107] In one scenario, the first obtaining module 820 is specifically used for: In response to an interactive operation on the selection window, the song time period corresponding to the selection window is adjusted; Obtain the song data corresponding to the time period of the song.
[0108] In one scenario, the first obtaining module 820 is specifically used for: In response to a trigger operation on the tag item, the song data corresponding to the triggered tag item is obtained, wherein the tag item is used to represent the song or a highlight segment in the song.
[0109] In one scenario, the song data is segmented to obtain the song fragment set based on a segmentation point corresponding to at least one of the following: musical structure, lyrics information, rhythm information, and duration rules. Based on the priority corresponding to the music structure, lyric boundaries, and rhythm information, the cut points corresponding to the music structure, lyric boundaries, and rhythm information are obtained, wherein the lyric boundaries include the start and end positions of the entire lyric line; The song data is segmented based on the cut points to obtain the first type of song fragments; The first type of song fragments whose duration does not meet the duration rules are split or merged to obtain the second type of song fragments; The song fragment set is obtained based on the first type of song fragments and the second type of song fragments whose durations conform to the duration rules.
[0110] In one scenario, the priority is positively correlated with the order in which the cut-off points are determined.
[0111] In one scenario, obtaining the cut point corresponding to the music structure, lyric boundaries, and rhythm information based on the priority of the music structure, lyric boundaries, and rhythm information includes: Based on the aforementioned music structure, the first type of cut points are obtained; Based on the lyrics boundary, the first type of cut points located inside the entire lyric line are corrected to obtain the second type of cut points; The third type of cut point is determined based on the rhythm information of the first song segment, wherein the first song segment is a song segment whose duration is greater than the upper limit of duration corresponding to the first type of cut point or the second type of cut point.
[0112] In one scenario, based on the first type of cut points corrected according to the lyrics boundary, multiple second type of cut points are obtained, including: Based on the deviation between the first type of cut point within the entire lyric and the lyric boundary, the first type of cut point is corrected to the start or end position of the lyric, thus obtaining the second type of cut point.
[0113] In one scenario, obtaining the cut point corresponding to the music structure, lyric boundaries, and rhythm information based on the priority of the music structure, lyric boundaries, and rhythm information includes: If a first type of cut point is not obtained based on the musical structure, a fourth type of cut point is determined based on the lyric boundary. The fifth type of cut point is determined based on the rhythm information of the second song segment, wherein the second song segment is a song segment whose duration is greater than the upper limit of the duration corresponding to the fourth type of cut point.
[0114] In one case, it also includes: In response to the absence or failure to meet the requirements of the cut points corresponding to the musical structure, lyric boundaries, and rhythm information, a sixth type of cut point is obtained based on the duration rule.
[0115] In one scenario, splitting or merging the first type of song fragments whose duration does not meet the duration rule to obtain the second type of song fragments includes: For the first type of song fragments whose duration exceeds the upper limit, the first type of song fragments are split based on the cut points corresponding to the musical structure, lyric boundaries, or rhythm information to obtain the second type of song fragments.
[0116] In one scenario, merging the first type of song fragments that do not meet the duration requirement to obtain the second type of song fragments includes: For the first type of song segment whose duration is less than the minimum duration limit, the first type of song segment is merged with the adjacent song segment; In response to a merging duration being less than or equal to the upper limit of the duration, the merging result of the first type of song segment and the adjacent song segment is taken as the second type of song segment; In response to a merge duration exceeding the stated duration limit, the merging of the first type of song segment with adjacent song segments is abandoned.
[0117] The above-described song processing device can execute the song processing method provided in any of the embodiments described herein, and has the corresponding functional modules and beneficial effects for executing the song processing method.
[0118] It is worth noting that the various units and modules included in the above-mentioned song processing device are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments described herein.
[0119] The following is for reference. Figure 9 This document illustrates a schematic diagram of an electronic device (e.g., a terminal device or server) 900 suitable for implementing the above-described methods. The terminal device referred to herein may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, personal digital assistants (PDAs), tablet computers (PADs), portable multimedia players (PMPs), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital televisions and desktop computers. Figure 9 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments described herein.
[0120] like Figure 9As shown, the electronic device 900 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage device 908 into a random access memory (RAM) 903. The RAM 903 also stores various programs and data required for the operation of the electronic device 900. The processing unit 901, ROM 902, and RAM 903 are interconnected via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0121] Typically, the following devices can be connected to I / O interface 905: input devices 906 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 907 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 908 including, for example, magnetic tapes, hard disks, etc.; and communication devices 909. Communication device 909 allows electronic device 900 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 9 An electronic device 900 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0122] In particular, according to embodiments of this document, the processes described in the above-referenced flowcharts can be implemented as computer software programs. For example, the technical solutions of this document include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device 909, or installed from a storage device 908, or installed from a ROM 902. When the computer program is executed by a processing device 901, it performs the functions defined in the methods of the embodiments of this document.
[0123] The names of messages or information exchanged between multiple devices in this document are for illustrative purposes only and are not intended to limit the scope of these messages or information.
[0124] The electronic device provided in this embodiment and the song processing method provided in the above technical solution belong to the same inventive concept. Technical details not described in detail in this document can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.
[0125] This article provides a computer storage medium storing a computer program that, when executed by a processor, implements the song processing method provided in the above embodiments.
[0126] It should be noted that the computer-readable medium mentioned above can be a computer-readable signal medium, a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory (EPROM, also known as flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this document, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. The transmitted data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0127] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol, such as HTTP (Hypertext Transfer Protocol), and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include Local Area Networks (LANs), Wide Area Networks (WANs), the Internet (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0128] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0129] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: Display the first control on the first page; In response to a trigger operation on the first control, obtain the song data corresponding to the first control; In response to the confirmation operation on the song data, a set of song segments corresponding to the song data is obtained. The set of song segments is used to generate a music video. The set of song segments is obtained by segmenting the song data based on a cut point corresponding to at least one of the following: music structure, lyrics information, rhythm information, and duration rules.
[0130] Computer program code for performing the operations described herein can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" programming language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including local area networks (LANs) or wide area networks (WANs), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0131] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this document. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0132] The modules or units described herein can be implemented in software or hardware. The names of modules or units do not necessarily constitute a limitation on the module or unit itself.
[0133] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include at least one of the following: Field-Programmable Gate Array (FPGA), Application-Specific Integrated Circuit (ASIC), Application-Specific Standard Product (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), etc.
[0134] In the context of this document, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory (flash memory), optical fibers, portable compact disk read-only memory, optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0135] The above description is merely a preferred embodiment and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of disclosure herein is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed herein that have similar functions.
[0136] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be interpreted as limiting the scope of this document. Certain features described in the context of individual implementations may also be implemented in combination in a single implementation. Conversely, various features described in the context of a single implementation may also be implemented individually or in any suitable sub-combination in multiple implementations.
[0137] Although the subject matter has been described using a programming language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative forms of implementing the claims.
Claims
1. A song processing method, comprising: Display the first control on the first page; In response to a trigger operation on the first control, obtain the song data corresponding to the first control; In response to the confirmation operation on the song data, a set of song segments corresponding to the song data is obtained. The set of song segments is used to generate a music video. The set of song segments is obtained by segmenting the song data based on a cut point corresponding to at least one of the following: music structure, lyrics information, rhythm information, and duration rules.
2. The method according to claim 1, wherein the first control includes at least one of a playback progress bar and a label item, wherein, The playback progress bar includes a selection window, which corresponds to a time segment of the song, and the label items are used to represent at least a portion of the song data.
3. The method according to claim 2, wherein obtaining the song data corresponding to the first control in response to a trigger operation on the first control includes: In response to an interactive operation on the selection window, the song time period corresponding to the selection window is adjusted; Obtain the song data corresponding to the time period of the song.
4. The method according to claim 2, wherein obtaining the song data corresponding to the first control in response to a trigger operation on the first control includes: In response to a trigger operation on the tag item, the song data corresponding to the triggered tag item is obtained, wherein the tag item is used to represent the song or a highlight segment in the song.
5. The method according to claim 1, wherein the song data is segmented to obtain the song fragment set based on a segmentation point corresponding to at least one of the following: music structure, lyrics information, rhythm information, and duration rules, comprising: Based on the priority corresponding to the music structure, lyric boundaries, and rhythm information, the cut points corresponding to the music structure, lyric boundaries, and rhythm information are obtained, wherein the lyric boundaries include the start and end positions of the entire lyric line; The song data is segmented based on the cut points to obtain the first type of song fragments; The first type of song fragments whose duration does not meet the duration rules are split or merged to obtain the second type of song fragments; The song fragment set is obtained based on the first type of song fragments and the second type of song fragments whose durations conform to the duration rules.
6. The method according to claim 5, wherein the priority is positively correlated with the order of determining the tangent point.
7. The method according to claim 6, wherein obtaining the cut point corresponding to the music structure, lyric boundaries, and rhythm information based on the priority corresponding to the music structure, lyric boundaries, and rhythm information includes: Based on the aforementioned music structure, the first type of cut points are obtained; Based on the lyrics boundary, the first type of cut points located inside the entire lyric line are corrected to obtain the second type of cut points; The third type of cut point is determined based on the rhythm information of the first song segment, wherein the first song segment is a song segment whose duration is greater than the upper limit of duration corresponding to the first type of cut point or the second type of cut point.
8. The method according to claim 7, wherein obtaining a plurality of second-type cut points based on the first-type cut points of the lyrics boundary correction portion includes: Based on the deviation between the first type of cut point within the entire lyric and the lyric boundary, the first type of cut point is corrected to the start or end position of the lyric, thus obtaining the second type of cut point.
9. The method according to claim 6, wherein obtaining the cut point corresponding to the music structure, lyric boundaries, and rhythm information based on the priority corresponding to the music structure, lyric boundaries, and rhythm information includes: If a first type of cut point is not obtained based on the musical structure, a fourth type of cut point is determined based on the lyric boundary. The fifth type of cut point is determined based on the rhythm information of the second song segment, wherein the second song segment is a song segment whose duration is greater than the upper limit of the duration corresponding to the fourth type of cut point.
10. The method of claim 5, further comprising: In response to the absence or failure to meet the requirements of the cut points corresponding to the musical structure, lyric boundaries, and rhythm information, a sixth type of cut point is obtained based on the duration rule.
11. The method according to claim 5, wherein splitting or merging the first type of song fragments whose duration does not meet the duration rule to obtain the second type of song fragments includes: For the first type of song fragments whose duration exceeds the upper limit, the first type of song fragments are split based on the cut points corresponding to the musical structure, lyric boundaries, or rhythm information to obtain the second type of song fragments.
12. The method according to claim 5, wherein merging the first type of song fragments whose duration does not meet the requirement to obtain the second type of song fragments includes: For the first type of song segment whose duration is less than the minimum duration limit, the first type of song segment is merged with the adjacent song segment; In response to a merging duration being less than or equal to the upper limit of the duration, the merging result of the first type of song segment and the adjacent song segment is taken as the second type of song segment; In response to a merge duration exceeding the stated duration limit, the merging of the first type of song segment with adjacent song segments is abandoned.
13. A song processing device, comprising: The display module is used to display the first control on the first page; The first obtaining module is used to obtain the song data corresponding to the first control in response to a trigger operation on the first control; The second obtaining module is used to obtain a set of song segments corresponding to the song data in response to a confirmation operation on the song data. The set of song segments is used to generate a music video. The set of song segments is obtained by segmenting the song data based on a cut point corresponding to at least one of the following: music structure, lyrics information, rhythm information, and duration rules.
14. An electronic device, the electronic device comprising: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the song processing method as described in any one of claims 1-12.
15. A storage medium comprising computer-executable instructions, which, when executed by a computer processor, are used to perform the song processing method as described in any one of claims 1-12.
16. A computer program product comprising a computer program that, when executed by a processor, implements the song processing method as described in any one of claims 1-12.