Song playing method and device, electronic equipment and storage medium

By analyzing song feature information to determine the target cut-out and cut-in time points, and applying transition processing parameters, the problem of abrupt sound during traditional song switching is solved, achieving a seamless transition effect.

CN121983007APending Publication Date: 2026-05-05HANGZHOU NETEASE CLOUD MUSIC TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HANGZHOU NETEASE CLOUD MUSIC TECH CO LTD
Filing Date
2025-12-31
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Traditional song transitions often involve incomplete splicing, resulting in abrupt sounds and disrupting the listening experience.

Method used

By analyzing the audio and structural features of the first and second songs, candidate cutout and cut-in time points are determined, and target cutout and cut-in time points are determined based on the difference information. Seamless connection is achieved by combining transition processing parameters.

Benefits of technology

It achieves a smooth and natural transition between the two songs, enhancing the listening experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121983007A_ABST
    Figure CN121983007A_ABST
Patent Text Reader

Abstract

The invention discloses a song playing method and device, electronic equipment and a storage medium, and the method comprises the steps: determining a target cut-out time point of a first song from candidate cut-out time points based on first difference information of audio feature information of the first song and a second song, determining a target cut-in time point of the second song from the candidate cut-in time points; determining transition processing parameters based on the audio feature information of the first song and the second song; and processing and playing a first music segment corresponding to the target cut-out time point of the first song and a second music segment corresponding to the target cut-in time point of the second song based on the transition processing parameters in response to the situation that the first song is played to the position corresponding to the target cut-out time point. Therefore, the transition processing parameters used for joining the two songs can be dynamically determined, and then corresponding song joining processing is executed, so that the seamless joining effect of rhythm coherence and natural listening feeling is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of information processing technology, and in particular to a song playback method, apparatus, electronic device, and storage medium. Background Technology

[0002] In traditional music playback, switching from song A to song B typically involves waiting for the last second of song A to play before starting the first second of song B. Among related methods, the transition between the two songs mainly falls into the following categories: natural transition, where the end of song A seamlessly connects to the beginning of song B; or, using energy detection methods to skip the beginning and end gaps in the data, or even overlapping some data for a fade-in / fade-out effect, making the song transition more continuous. However, these methods may result in incomplete transitions, abrupt changes in sound or music theory, and noticeable auditory breaks at the moment of switching, disrupting the overall listening experience. Summary of the Invention

[0003] This application provides a song playback method, apparatus, electronic device, and storage medium. By using the feature information between the songs to be connected, the transition processing parameters of the two songs are determined to achieve a seamless connection effect that conforms to the characteristics of the songs.

[0004] In a first aspect, embodiments of this application provide a song playback method, the method comprising: Obtain feature information of a first song and a second song, the feature information including audio feature information and structural feature information, wherein the second song is a song played after the first song; Based on the structural feature information of the first song, the candidate cut-out time point of the first song is determined, and based on the structural feature information of the second song, the candidate cut-in time point of the second song is determined. Based on the first difference information of the audio feature information of the first song and the second song, the target cut-out time point of the first song is determined from the candidate cut-out time point, and the target cut-in time point of the second song is determined from the candidate cut-in time point; Based on the audio feature information of the first song and the second song, the transition processing parameters of the first music segment corresponding to the target cut-out time point of the first song and the second music segment corresponding to the target cut-in time point of the second song are determined. In response to the first song playing to the position corresponding to the target cut-out time point, based on the transition processing parameters, the first music segment corresponding to the target cut-out time point of the first song and the second music segment corresponding to the target cut-in time point of the second song are processed and played.

[0005] Secondly, embodiments of this application provide a song playback device, the device comprising: The acquisition module is used to acquire feature information of a first song and a second song. The feature information includes audio feature information and structural feature information. The second song is a song played after the first song. The first determining module is used to determine the candidate cut-out time point of the first song based on the structural feature information of the first song, and to determine the candidate cut-in time point of the second song based on the structural feature information of the second song. The second determining module is used to determine the target cut-out time point of the first song from the candidate cut-out time points based on the first difference information of the audio feature information of the first song and the second song, and to determine the target cut-in time point of the second song from the candidate cut-in time points. The third determining module is used to determine the transition processing parameters of the first music segment corresponding to the target cut-out time point of the first song and the second music segment corresponding to the target cut-in time point of the second song based on the audio feature information of the first song and the second song. The playback processing module is used to respond to the first song playing to the position corresponding to the target cut-out time point, and to process and play the first music segment corresponding to the target cut-out time point of the first song and the second music segment corresponding to the target cut-in time point of the second song based on the transition processing parameters.

[0006] Thirdly, embodiments of this application also provide an electronic device, which includes a processor and a memory, wherein the memory stores a computer program that, when executed by the processor, causes the processor to perform steps of a method for playing any song.

[0007] Fourthly, embodiments of this application also provide a computer-readable storage medium including a processor and a memory, wherein the memory stores a computer program that, when executed by the processor, causes the processor to perform steps of a method for playing any song.

[0008] Fifthly, embodiments of this application also provide a computer program product, including a computer program stored in a computer-readable storage medium; when a processor of an electronic device reads the computer program from the computer-readable storage medium, the processor executes the computer program, causing the electronic device to perform the steps of any of the song playback methods provided in embodiments of this application.

[0009] By comprehensively analyzing the audio and structural features of the first and second songs, the solution of this application can identify the target cut-out and target cut-in time points that conform to the song features and listening habits. The transition processing parameters are dynamically determined by combining the audio feature information of the first and second songs, thereby achieving a smooth and natural seamless transition between the two songs. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 This is a schematic diagram of an implementation environment scenario for the song playback method provided in this application embodiment; Figure 2 This is a flowchart illustrating the song playback method provided in the embodiments of this application; Figure 3 This is a schematic diagram of voice separation in the song playback method provided in the embodiments of this application; Figure 4 This is a schematic diagram of drum sound separation in the song playback method provided in the embodiments of this application; Figure 5 This is a schematic diagram of a music segment in the song playback method provided in the embodiments of this application; Figure 6 This is a schematic diagram of an architecture of the song playback method provided in the embodiments of this application; Figure 7 This is a schematic diagram of the Gaussian curve of the song playback method provided in the embodiments of this application; Figure 8 This is a schematic diagram of the structure of the song playback device provided in the embodiments of this application; Figure 9 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0012] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application. At the same time, in the description of the embodiments of this application, the terms "first," "second," etc., are only used to distinguish descriptions and should not be construed as indicating or implying relative importance. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more features. In the description of the embodiments of this application, "multiple" means two or more, unless otherwise explicitly specified.

[0013] This application provides a song playback method, apparatus, electronic device, and computer-readable storage medium.

[0014] Specifically, this embodiment will be described from the perspective of a song playback device, which can be integrated into an electronic device, meaning that the song playback method of this application embodiment can be executed by an electronic device. This electronic device can be a server, a terminal, or other similar device.

[0015] The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery network (CDN) acceleration services, and big data and artificial intelligence platforms. The terminal can be a device that includes both receiving and transmitting hardware, i.e., a device with receiving and transmitting hardware capable of performing bidirectional communication over a bidirectional communication link. Terminals can include, but are not limited to, mobile phones, computers, smart voice interaction devices, smart home appliances, vehicle terminals, and aircraft. Terminals and servers can be directly or indirectly connected via wired or wireless communication methods; this application does not impose any restrictions.

[0016] For example, such as Figure 1As shown, the terminal device obtains the playlist playback request from the target user and sends the playlist playback request to the server. The server obtains the first song and the second song from the target user's playlist; it extracts features from the first and second songs to obtain the song's feature information and sends it to the terminal device. The terminal device uses this feature information to obtain the feature information of the first and second songs, including audio feature information and structural feature information. The second song is the song played after the first song. Based on the structural feature information of the first song, the candidate cut-out time point of the first song is determined, and based on the structural feature information of the second song, the candidate cut-in time point of the second song is determined. Based on the first song... The system uses the first difference information of the audio feature information of the first song and the second song to determine the target cut-out time point of the first song from the candidate cut-out time point, and the target cut-in time point of the second song from the candidate cut-in time point; based on the audio feature information of the first song and the second song, it determines the transition processing parameters of the first music segment corresponding to the target cut-out time point of the first song and the second music segment corresponding to the target cut-in time point of the second song; in response to the first song playing to the position corresponding to the target cut-out time point, it processes and plays the first music segment corresponding to the target cut-out time point of the first song and the second music segment corresponding to the target cut-in time point of the second song based on the transition processing parameters.

[0017] It should be noted that, Figure 1 The illustrated scenario of the song playback method is merely an example. The implementation environment of the song playback method described in this application is intended to more clearly illustrate the technical solutions of this application and does not constitute a limitation on the technical solutions provided in this application. Those skilled in the art will recognize that, with the evolution of data processing and the emergence of new business scenarios, the technical solutions provided in this application are also applicable to similar technical problems.

[0018] The solutions provided in this application are specifically illustrated through the following embodiments. It should be noted that the order of description of the following embodiments is not intended to limit the preferred order of the embodiments.

[0019] This embodiment will be described from the perspective of a song playback device, which can be integrated into an electronic device, which can be a terminal and / or a server, and this application does not impose any limitations on it.

[0020] This application provides a method for playing songs. Please refer to [link / reference]. Figure 2 , Figure 2 The specific process of the song playback method provided in this application embodiment can be summarized in the following steps 201 to 205: Step 201: Obtain the feature information of the first song and the second song. The feature information includes audio feature information and structural feature information. The second song is the song played after the first song.

[0021] The terms "first song" and "second song" are used to describe two adjacent songs playing consecutively. "First song" refers to the song currently playing and about to be switched to (the preceding song). "Second song" refers to the song immediately following the first song (the following song).

[0022] In one implementation scenario, the first and second songs originate from the same playlist (or song list) preset by the user or generated by the system, and the two songs have a pre-set playback order. In another implementation scenario, the first and second songs originate from different playlists or music collections. For example, different playlists that the user actively switches between during listening, or songs selected by the system from different sources for continuous playback based on recommendation algorithms, random playback logic, etc.

[0023] Audio feature information refers to the characteristic information extracted from audio signals that reflects the inherent musical attributes of a song. For example, audio feature information may include, but is not limited to, rhythmic features (such as beats per minute, time signature, and beat timestamp), tonality features (such as key type (e.g., C major, A minor), energy and spectral features (such as loudness sequences across the entire frequency band and sub-bands, measured in LUFS), and auxiliary information such as genre labels and instrument types. Structural feature information refers to information about the internal organization of a song. For example, structural feature information may include the start and end times of the vocals (i.e., the starting and ending positions of the vocals on the timeline), the start and end times of the playing of a specific instrument (such as drums, bass, or other instruments with rhythmic guidance), and information about the structure of musical sections. The music segment structure information describes the overall segmentation of the song, such as the types of music segments like intro, verse, chorus, bridge, and outro, as well as their corresponding start and end times.

[0024] In one implementation scenario, the steps of obtaining feature information for the first and second songs are pre-completed by the server. After standardizing the original audio files, the server inputs them into multiple analysis models to obtain audio feature information. For example, rhythmic features can be extracted using a beat tracking model (e.g., BeatNet based on a CRNN architecture); the key signature can be output using a tonality recognition model and mapped to Camelot encoding; corresponding audio tracks can be separated using a vocal and drum separation model (e.g., Demucs), and the start and end times of vocals and drums can be determined using an energy threshold detection algorithm; simultaneously, the songs are input into a music structure analysis tool (e.g., All-In-One) to obtain the music segment structure information for each musical section.

[0025] In one possible implementation, the music segment structure information can be a data structure that represents the music segments of a song on the timeline as ordered pairs (timestamps, music segment labels). One possible representation is: [(0.0s, intro), (23.5s, verse), (48.2s, chorus), (72.8s, verse), (97.1s, chorus), (121.6s, bridge), (139.3s, chorus), (163.9s, outro)].

[0026] This music segment structure information indicates that a song starts from the beginning, with 0.0 seconds to 23.5 seconds as the intro, 23.5 seconds to 48.2 seconds as the verse, 48.2 seconds to 72.8 seconds as the chorus, and so on, until the outro. Each timestamp marks the start time of the corresponding music segment, and the interval between adjacent timestamps is the duration of the music segment described by that label.

[0027] Full-band and frequency-band loudness sequences refer to numerical sequences formed by quantizing the energy intensity of audio signals over time, typically expressed in loudness units relative to full scale (LUFS). One specific representation is as follows: Full-band loudness sequence: [ 24.1, 22.8, 20.3, 18.7, 19.2,…]LUFS, with a sampling interval of one data point every 100 milliseconds; Low-frequency (20Hz–250Hz) loudness sequence: [ 30.5, 28.9, 25.1, 22.4, 23.0,…]LUFS; Mid-frequency (250Hz–4kHz) loudness sequence: [ 26.3, 24.7, 21.8, 19.5, 20.1,…]LUFS; High-frequency loudness sequence (4kHz–20kHz): [ 35.2, 33.6, 30.4, 28.9, 29.3,…]LUFS.

[0028] In some embodiments, the above four loudness sequences can be synchronously aligned with the same time axis to reflect the energy change trends of the entire song in the whole and in different frequency bands.

[0029] In one possible implementation, the song can be used as input, analyzed by a beat and downbeat detection algorithm (such as the BeatNet model based on the CRNN architecture), and the output can be the song's audio feature information, including the number of beats per unit time (the unit time can be 1 minute, 1 second, etc., and the number of beats per unit time can be BPM), the time signature (such as 4 / 4, 3 / 4, etc.), and the beat sequence arranged in chronological order. Taking a song with a time signature of 4 / 4 as an example, its output beat sequence can be represented as: [(time1,1),(time2,2),(time3,3),(time4,4),(time5,1),(time6,2),(time7,3),(time8,4),…]. Here, time1, time2, etc. are the specific timestamps of the beats (in seconds), the numbers in parentheses indicate the position of the beat in the current measure, and the number 1 represents an downbeat (i.e., the first beat of each measure).

[0030] Step 202: Based on the structural feature information of the first song, determine the candidate cut-out time point of the first song, and based on the structural feature information of the second song, determine the candidate cut-in time point of the second song.

[0031] Candidate cutout time points refer to one or more time positions in the first song that are suitable as cutout points, identified based on its structural feature information.

[0032] In some embodiments, for the first song, based on its structural feature information, several suitable cut-out points are identified on the timeline. For example, the cut-out point can be located at a point where the vocals end, the performance ends, or other easily connected locations. For instance, if the vocals in the first song end at 180.0 seconds and there is a purely instrumental outro afterward, the accented beat of the first measure after the vocals end (e.g., 181.2 seconds) can be considered as a candidate cut-out point; similarly, if the drums stop at 175.0 seconds, the first accented beat after the drums stops can be considered as another candidate cut-out point. Furthermore, the system can search based on all preset key time point types (e.g., vocal key time points, performance key time points, structural key time points, measure key time points, etc.) to generate a set of candidate cut-out points within the effective playback range of the first song.

[0033] Candidate entry points refer to one or more potential locations in the second song that are suitable as entry points, identified based on its structural features. In some embodiments, for the second song, several suitable entry points are identified on the timeline based on its structural features. The entry point can be located at a time point that is easily connected, such as the start of the vocals or the start of the instrumental performance. For example, if the intro of the second song lasts for 8 seconds and the vocals begin at the 9th second, the end of the 8th second (such as the downbeat of the next measure corresponding to the 8th second) can be considered as a candidate entry point; if the drum sound first appears at the 5th second, the start of that drum sound can also be considered as another candidate entry point. The system can search based on all preset key time point types (such as vocal key time points, instrumental key time points, structural key time points, measure key time points, etc.) to generate a set of candidate entry points within the intro section of the second song.

[0034] Step 203: Based on the first difference information of the audio feature information of the first song and the second song, determine the target cut-out time point of the first song from the candidate cut-out time points, and determine the target cut-in time point of the second song from the candidate cut-in time points.

[0035] The first difference information refers to the difference obtained by comparing the audio feature information of the first song and the second song, and is used to assess the degree of auditory conflict that may occur when the two songs are connected. The difference information includes at least: rhythmic difference information, such as the absolute difference in BPM between the two songs (ΔBPM=|BPM1). BPM2; Tonal compatibility, such as using the Camelot wheel to determine whether two songs belong to the same key, related key, or closely related key; if they belong to the same key, they are considered compatible; otherwise, they are considered conflicting; Loudness difference, such as the absolute value of the difference in average loudness between two songs near their respective candidate cutout time points and candidate cutin time points (ΔLUFS=|LUFS1). LUFS2|).

[0036] It should be noted that the first difference information can be global audio feature difference information, that is, the difference between audio feature information extracted from the entire song. For example, the difference in beats per minute (BPM) between the first and second songs, whether the tonal compatibility is the same (same tone, related tone, or distant related tone), loudness (LUFS) difference, and whether the genre labels are consistent, etc. Alternatively, the first difference information can also be local audio feature difference information, that is, calculating the difference in local audio feature information between the audio segments of the first song near the candidate cutout time point (e.g., 1-2 seconds before and after) and the audio segments of the second song near the candidate cutin time point (e.g., 1-2 seconds before and after). For example, under the combination of key drum time points, comparing the difference in energy and beat rate near the two key drum time points.

[0037] The target cut-out time point for the first song refers to the specific moment at which playback actually stops (or transition processing begins) from the first song. This time point is not arbitrarily assigned, but rather obtained by matching multiple candidate cut-out time points determined based on the structural feature information of the first song.

[0038] The target entry point for the second song refers to the specific moment at which the second song begins playing (or is incorporated into the transition process). This time point is also obtained by matching from a set of candidate entry points generated based on the structural features of the second song itself.

[0039] In some embodiments, candidate cutout time points and candidate cutin time points can be two sets of data: one set consists of multiple candidate cutout time points for the first song, and the other set consists of multiple candidate cutin time points for the second song. The output is the determined target cutout time point and target cutin time point. Specifically, multiple type combinations formed by each candidate cutout time point of the first song and each candidate cutin time point of the second song can be traversed. For each type combination, audio feature information of the first song within a time window centered on the candidate cutout time point is extracted, and audio feature information of the second song within a time window centered on the candidate cutin time point is also extracted. Subsequently, a first difference information between these two audio feature information is calculated. During the calculation, corresponding weights can be assigned according to the importance of different features; for example, higher weights can be assigned to beat rate differences and tonality compatibility.

[0040] After calculating the first difference information corresponding to all type combination pairings, the optimal type combination is selected from all type combinations according to a predetermined optimization strategy (e.g., finding the type combination with the smallest sum of first difference information, or requiring that all differences are below a threshold). The candidate cut-out time point of the first song in the optimal type combination is finally determined as the target cut-out time point, and the candidate cut-in time point of the second song in the optimal type combination is determined as the target cut-in time point.

[0041] Step 204: Based on the audio feature information of the first song and the second song, determine the transition processing parameters of the first music segment corresponding to the target cut-out time point of the first song and the second music segment corresponding to the target cut-in time point of the second song.

[0042] The first music segment corresponding to the target cut-out time point of the first song refers to an audio clip of a certain duration extending forward or backward from that target cut-out time point, used to connect transition areas in the processing. For example, the audio within 500 milliseconds before the target cut-out time point can be taken as the first music segment, representing the part of the first song that is about to end. The second music segment corresponding to the target cut-in time point of the second song refers to an audio clip extending backward or forward from the target cut-in time point, such as the audio from the start of that cut-in time point to 500 milliseconds thereafter, representing the part of the second song that is about to begin. These two music segments together constitute the object to be processed by audio fusion.

[0043] In an optional embodiment, the first music segment and the second music segment are audio segments of fixed duration, or the first music segment and the second music segment are not audio segments of fixed duration, and the specific duration of the first music segment and the second music segment can be dynamically determined, for example, until the end time of the current music segment (such as a musical phrase).

[0044] Transition processing parameters are a set of configuration parameters used to control the audio signal processing of the two music segments during overlapping playback, so that the two songs, which may have different styles, rhythms or energy, can be connected smoothly and naturally in terms of sound.

[0045] In some embodiments, the transition processing parameters include at least transition time and song processing parameters. The transition time determines the duration of the overlapping playback; the song processing parameters can be further subdivided into volume control parameters, frequency control parameters, and rhythm synchronization parameters, which are used to adjust volume, spectral energy distribution, and beat rate, respectively.

[0046] In some embodiments, transition processing parameters can be determined in various ways. For example, a gain compensation parameter can be calculated by comparing the average loudness or peak level of the first and second music segments to ensure that there are no abrupt volume jumps during the transition. Another example is that the transition duration and the shape of the volume control curve can be determined by comparing the tempo and beat positions of the first and second music segments. Yet another example is that the transition processing parameters can include a filter that reduces the low-frequency components of the first music segment if the low-frequency energy of the first music segment is higher than that of the second music segment, by comparing the energy distribution of the two music segments (such as the loudness sequence of frequency divisions).

[0047] Step 205: In response to the first song playing to the position corresponding to the target cut-out time point, based on the transition processing parameters, process and play the first music segment corresponding to the target cut-out time point of the first song and the second music segment corresponding to the target cut-in time point of the second song.

[0048] In an optional implementation, when the current playback position reaches the target cut-out time point during the playback of the first song, the overlay playback process is triggered, that is, the playback synchronously jumps to the target cut-in time point of the second song and reads the second music segment. Simultaneously, the subsequent music segments of the first song from the target cut-out time point, i.e., the first music segment, continue playing. During the subsequent transition time, the first and second music segments undergo audio mixing processing: based on volume control parameters, a gradually decaying gain curve (such as Gaussian, linear, or exponential fade-out) is applied to the first music segment, while a gradually increasing gain curve (Gaussian, linear, or exponential fade-in) is applied to the second music segment; based on frequency control parameters, filtering is applied to one or two music segments, for example, performing a low-pass sweep on the first music segment to gradually reduce high-frequency components, or performing a high-pass sweep on the second music segment to gradually release low-frequency components; the beat rate can also be adjusted according to rhythm synchronization parameters, by using a time-stretching algorithm (such as PSOLA or Phase Vocoder) to slightly speed up one or two music segments of one or two songs while maintaining the pitch, so that their beats are aligned.

[0049] After the transition period ends, playback of the first song stops, and playback of the second song continues normally, thus completing a smooth transition from the first song to the second song. At this point, even if there is still some unplayed music in the first song after the transition period, the first song will not be played again.

[0050] It should be understood that the transition point between adjacent songs, that is, the position corresponding to the target cut-out time point in the first song, refers to the start time point corresponding to the first musical segment in the first song, which depends on the target cut-out time point and the transition time. Reading the second musical segment refers to reading the song content corresponding to the start time point of the second musical segment, which depends on the target cut-in time point and the transition time.

[0051] It should be understood that the types of information included in the feature information may include more or fewer than those shown above, and this application does not limit this.

[0052] In some embodiments, determining candidate cut-out time points for the first song based on structural feature information of the first song, and determining candidate cut-in time points for the second song based on structural feature information of the second song, includes: Based on the structural feature information of the first song and multiple preset key time point types, the first key time point under each preset key time point type is searched in the first song as a candidate cutout time point. Based on the structural features of the second song and multiple preset key time point types, the second key time point under each preset key time point type is searched in the second song as a candidate entry time point.

[0053] Predefined key time point types refer to a predefined set of time position categories with musical semantic significance, used to guide the extraction of potential switching points from the structural features of a song that conform to auditory logic and musical performance rules. For example, preset key time point types include, but are not limited to, vocal key time points, instrumental key time points, structural key time points, and measure key time points.

[0054] In an optional implementation, the first song can be parsed to extract its structural feature information, including the start and end times of the vocals, the start and end times of the performance of a specified instrument, the division of musical sections (such as intro, verse, chorus, coda, etc.) and the start and end times of each musical section, as well as the beat sequence and accent positions. Subsequently, based on the above information, according to multiple preset key time point types, the timeline of the first song is searched one by one to find whether a corresponding preset key time point type exists, thus obtaining the first key time point. For example, regarding the type of key timing points for vocals, it could be the end time of the vocals in the first song, or the downbeat of the measure following the end time of the vocals. Regarding the type of key timing points for performance, the nearest end time of the drum can be found within the time range after the end time of the vocals; the end time of the drum in the first song, or the downbeat of the measure following the end time of the drum, can be used as the performance key timing point. Regarding the type of key timing points for structure, the beginning time of the coda, or the downbeat of the measure following the beginning time of the coda, can be selected as the structural key timing point. Regarding the type of key timing points for a measure, the first downbeat within the time interval after the end time of the vocals but before the end time of the song can be selected as the measure key timing point. All successfully identified first key timing points together constitute the candidate cut-out timing point set for the first song.

[0055] In an optional implementation, the second song can be parsed to extract its structural feature information, including the start and end times of the vocals, the start and end times of the performance of a specified instrument, the division of musical sections (such as intro, verse, chorus, coda, etc.) and the start and end times of each musical section, as well as the beat sequence and accent positions. Subsequently, based on the above information, according to multiple preset key time point types, the timeline of the second song is searched one by one to find whether there is a corresponding preset key time point type, thus obtaining the second key time point. For example, regarding the type of key vocal timing, it could be the start time of the vocals in the second song, or the downbeat of the measure preceding the start time of the vocals; regarding the type of key performance timing, the nearest drum end time can be found within the time range before the end time of the vocals, and the drum end time of the second song, or the downbeat of the measure preceding the drum end time, can be used as the performance key timing; regarding the type of key structural timing, the end time of the intro section, or the downbeat of the measure preceding the end time of the intro section, can be selected as the structural key timing; regarding the type of key measure timing, the last downbeat time within the time interval before the start time of the vocals and after the start time of the song can be selected as the measure key timing. All successfully identified second key timing points together constitute the candidate entry timing set for the second song.

[0056] It should be noted that the above search process is not a simple enumeration, but rather introduces temporal relationships to ensure the musicality of key time points. For example, key performance time points must be after (for the first song) or before (for the second song) key vocal time points, and key time points in a measure must not be too close to the start or end of the song, in order to filter out pseudo-key time points that exist but do not conform to the logical connection.

[0057] It should be understood that the types of multiple preset key time points may include more or fewer than those shown above, and this application does not limit this.

[0058] In some embodiments, the structural feature information includes at least one of the following: the start and end times of the singing, the start and end times of the performance of a specified instrument, and the structural information of a musical segment; the structural information of a musical segment includes the type of at least one musical segment and the start and end times of each musical segment. Preset key time point types include: vocal key time points, specified instrumental performance key time points, structural key time points, and measure key time points; The key moments of the singing are determined based on the start and end times of the singing. The key timing points for the performance of a specified instrument are determined based on the start and end times of the performance; The key time points in the structure are determined based on the start and end time information of the specified music segments; Key time points in each section are determined based on preset retake time points.

[0059] The start and end times of the singing include the start and end times of the vocals on the timeline; the start and end times of the performance of a designated instrument include the start and end times of rhythmic percussion instruments such as the bass drum and snare drum.

[0060] Musical segment structure information includes the type of each musical segment, such as intro, verse, chorus, interlude, and coda, as well as its corresponding start and end times.

[0061] In an optional implementation, if there are start and end times for the singing, the key timing points of the singing are determined based on the start and end times of the singing. The key timing points for the performance of a designated instrument are determined based on the start and end times of the performance and the start and end times of the singing. The key time points in the structure are determined based on the start and end time information of the specified music segment and the start and end time points of the singing. Key timing points in a section are determined based on the replay timing and the start and end times of the song.

[0062] In an optional embodiment, if there is no start and end time point for the singing, the candidate cutout time point may include: a key singing time point determined based on the end time point of the singing in the first song; a key performance time point determined based on the end time point of the performance in the first song; a key structural time point determined based on the start time point of the coda in the first song; and a key measure time point determined based on the first accented beat time point of a first preset number of measures (e.g., the penultimate measure) starting from the end time point of the first song.

[0063] In an optional embodiment, if there is no singing start and end time point, the candidate cut-in time point may include: a singing key time point determined based on the singing start time point of the second song; a performance key time point determined based on the performance start time point of the second song; a structural key time point determined based on the end time point of the intro section of the second song; and a measure key time point determined based on the first downbeat time point of a second preset number of measures (e.g., the second measure) starting from the singing start time point of the second song.

[0064] In an optional embodiment, if there are start and end times for the singing, the candidate cutout times may include: a key singing time point determined based on the end time of the singing of the first song; a key performance time point corresponding to the end time of the performance that occurs after the key singing time point; a key structural time point corresponding to the start time of the coda that occurs after the key singing time point; and a key measure time point corresponding to the first accented beat time point of the first preset number of measures (i.e., the strong beat of the measure) that occurs after the key singing time point and starts from the end time of the song.

[0065] In an optional embodiment, if there are start and end times for the singing, the candidate cut-in times may include: a key singing time point determined based on the start time of the second song; a key performance time point corresponding to the start time of the performance that occurs before the key singing time point; a structural key time point corresponding to the end time of the intro section that occurs before the key singing time point; and a key measure time point corresponding to the first accent time point of the second preset number of measures that occurs before the key singing time point and starts from the start time of the song.

[0066] When the end time of the song is identified in the first song, the remaining candidate cut-out time points of the first song can be time points that appear after the end time point of the song in the first song; or when the start time point of the song is identified in the second song, the remaining candidate cut-out time points of the second song can be time points that appear before the start time point of the song in the second song.

[0067] In some embodiments, the first key time point of the first song includes at least one of the following: the end time point of the singing of the first song, and / or the end time point of the performance after the end time point of the singing of the first song, the start time point of the coda, and the first downbeat time point earlier than the end time point of the first song. The second key time point of the second song is the start time of the second song's vocals, and / or at least one of the following: the start time of the performance before the start time of the second song's vocals, the end time of the intro, or the second downbeat time point later than the start time of the second song.

[0068] In some embodiments, the type combination of the preset key time point types of candidate cutout time points and candidate cut-in time points is set with corresponding priorities. The priority order from high to low is: all are performance key time points, all are structural key time points, a type combination of vocal key time points and measure key time points, and other type combinations. Based on the first difference information of the audio feature information of the first song and the second song, the target cut-out time point of the first song is determined from the candidate cut-out time points, and the target cut-in time point of the second song is determined from the candidate cut-in time points, including: Based on the preset key time point type to which the candidate cutout time point belongs and the preset key time point type to which the candidate cutin time point belongs, determine at least one type combination composed of the candidate cutout time point and the candidate cutin time point. Based on the highest priority type combination and the first difference information of the audio feature information of the first song and the second song, the target cut-out time point of the first song and the target cut-in time point of the second song are determined.

[0069] The preset key timing point types include vocal key timing points, instrumental key timing points, structural key timing points, and measure key timing points. Different type combinations have different priorities, in descending order: combinations where all are instrumental key timing points (e.g., the end of the drumbeat in the first song and the beginning of the drumbeat in the second song); combinations where all are structural key timing points (e.g., the beginning of the coda in the first song and the end of the intro in the second song); combinations of vocal and measure key timing points (e.g., the end of the vocal part in the first song and the measure key timing point in the second song); and other combinations not covered above (e.g., mixed cases of instrumental and structural key timing points, measure key timing points and vocal key timing points, etc.).

[0070] In one specific embodiment, all candidate cutout time points of the first song and all candidate cutin time points of the second song are traversed to obtain the preset key time point type corresponding to each time point, and all possible type combinations are generated. For example, if the first song contains one performance key time point and one structural key time point, and the second song contains one performance key time point and one measure key time point, then type combinations such as (performance key time point, performance key time point), (performance key time point, measure key time point), (structural key time point, performance key time point), and (structural key time point, measure key time point) can be formed. Subsequently, according to the priority order of the above type combinations, the type combination with the highest priority is selected.

[0071] In one specific embodiment, the candidate cut-out time point of the first song in the highest priority type combination is finally determined as the target cut-out time point, and the candidate cut-in time point of the second song in the optimal type combination is determined as the target cut-in time point.

[0072] In one specific embodiment, the target cut-out time point for the first song and the target cut-in time point for the second song can be determined based on first difference information, such as indicators of rhythmic difference information (e.g., ΔBPM), tonality compatibility information (based on Camelot wheel to determine whether they are in the same key, related key, or closely related key), and loudness difference information (ΔLUFS). For example, in a type combination where both songs are key performance time points, if both songs have a BPM greater than 80 and are tonally compatible, then this highest priority type combination is selected as the target cut-out time point for the first song and the target cut-in time point for the second song.

[0073] It should be understood that the types of combinations may include more or fewer than those shown above, and this application does not impose any limitation on this. The priority of each type combination can also be flexibly configured according to the actual situation.

[0074] In some embodiments, the first difference information includes the difference between the number of beats per unit time and the preset beat threshold between the first song and the second song, and / or, the tonality compatibility information between the first song and the second song; Based on the highest priority type combination and the first difference information of the audio feature information of the first song and the second song, the target cut-out time point of the first song is determined from the candidate cut-out time points, and the target cut-in time point of the second song is determined from the candidate cut-in time points, including: When the highest priority type combination has key performance time points and / or the number of beats per unit time of the first song and the second song both exceed the preset beat number threshold, the target cut-out time point is determined based on the key singing time point of the first song, and the target cut-in time point is determined based on the key performance time point of the second song. When the highest priority type combination is that both have structural key time points and / or tonality compatibility information indicates that the first song and the second song are tonally compatible, the target cut-out time point is determined based on the vocal key time point of the first song, and the target cut-in time point is determined based on the vocal key time point of the second song. When the highest priority type combination is the combination of vocal key time point and measure key time point, the target cutout time point is determined in the first song based on the vocal key time point and measure key time point, and the target cutin time point is determined in the second song. When the highest priority type combination is another type combination, the target cut-out time point is determined from the candidate cut-out time points of the first song according to the preset determination order, and the target cut-in time point is determined from the candidate cut-in time points of the second song.

[0075] In some embodiments, when the highest priority type combination currently identified is a key performance time point, or when it is further determined by combining audio features that the number of beats per unit time (i.e., BPM) of the first song and the second song both exceed a preset number of beats per unit time threshold (e.g., 80 BPM), the target cutout time point is determined based on the key singing time point detected in the first song (e.g., the downbeat time point of a measure after the end of the singing), and the target cutin time point is determined based on the corresponding key performance time point in the second song (e.g., the downbeat time point of a measure before the start of the drumbeat).

[0076] In some embodiments, when the highest priority type combination currently identified is a structural key time point, and the tonality compatibility information indicates that the two songs are tonally compatible (such as belonging to C major, A minor, or relative key, near fifth key, etc.), the target cut-out time point is determined based on the vocal key time point detected in the first song (such as the downbeat time point of a measure after the end of the vocalization), and the target cut-in time point is determined based on the corresponding vocal key time point in the second song (such as the downbeat time point of a measure before the start of the vocalization).

[0077] In some embodiments, when the highest priority type combination currently identified is a combination of vocal key timing point and measure key timing point, specifically, if the first song has a vocal key timing point (e.g., the end of vocals) and the second song has a measure key timing point, then the target cutout point is determined based on the vocal key timing point (e.g., the downbeat of the measure after the end of vocals), and the target cutin point of the second song is determined based on the measure key timing point of the second song (e.g., the downbeat of the second measure from the beginning). Conversely, if the first song does not have a vocal key timing point (e.g., purely instrumental music), then its target cutout point is the measure key timing point of the first song (e.g., the downbeat of the second to last measure), and if the second song has a vocal key timing point, then the target cutin point is determined based on the vocal key timing point of the second song (e.g., the downbeat of the measure before the start of vocals).

[0078] In some embodiments, when the highest priority type combination currently identified is another type combination, target points are attempted to be selected from the candidate set in a preset order. This preset order can be: vocal key time point > performance key time point > structural key time point > measure key time point. The candidate cut-out time points of the first song are traversed, and the first existing time point corresponding to the preset key time point type is selected as the target cut-out time point in this preset order. Similarly, for the second song, the first existing time point corresponding to the preset key time point type is selected as the target entry time point.

[0079] It should be noted that the above descriptions of determining the target cut-out time point of the first song from candidate cut-out time points and the target cut-in time point of the second song from candidate cut-in time points are for illustrative purposes only and are not restrictive. The target cut-out time point can be a candidate cut-out time point, a time point near a candidate cut-out time point, or a stressed beat time point in a measure near a candidate cut-out time point, etc. Similarly, the target cut-in time point can be a candidate cut-in time point, a time point near a candidate cut-out time point, or a stressed beat time point in a measure near a candidate cut-in time point, etc.

[0080] It should be understood that there are other ways to determine the target cut-out time and the target cut-in time, and this application does not limit them.

[0081] In some embodiments, the transition processing parameters include transition time and song processing parameters. In response to the first song reaching the position corresponding to the target cut-out time point, based on transition processing parameters, the first music segment corresponding to the target cut-out time point of the first song and the second music segment corresponding to the target cut-in time point of the second song are processed and played, including: Based on the transition time, the target cut-out time of the first song, and the target cut-in time of the second song, the music segment of the transition time in the first song where the target cut-in time is located is taken as the first music segment, and the music segment of the transition time in the second song where the target cut-out time is located is taken as the second music segment. In response to the position corresponding to the target cut-out time point when the first song is played, the first music segment and the second music segment are processed and played in an overlapping manner based on the song processing parameters.

[0082] The transition time refers to the duration of overlapping playback between the first and second songs during the transition process, usually expressed in seconds or measures. Its value is determined based on the tempo information, tonality compatibility information, and energy difference information of the two songs. Song processing parameters may include volume control parameters, frequency control parameters, and rhythm synchronization parameters, which are used to regulate the audio signal processing during the transition period to optimize the continuity of the listening experience.

[0083] In an optional embodiment, when performing overlapping playback, two audio segments of equal length can be extracted based on the transition time, the target cut-out time of the first song, and the target cut-in time of the second song: the audio segment in the first song that starts from its target cut-out time and lasts for one transition time is designated as the first music segment; the audio segment in the second song that starts from its target cut-in time and lasts for the same transition time is designated as the second music segment. For example, if the target cut-out time is 190.5 seconds, the target cut-in time is 10.305 seconds, and the transition time is 3.0 seconds, then the first music segment corresponds to the interval from 190.5 seconds to 193.5 seconds in song A, and the second music segment corresponds to the interval from 10.305 seconds to 13.305 seconds in song B. These two music segments are time-aligned and will be processed and overlapped for playback in subsequent steps.

[0084] In some embodiments, the song processing parameters include volume control parameters for the first song and the second song. The volume control parameters are used to control the volume of the first music segment to decrease from the target cut-out time point and to control the volume of the second music segment to increase at least from the overlapping playback position.

[0085] In some embodiments, the song processing parameters include frequency control parameters for a first song and a second song, the frequency control parameters being used to process specified frequency components of the first music segment and / or the second music segment during a transition time.

[0086] In some embodiments, for the first musical segment (i.e., an audio segment that begins at the target cut-out time of the first song and lasts for a transition time), the volume control parameters include a function that controls the volume of the first musical segment to decrease over time, such that the first musical segment exhibits a decreasing volume trend from the target cut-out time and eventually decays to zero or near silence at the end of the overlap. Common curves include linear, exponential, logarithmic, or Gaussian functions. Correspondingly, for the second musical segment (i.e., an audio segment that begins at the target cut-in time of the second song and also lasts for a transition time), the volume control parameters include a function that increases the volume of the second musical segment over time, such that both musical segments exhibit an increasing volume trend at least from the beginning of the overlap (i.e., the target cut-in time) and gradually increase to their full amplitude. To ensure auditory balance, the curve of the second musical segment is mathematically symmetrical or complementary to the fade-out curve of the first musical segment to avoid fluctuations or dips in overall loudness during the overlap.

[0087] In some embodiments, the song processing parameters further include frequency control parameters for selectively processing specified frequency components in the first and / or second music segments during the transition time. These frequency control parameters may include filter type, cutoff frequency, and sweep path. For example, a high-pass filter (filtering out low frequencies) may be applied to the first song at the beginning of the transition, while a low-pass filter (filtering out high frequencies) may be applied to the second song. Then, during the transition time, the low-pass cutoff frequency of the first song may be gradually increased, and the high-pass cutoff frequency of the second song may be gradually decreased, ultimately restoring the full spectrum of both music segments.

[0088] In some embodiments, based on the audio feature information of the first song and the second song, the transition processing parameters for the first music segment corresponding to the target cut-out time point of the first song and the second music segment corresponding to the target cut-in time point of the second song are determined, including: Based on the audio feature information of the first song and the audio feature information of the second song, determine the second difference information between the first song and the second song; The transition time is determined based on the second difference information between the first and second songs; Based on the second difference information between the first and second songs, the target cutout time of the first song, and the target cutin time of the second song, the song processing parameters are determined.

[0089] In some embodiments, the second difference information includes at least one of the following: beat difference information between the first song and the second song, tonality compatibility information between the first song and the second song, and energy difference information between the first song and the second song.

[0090] In one alternative implementation, the transition time can be determined based on a machine learning model using second difference information between the first and second songs. Specifically, a transition duration prediction model can be pre-built, which takes the second difference information as input feature vectors and outputs a recommended transition time (in seconds or measures). The second difference information can include local rhythm difference information (ΔBPM_local), local tonality compatibility information (such as same key, closely related keys, conflicts, etc.), local loudness difference information (ΔLUFS_local), and multi-dimensional features such as music tags. This transition duration prediction model can be trained using supervised learning: the training dataset consists of a large number of manually annotated professional DJ mix sample pairs, each pair labeled with the actual transition duration used and the corresponding audio features of the two songs near key time points; the transition duration prediction model can employ a multilayer perceptron (MLP) or lightweight Transformer structure. During actual operation, the client inputs the second difference information calculated in real time into the transition duration prediction model, which then outputs the matched transition time.

[0091] The local rhythm difference information (ΔBPM_local) can be obtained by taking the BPM values ​​of local music segments with the target cut-out time point and the target cut-in time point as the center (or starting point) and cutting forward or backward for a time window (e.g., 2 to 4 seconds), and calculating the absolute value of the difference between the two.

[0092] Local tonality compatibility information is obtained by performing tonal analysis on the musical segments of the two local time windows using a tonality recognition model (such as Key-Transformer). Based on music theory rules (such as the Camelot roulette system), it is determined whether the two belong to the same key (such as both being in C major), related keys (such as C major and A minor), closely related keys (such as C major and G major or F major), or conflicting keys (such as C major and F# minor).

[0093] Local loudness difference information (ΔLUFS_local) can be obtained by calculating the loudness (in LUFS) of two musical segments within the same time window and taking the absolute value of the difference. This indicator reflects the energy level difference between the two songs at the actual transition point. If ΔLUFS_local is too large (e.g., exceeding 6 LUFS), direct superposition will cause a volume jump.

[0094] Music tags are a set of semantic features, including but not limited to genre (such as electronic, rock, jazz, etc.), mood (such as cheerful, soothing, exciting, etc.), and instrumental composition (such as whether it contains drums, vocals, strings, etc.). Music tags can be output by pre-trained music classification models (such as VGGish or OpenL3) to help determine whether two songs are compatible in style or atmosphere.

[0095] In another alternative implementation, the song processing parameters can also be determined using a machine learning-assisted model. For example, a parameter generation network can be constructed, whose inputs include second difference information, contextual features at the target cutout time point of the first song (such as audio feature information near the target cutout time point, the type of the target cutout time point, etc.), and contextual features at the target cutin time point of the second song (such as audio feature information near the target cutin time point, the type of the target cutin time point, etc.). This network can output structured song processing parameters, including the volume control curve type (such as Gaussian, exponential, S-shaped) and its parameter values ​​(such as μ, σ), whether a filter is enabled, the filter type (high-pass / low-pass / band-pass, etc.), the sweep start and end frequencies and rate of change, and the tempo ratio, etc. This parameter generation network can be trained offline using reinforcement learning: using professional DJ mix samples as expert strategies, the parameters of the parameter generation network are optimized by minimizing the difference between the generated audio and the reference mix in perceptual quality (such as using DNSMOS or PESQ metrics) and tempo alignment error. When deployed online, this parameter generation network can automatically generate appropriate song processing parameters based on the contextual features at the junction of the two songs.

[0096] In some embodiments, the transition time is negatively correlated with the second difference information.

[0097] In some embodiments, the song processing parameters further include rhythm synchronization parameters for adjusting the number of beats per unit time for the first and second musical segments, and the method further includes; When the beat difference information between the first song and the second song does not exceed the first preset difference threshold, the rhythm synchronization parameter is determined to be the first preset value. The first preset value is used to indicate that the number of beats per unit time of the first music segment and the second music segment will not be adjusted. When the beat difference between the first song and the second song exceeds the first preset difference threshold, the rhythm synchronization parameter is determined to be the second preset value, so as to adjust the number of beats per unit time of the first music segment and the second music segment based on the second preset value.

[0098] For example, the transition time is not fixed but negatively correlated with the second difference information. That is, the greater the difference in local audio features between the first and second songs near the target cut-out and target cut-in points, the shorter the determined transition time; conversely, the smaller the difference in local audio features near the target cut-out and target cut-in points, the longer the determined transition time. In certain scenarios (such as tonality dissonance, drastic rhythmic jumps, or sudden energy changes), longer overlapping playback amplifies dissonance, leading to auditory discomfort; therefore, a shorter transition time is needed to quickly complete the switch. In other scenarios, a longer transition time can create a smooth, immersive blending effect, improving auditory coherence. For instance, if the second difference information shows ΔBPM_local as 6, tonality conflict, and ΔLUFS_local as 8 LUFS, the transition time can be set to 1 beat.

[0099] In some embodiments, the song processing parameters further include a rhythm synchronization parameter to control whether and how the beats per unit time (BPM) of the first song and / or the second song are fine-tuned to achieve beat alignment. The determination of this rhythm synchronization parameter depends on the calculation of beat difference information between the first and second songs. For example, firstly, the music segments (e.g., the first music segment, the second music segment, etc.) of the two songs are calculated, taking a time window forward or backward from the target cut-out time point and the target cut-in time point as the center (or starting point). For each music segment, its local BPM value (BPM1_local and BPM2_local) is calculated, and their absolute difference is taken as the local beat difference information. If the local beat difference information does not exceed a first preset difference threshold (e.g., 2 BPM), it can be considered that the rhythms are basically consistent. In this case, the rhythm synchronization parameter is set to a first preset value (e.g., 1.0 or a disabled flag), indicating that no speed adjustment is performed on either song, and they are played overlapping at the original rate, thereby avoiding unnecessary sound quality loss.

[0100] When the beat difference exceeds the first preset difference threshold, it is considered that there is a risk of rhythm misalignment. At this time, the rhythm synchronization parameter is set to a second preset value, specifically a non-unit speed ratio (such as 0.9697 or 1.032), which is used to trigger the time stretching algorithm to adjust the speed of one or two music segments in real time. For example, the speed ratio (such as 0.9697 or 1.032) can be applied to the first song or the second song, for example: If the tempo of the second song is adjusted to match that of the first song, then the speed ratio = BPM1 / BPM2.

[0101] For example, if the local BPM (denoted as BPM1_local) of the first song near the target cut-out point is 124, and the local BPM (BPM2_local) of the second song near the target cut-in point is 128, then the speed change ratio = 124 ÷ 128 ≈ 0.9697. This 0.9697 is applied to the second song, causing it to play at approximately 96.97% of its original speed during the transition, thereby reducing its effective BPM from 128 to 124 to align with the first song.

[0102] If the tempo of the first song is adjusted to match that of the second song, then the speed ratio = BPM2 / BPM1.

[0103] In the example above, the speed ratio = 128 ÷ 124 ≈ 1.0323. This ratio is applied to the first song, causing it to speed up slightly during the transition (approximately 103.23% speed), increasing its effective BPM from 124 to 128, aligning it with the second song.

[0104] In practice, the first song (i.e. the song that is about to end) can be given priority for speed adjustment. At this time, the first song is nearing its end, usually in the outro, fade-out or rhythmic climax stage, and its musical information density is low, so it has a higher tolerance for speed adjustment.

[0105] It should be noted that when the tempo of the first or second song increases, the duration of the audio signal in the first or second musical segment is shortened. Due to this time compression, the duration of each note, melody, and rhythmic element in the audio signal of the first or second musical segment will also be shortened accordingly.

[0106] Similarly, when the tempo of the first or second song slows down, the duration of the audio signal in the first or second musical segment is lengthened. Because of this time stretching, the duration of each note, melody, and rhythmic element in the audio signal of the first or second musical segment will also be extended accordingly.

[0107] In one alternative approach, when the tempo (also known as the number of beats per unit time) of the first musical segment during the transition time of the first song is increased, the musical segment that would normally take longer to play can be completed in a shorter time due to time compression. This means that after increasing the tempo, if the transition time is fixed, the first musical segment of the first song may finish playing earlier within the remaining transition time. In this case, the first song can be omitted during the remaining transition time, or other content, such as transitional sound effects, can be played instead.

[0108] In an alternative approach, instead of slowing down the tempo of the first musical segment during the transition time of the first song, the originally shorter segment requires a longer playback time due to time stretching. In this case, if the transition time is fixed, the entire musical segment cannot be played within the fixed transition time. Therefore, only a portion of the first musical segment is played during the transition time.

[0109] To facilitate understanding, the following will describe in detail the specific execution process of the song playback method provided in this embodiment, using concrete examples.

[0110] Please see Figure 6 ,like Figure 6 As shown, this application embodiment can be divided into three stages: the server-side extraction stage, the client-side decision-making stage, and the player execution stage. The server-side extraction stage includes an audio feature extraction module and a structure extraction module. By extracting audio feature information such as song tempo, key, and energy, as well as structural feature information, it provides data support for subsequent connection point decisions and connection strategy decisions. The client-side decision-making stage, based on the audio feature information and structural feature information of the two songs to be connected, first determines the optimal connection time point (e.g., ...). Figure 6 The decision-making process for transition points (such as the target cutout time point and target cutin time point mentioned above) is combined with the transition time point and the audio feature information of the song to intelligently determine the best transition strategy (such as the transition processing parameters mentioned above). Figure 6 (The transition strategy decision is shown). Finally, the transition point and transition strategy are synchronized to the player, and the player performs the transition processing to achieve seamless DJ-style playback with a continuous and natural listening experience without manual intervention.

[0111] The audio feature extraction module generates a standardized audio feature library for each song, serving as data for all subsequent connection decisions and processing. The extracted feature information includes, but is not limited to, rhythm features, tonality features, energy and spectral features, and other music tag information. Rhythm features are extracted using audio analysis algorithms to obtain BPM (beats per minute), time signature (e.g., 4 / 4 time), and beat timestamps (the specific time position of each beat). In a specific embodiment, methods such as BeatNet can be used: the pre-processed audio is input into a trained CRNN model, which outputs beat and downbeat activation signals. These signals are then used for particle filtering and tracking inference to lock the beat and downbeat positions in real time. For example, a song with a BPM of 120 and a time signature of 4 / 4 contains 120 beats per minute (each beat is 0.5 seconds apart), with each measure consisting of 4 beats, and the first beat being an downbeat. Tonal characteristics are used to identify the tonic key of a song and determine the tonal relationship between two songs based on a universal Camelot wheel. This includes homophony (e.g., C major and C major), relative keys (e.g., C major and A minor), closely related keys (difference of one key signature, e.g., C major and G major or F major), moderately related keys (difference of two key signatures), and distantly related keys (difference of three or more key signatures), to assess tonal compatibility when connecting songs. Energy characteristics include a full-band loudness sequence (in LUFS) and a loudness sequence showing the energy change over time, divided into low, mid, and high frequencies. Other information includes genre labels, instrument types, etc. Genre labels can include electronic, R&B, slow DJ, light music, hip-hop, children's music, traditional Chinese music, folk, jazz, rock, etc., and instrument labels can include time-series labels for instruments such as vocals, drums, guitar, and piano.

[0112] The music structure extraction module extracts structural time points from music from different dimensions. The structural feature information that this module can extract includes, but is not limited to, vocal entry / exit points, drum entry / exit points, and music segment structure points. Vocal entry / exit points can be calculated by extracting vocal tracks using a vocal track separation algorithm (such as Demucs) and then calculating the start and end times of each vocal track. Similarly, for drum entry / exit points, the start and end times of each drum sound are determined based on the drum track separation results. Music segment structure points include the start and end times of the intro, verse, chorus, bridge, and outro. In a specific embodiment, the open-source tool All-In-One can be used for analysis. All-In-One is based on a deep learning model, trained on a large amount of music data, and can accurately identify various segments in the audio. After inputting the main audio track or other combined audio tracks into All-In-One, it outputs the start and end times of each segment and the segment label.

[0113] Please see Figure 3 ,like Figure 3 As shown, the waveform above is the waveform of the original song's audio signal, and the waveform below is the waveform of the audio signal of the human voice track after processing by the human voice track separation algorithm.

[0114] Please see Figure 4 ,like Figure 4 As shown, the waveform above is the waveform of the original song's audio signal, and the waveform below is the waveform of the drum track's audio signal processed by the drum track separation algorithm.

[0115] Please see Figure 5 ,like Figure 5 As shown, the musical structure of a song includes various musical sections such as intro, verse, chorus, bridge, and outro.

[0116] After receiving the music structure information of the first song (Song A) and the second song (Song B) from the server on the client side, the decision module further processes the structural feature information and searches for four preset key time point types. The preset key time point types are predefined types of key time points. If no matching preset key time point type is found, it is considered that the preset key time point type does not exist in the song, as shown in Table 1: Table 1

[0117] Based on the aforementioned preset key time point types, candidate cutout time points for the first song and candidate cutin time points for the second song are searched, and the final target cutout time point and target cutin time point are selected according to the following priority: The first priority is the combination of drum key timing points: when both song A and song B have effective drum key timing points and both have a BPM greater than 80, the target cut-out point for song A is set to the downbeat of a measure after the vocals end, and the target cut-in point for song B is set to the drum start time, in order to maintain rhythmic continuity.

[0118] The second priority is the combination of structural key time points: when both songs A and B have effective structural key time points, and their tones are the same, related, or closely related, the target cut-out point for song A is set to the downbeat of a measure after the vocals end, and the target cut-in point for song B is set to the downbeat of a measure before the vocals begin.

[0119] The third priority is the combination of vocal key moments and measure key moments: when song A has vocal key moments, its target cut-out point is the downbeat of a measure one measure after the vocal end time, and the target cut-in point for song B is its measure key moment; when song A has no vocal key moments, its target cut-out point is its measure key moment, and the target cut-in point for song B is the downbeat of a measure one measure before the vocal start time.

[0120] The fourth priority is other types of combinations: In this case, for song A, the key time points are selected in the following order: one measure after the key time point of the vocals, one measure after the key time point of the drums, the key time point of the structure, and the key time point of the measure; for song B, the key time points are selected in the following order: one measure before the key time point of the vocals, one measure before the key time point of the drums, the key time point of the structure, and the key time point of the measure.

[0121] The decision-making module, based on the transition time and audio feature information of the two songs, decides whether to enable speed adjustment, determine the transition time, and the transition method. The decision-making module calculates the following features: Speed ​​difference: ΔBPM = |BPM_A BPM_B|; Tonal compatibility: Based on the circle of fifths, tonality compatibility includes consonant, related, and closely related keys; all others are conflicting. Volume difference: ΔLUFS=|LUFS_A LUFS_B|, where LUFS_A is the local loudness near the exit point of song A, and LUFS_B is the local loudness near the entry point of song B.

[0122] The transition time is determined according to the following preset rules: When ΔBPM≤±2, tonality is compatible, and ΔLUFS≤3, the transition time is determined to be 4 measures. Since the song conflicts are minimal, a longer transition time can be supported. When ΔBPM≤±5 or 3<ΔLUFS≤6 and the tone is compatible, the transition time is set to 3 measures to balance naturalness and smoothness; When there is a tone conflict, or ΔLUFS>6, or ΔBPM>±5, the transition time is set to 2 measures to reserve a buffer space to cover the differences; When a tone conflict, ΔBPM>±5 and ΔLUFS>6 are simultaneously met, the transition time is set to one measure to ensure a quick transition and avoid deterioration in sound quality.

[0123] Transition methods include volume control curves, filters, speed changes, and beat matching.

[0124] In some embodiments, the shot-matching is used as a basic operation, and by default the target cut-out point and the target cut-in point are aligned with the retake time points of their respective measures.

[0125] In some embodiments, variable speed operation is enabled only when ΔBPM>±2 to avoid unnecessary sound quality degradation.

[0126] In some embodiments, the specific transition method is as follows: When the style and tone are the same, ΔBPM≤±2, and ΔLUFS≤3, the S-curve is used and the filter is not enabled. Because the conflict is minimal, the S-curve conforms to the characteristics of human hearing. When the energy difference is large (ΔLUFS>3), an exponential curve is used, the filter is not enabled, and the dynamic balance of the curve itself is used to avoid sudden changes in volume. When there are large rhythm differences (ΔBPM>±3), a logarithmic curve is used, and a high-pass filter is applied to song A to gradually filter out low frequencies. Because the logarithmic curve changes rapidly in the early stage, it helps to quickly align the rhythm, and the high-pass filter can avoid low-frequency muddiness. When there is a conflict in tonality or a large difference in style (such as a rock song followed by a classical song), a Gaussian curve is used and frequency sweeping is performed: a low-pass frequency sweep is applied to song A and a high-pass frequency sweep is applied to song B to avoid tonality conflict. When transitioning between instrumental or ambient music, a Gaussian curve is used with a 1 / 4 beat feedback delay to enhance the delicate blending and preserve the song's ending notes. When combining key drum beats at crucial moments, a linear curve is used and drum beats are layered to enhance rhythmic continuity. The linear curve does not mask the impact of the drum beats.

[0127] The transition playback module is executed by the player, which starts transition playback after receiving the transition parameters output by the decision module.

[0128] For example, the parameters include: the target cut-out time for song A is 190.5 seconds, the target cut-in time for song B is 10.305 seconds, the transition time is 3.0 seconds, the transition method is a Gaussian curve combined with frequency sweep (200Hz and 20000Hz), the rhythm synchronization parameter is a speed ratio of 0.9697, and the volume control parameters are a fade-in curve with μ=0.5 and σ=0.2, and the fade-out curve is similar. During playback, when song A plays to 190.5 seconds, song B starts playing from its 10.305-second mark. Song A undergoes real-time speed adjustment according to the speed ratio of 0.9697, while the volume control and frequency sweep filtering of the Gaussian curves superimposed on songs A and B are applied to achieve a smooth and natural song transition.

[0129] Please refer to Figure 7 ,like Figure 7 As shown, the Gaussian curve is truncated in the left half during fade-in processing, and the volume of the first song is gradually increased from 0 to 1 using a function in the left half; in the fade-out processing, the right half is truncated, and the volume of the second song is gradually decreased from 1 to 0 using a function in the right half. The parameter μ (mean) controls the position of the fastest volume change and is set to 0.5; σ (standard deviation) controls the steepness of the curve, i.e., the rate of volume change, and is set to 0.2. Frequency sweep refers to dynamically adjusting the cutoff frequency of the filter within a continuous frequency range, for example, increasing the cutoff frequency from 200Hz to 20000Hz: if it is a high-pass sweep, the high-frequency components are gradually released, making the sound brighter; if it is a low-pass sweep, the low-frequency components are gradually released, making the sound thicker. This mechanism effectively masks potential conflicts in the transition process through spectral crossover.

[0130] This embodiment also provides a song playback device, which can be integrated into a terminal device. For example, such as... Figure 8 As shown, the song playback device may include: The acquisition module 301 is used to acquire feature information of the first song and the second song. The feature information includes audio feature information and structural feature information. The second song is the song played after the first song. The first determining module 302 is used to determine the candidate cut-out time point of the first song based on the structural feature information of the first song, and to determine the candidate cut-in time point of the second song based on the structural feature information of the second song. The second determining module 303 is used to determine the target cut-out time point of the first song from the candidate cut-out time points and the target cut-in time point of the second song from the candidate cut-in time points based on the first difference information of the audio feature information of the first song and the second song. The third determining module 304 is used to determine the transition processing parameters of the first music segment corresponding to the target cut-out time point of the first song and the second music segment corresponding to the target cut-in time point of the second song based on the audio feature information of the first song and the second song. The playback processing module 305 is used to process and play the first music segment corresponding to the target cut-out time point of the first song and the second music segment corresponding to the target cut-in time point of the second song based on the transition processing parameters, in response to the first song playing to the position corresponding to the target cut-out time point.

[0131] The acquisition module 301, the first determination module 302, the second determination module 303, the third determination module 304, and the playback processing module 305 can be used to execute the embodiments corresponding to the above-mentioned song playback method. For the specific implementation methods of these modules and more details, please refer to the corresponding method section, which will not be elaborated here.

[0132] By comprehensively analyzing the audio and structural features of the first and second songs, the song playback device provided in this application can identify the target cut-out and target cut-in time points that conform to the song features and listening habits. It can also dynamically determine the transition processing parameters by combining the audio feature information of the first and second songs, thereby achieving a smooth, natural and seamless connection effect between the two songs that conforms to professional mixing logic.

[0133] Accordingly, this application also provides an electronic device, which can be a terminal, such as a smartphone, tablet computer, laptop computer, touch screen, game console, personal computer (PC), personal digital assistant (PDA), or other terminal device. Alternatively, the electronic device can be a server.

[0134] like Figure 9 As shown, Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device 400 includes a processor 401 with one or more processing cores, a memory 402 with one or more computer-readable storage media, and a computer program stored in the memory 402 and executable on the processor. The processor 401 and the memory 402 are electrically connected. Those skilled in the art will understand that the electronic device structure shown in the figure does not constitute a limitation on the electronic device, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0135] The processor 401 is the control center of the electronic device 400. It connects various parts of the electronic device 400 via various interfaces and lines. By running or loading software programs and / or units stored in the memory 402, and by calling data stored in the memory 402, it executes various functions and processes data of the electronic device 400, thereby providing overall monitoring of the electronic device 400. The processor 401 can be a central processing unit (CPU), a graphics processing unit (GPU), a network processor (NP), etc., and can implement or execute the methods, steps, and logic diagrams disclosed in the embodiments of this application.

[0136] In this embodiment, the processor 401 in the electronic device 400 loads the instructions corresponding to the processes of one or more applications into the memory 402 according to the following steps, and the processor 401 runs the applications stored in the memory 402 to realize various functions, such as: Obtain feature information of the first song and the second song. The feature information includes audio feature information and structural feature information. The second song is the song played after the first song. Based on the structural feature information of the first song, the candidate cut-out time point of the first song is determined, and based on the structural feature information of the second song, the candidate cut-in time point of the second song is determined. Based on the first difference information of the audio feature information of the first song and the second song, the target cut-out time point of the first song is determined from the candidate cut-out time point, and the target cut-in time point of the second song is determined from the candidate cut-in time point. Based on the audio feature information of the first song and the second song, the transition processing parameters of the first music segment corresponding to the target cut-out time point of the first song and the second music segment corresponding to the target cut-in time point of the second song are determined. In response to the first song reaching the position corresponding to the target cut-out time point, based on the transition processing parameters, the first music segment corresponding to the target cut-out time point of the first song and the second music segment corresponding to the target cut-in time point of the second song are processed and played.

[0137] By using the electronic device provided in this application embodiment, through comprehensive analysis of the audio feature information and structural feature information of the first song and the second song, the target cut-out time point and the target cut-in time point that conform to the song features and listening habits can be identified. The transition processing parameters are dynamically determined by combining the audio feature information of the first song and the second song, thereby achieving a smooth, natural and seamless connection effect between the two songs that conforms to professional mixing logic.

[0138] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0139] Optional, such as Figure 9 As shown, the electronic device 400 also includes: a touch display screen 403, a radio frequency circuit 404, an audio circuit 405, an input unit 406, and a power supply 407. The processor 401 is electrically connected to the touch display screen 403, the radio frequency circuit 404, the audio circuit 405, the input unit 406, and the power supply 407. Those skilled in the art will understand that... Figure 9 The electronic device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0140] The touch display screen 403 can be used to display a graphical user interface (GUI) and receive operation commands generated by the user interacting with the GUI. The touch display screen 403 may include a display panel and a touch panel. The display panel can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces of the electronic device. These graphical user interfaces can be composed of graphics, text, icons, video, and any combination thereof. Optionally, the display panel can be configured using a liquid crystal display (LCD), organic light-emitting diode (OLED), or other similar technologies. The touch panel can be used to collect touch operations performed by the user on or near it (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or near the touch panel), generate corresponding operation commands, and execute the corresponding program according to the operation commands. Optionally, the touch panel may include two parts: a touch detection device and a touch controller. The touch detection device detects the user's touch location and the signal generated by the touch operation, transmitting the signal to the touch controller. The touch controller receives touch information from the touch detection device, converts it into touch point coordinates, and sends it to the processor 401. It can also receive and execute commands from the processor 401. The touch panel can cover the display panel. When the touch panel detects a touch operation on or near it, it transmits the information to the processor 401 to determine the type of touch event. Subsequently, the processor 401 provides corresponding visual output on the display panel based on the type of touch event. In this embodiment, the touch panel and the display panel can be integrated into the touch display screen 403 to achieve input and output functions. However, in some embodiments, the touch panel and the touch display screen 403 can be implemented as two independent components to achieve input and output functions. That is, the touch display screen 403 can also be used as part of the input unit 406 to achieve input functions.

[0141] The radio frequency circuit 404 can be used to transmit and receive radio frequency signals to establish wireless communication with network devices or other electronic devices, and to transmit and receive signals with network devices or other electronic devices.

[0142] Audio circuit 405 can be used to provide an audio interface between a user and an electronic device via a speaker and a microphone. Audio circuit 405 can convert received audio data into electrical signals and transmit them to the speaker, where the speaker converts them into sound signals for output. Conversely, the microphone converts collected sound signals into electrical signals, which are then received by audio circuit 405, converted back into audio data, and then processed by processor 401 before being transmitted via radio frequency circuit 404 to, for example, another electronic device, or output to memory 402 for further processing. Audio circuit 405 may also include an earphone jack to provide communication between peripheral headphones and electronic devices.

[0143] The input unit 406 can be used to receive source audio, reference audio, etc.

[0144] Power supply 407 is used to supply power to various components of electronic device 400. Optionally, power supply 407 can be logically connected to processor 401 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. Power supply 407 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0145] although Figure 9 As not shown in the diagram, the electronic device 400 may also include a camera, sensor, wireless fidelity module, Bluetooth module, etc., which will not be described in detail here.

[0146] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0147] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0148] Therefore, embodiments of this application provide a computer-readable storage medium storing multiple computer programs that can be loaded by a processor to execute any of the song playback methods provided in this application. The computer program can execute the steps of the following song playback method: Obtain feature information of the first song and the second song. The feature information includes audio feature information and structural feature information. The second song is the song played after the first song. Based on the structural feature information of the first song, the candidate cut-out time point of the first song is determined, and based on the structural feature information of the second song, the candidate cut-in time point of the second song is determined. Based on the first difference information of the audio feature information of the first song and the second song, the target cut-out time point of the first song is determined from the candidate cut-out time point, and the target cut-in time point of the second song is determined from the candidate cut-in time point. Based on the audio feature information of the first song and the second song, the transition processing parameters of the first music segment corresponding to the target cut-out time point of the first song and the second music segment corresponding to the target cut-in time point of the second song are determined. In response to the first song reaching the position corresponding to the target cut-out time point, based on the transition processing parameters, the first music segment corresponding to the target cut-out time point of the first song and the second music segment corresponding to the target cut-in time point of the second song are processed and played.

[0149] By using the computer-readable storage medium provided in the embodiments of this application, and by comprehensively analyzing the audio feature information and structural feature information of the first song and the second song, it is possible to identify the target cut-out time point and the target cut-in time point that conform to the song features and listening habits, and dynamically determine the transition processing parameters by combining the audio feature information of the first song and the second song, thereby achieving a smooth, natural and seamless connection effect between the two songs that conforms to professional mixing logic.

[0150] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0151] The computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0152] Since the computer program stored in the computer-readable storage medium can execute any of the song playback methods provided in the embodiments of this application, it can achieve the beneficial effects that any of the song playback methods provided in the embodiments of this application can achieve, as detailed in the preceding embodiments, and will not be repeated here.

[0153] According to one aspect of this application, a computer program product or computer program is also provided, comprising computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the electronic device to perform the methods provided in the various optional implementations of the above embodiments.

[0154] In the above embodiments of the song playback device, computer-readable storage medium, electronic device, and computer program product, the descriptions of each embodiment have different focuses. Parts not described in detail in a particular embodiment can be referred to in the relevant descriptions of other embodiments. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes and beneficial effects of the song playback device, computer-readable storage medium, computer program product, electronic device, and their corresponding units described above can be referred to the description of the song playback method in the above embodiments, and will not be repeated here.

[0155] The foregoing has provided a detailed description of a song playback method, apparatus, electronic device, computer-readable storage medium, and computer program product provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for playing songs, characterized in that, The method includes: Obtain feature information of a first song and a second song, the feature information including audio feature information and structural feature information, wherein the second song is a song played after the first song; Based on the structural feature information of the first song, the candidate cut-out time point of the first song is determined, and based on the structural feature information of the second song, the candidate cut-in time point of the second song is determined. Based on the first difference information of the audio feature information of the first song and the second song, the target cut-out time point of the first song is determined from the candidate cut-out time point, and the target cut-in time point of the second song is determined from the candidate cut-in time point; Based on the audio feature information of the first song and the second song, the transition processing parameters of the first music segment corresponding to the target cut-out time point of the first song and the second music segment corresponding to the target cut-in time point of the second song are determined. In response to the first song playing to the position corresponding to the target cut-out time point, based on the transition processing parameters, the first music segment corresponding to the target cut-out time point of the first song and the second music segment corresponding to the target cut-in time point of the second song are processed and played.

2. The song playback method according to claim 1, characterized in that, The step of determining the candidate cut-out time point of the first song based on the structural feature information of the first song, and determining the candidate cut-in time point of the second song based on the structural feature information of the second song, includes: Based on the structural feature information of the first song and multiple preset key time point types, the first key time point under each preset key time point type is searched in the first song as a candidate cutout time point; Based on the structural feature information of the second song and multiple preset key time point types, the second key time point under each preset key time point type is searched in the second song as a candidate entry time point.

3. The song playback method according to claim 2, characterized in that, The structural feature information includes at least one of the following: the start and end times of the singing, the start and end times of the performance of a specified instrument, and the structural information of a musical segment; the structural information of the musical segment includes the type of at least one musical segment and the start and end times of each musical segment. The preset key time point types include: key time points for singing, key time points for the performance of the specified instrument, key time points for structure, and key time points for measures. The key timing points of the singing are determined based on the start and end times of the singing. The key performance points of the specified instrument are determined based on the start and end times of the performance. The key time points of the structure are determined based on the start and end time information of a specified music segment; The key time points of each segment are determined based on preset retake time points.

4. The song playback method according to claim 3, characterized in that, The preset key time point type combinations of the candidate cutout time points and the candidate cut-in time points have corresponding priorities, and the priority order from high to low is: all are performance key time points, all are structural key time points, a combination of vocal key time points and measure key time points, and other type combinations. The step of determining the target cut-out time point of the first song from the candidate cut-out time points and the target cut-in time point of the second song from the candidate cut-in time points based on the first difference information of the audio feature information of the first song and the second song includes: Based on the preset key time point type to which the candidate cutout time point belongs and the preset key time point type to which the candidate cutin time point belongs, determine at least one type combination composed of the candidate cutout time point and the candidate cutin time point. Based on the highest priority type combination and the first difference information of the audio feature information of the first song and the second song, the target cut-out time point of the first song and the target cut-in time point of the second song are determined.

5. The song playback method according to any one of claims 1 to 4, characterized in that, The transition processing parameters include transition time and song processing parameters. In response to the first song playing to the position corresponding to the target cut-out time point, based on the transition processing parameters, the first music segment corresponding to the target cut-out time point of the first song and the second music segment corresponding to the target cut-in time point of the second song are processed and played, including: Based on the transition time, the target cut-out time of the first song, and the target cut-in time of the second song, the music segment of the transition time in the first song where the target cut-in time is located is taken as the first music segment, and the music segment of the transition time in the second song where the target cut-out time is located is taken as the second music segment. In response to the first song playing to the position corresponding to the target cut-out time point, the first music segment and the second music segment are processed and played in an overlapping manner based on the song processing parameters.

6. The song playback method according to claim 5, characterized in that, The step of determining the transition processing parameters for the first music segment corresponding to the target cut-out time point of the first song and the second music segment corresponding to the target cut-in time point of the second song based on the audio feature information of the first song and the second song includes: Based on the audio feature information of the first song and the audio feature information of the second song, a second difference information between the first song and the second song is determined. The transition time is determined based on the second difference information between the first song and the second song; Based on the second difference information between the first song and the second song, the target cut-out time point of the first song, and the target cut-in time point of the second song, the song processing parameters are determined.

7. The song playback method according to claim 5, characterized in that, The song processing parameters also include rhythm synchronization parameters for adjusting the number of beats per unit time for the first music segment and the second music segment, and the method further includes; When the beat difference information between the first song and the second song does not exceed the first preset difference threshold, the rhythm synchronization parameter is determined to be the first preset value. The first preset value is used to indicate that the number of beats per unit time of the first music segment and the second music segment is not adjusted. When the beat difference information between the first song and the second song exceeds the first preset difference threshold, the rhythm synchronization parameter is determined to be a second preset value, so as to adjust the number of beats per unit time of the first music segment and the second music segment based on the second preset value.

8. A song playback device, characterized in that, The device includes: The acquisition module is used to acquire feature information of a first song and a second song. The feature information includes audio feature information and structural feature information. The second song is a song played after the first song. The first determining module is used to determine the candidate cut-out time point of the first song based on the structural feature information of the first song, and to determine the candidate cut-in time point of the second song based on the structural feature information of the second song. The second determining module is used to determine the target cut-out time point of the first song from the candidate cut-out time points based on the first difference information of the audio feature information of the first song and the second song, and to determine the target cut-in time point of the second song from the candidate cut-in time points. The third determining module is used to determine the transition processing parameters of the first music segment corresponding to the target cut-out time point of the first song and the second music segment corresponding to the target cut-in time point of the second song based on the audio feature information of the first song and the second song. The playback processing module is used to respond to the first song playing to the position corresponding to the target cut-out time point, and to process and play the first music segment corresponding to the target cut-out time point of the first song and the second music segment corresponding to the target cut-in time point of the second song based on the transition processing parameters.

9. An electronic device, characterized in that, It includes a processor and a memory, wherein the memory stores a computer program that, when executed by the processor, causes the processor to perform the steps of the song playback method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, It includes a computer program that, when run on an electronic device, causes the electronic device to perform the steps of the song playback method according to any one of claims 1 to 7.