Song transition method and device, electronic equipment, storage medium and program product

By acquiring the audio features of songs and selecting appropriate transition strategies, audio control parameters are generated, solving the problem of poor song transition effects and achieving smooth and natural transitions between songs, thus improving the listening experience.

CN122412643APending Publication Date: 2026-07-17HANGZHOU NETEASE CLOUD MUSIC TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HANGZHOU NETEASE CLOUD MUSIC TECH CO LTD
Filing Date
2026-04-13
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing song transition methods lack flexibility and adaptability when implementing continuous song playback, resulting in abrupt transition effects that cannot dynamically adjust according to user operations and song characteristics, thus affecting the listening experience.

Method used

By acquiring the audio features of a song, a transition strategy is dynamically selected and audio control parameters are generated to achieve a smooth transition between songs. This includes strategies such as rhythm alignment and avoiding vocal overlap, and a detailed set of audio control instructions is generated to ensure the naturalness of the transition effect.

Benefits of technology

It significantly improves the smoothness and naturalness of song transitions, providing a smoother and more natural continuous music playback experience, and can respond to users' real-time adjustment needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122412643A_ABST
    Figure CN122412643A_ABST
Patent Text Reader

Abstract

This application discloses a song transition method, apparatus, electronic device, storage medium, and program product, relating to the field of multimedia technology. The method includes: acquiring a first audio feature of a currently playing first song and a second audio feature of a second song to be played; determining a target song transition strategy from a variety of preset song transition strategies based on the first and second audio features; generating audio control parameters for transitioning from the first song to the second song based on the target song transition strategy; and transitioning from the first song to the second song under the audio control parameters. By implementing the technical solution of this application, the optimal transition strategy can be intelligently selected based on the song's audio features, achieving a smooth and adaptable playback transition, thereby improving the continuity of music playback and enhancing user immersion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of multimedia technology, specifically to song transition methods, devices, electronic devices, storage media, and program products. Background Technology

[0002] In the field of music playback, especially in applications such as music streaming services and smart terminal players, achieving smooth and natural transitions between songs is one of the key technologies for improving user experience. Currently, while common song transition methods achieve continuous playback to some extent, they still face many challenges in practical applications. For example, pre-built fixed mixing schemes lack the flexibility to dynamically adjust based on real-time user actions and the characteristics of different songs; while the widely adopted simple audio crossfade strategy often neglects the rhythm and structure of the song itself, resulting in abrupt transitions and a discordant listening experience. Summary of the Invention

[0003] In view of this, this application provides a song transition method, apparatus, electronic device, storage medium, and program product to solve the problem of poor song transition effect.

[0004] In a first aspect, this application provides a song transition method, comprising: obtaining a first audio feature of a currently playing first song and a second audio feature of a second song to be played; determining a target song transition strategy from a variety of preset song transition strategies based on the first audio feature and the second audio feature; generating audio control parameters for transitioning from the first song to the second song based on the target song transition strategy; and transitioning from the first song to the second song under the audio control parameters.

[0005] In some optional implementations, a target song transition strategy is determined from a variety of preset song transition strategies based on a first audio feature and a second audio feature, including: obtaining the transition priority corresponding to each preset song transition strategy; performing feature matching between the first audio feature and the second audio feature and each preset song transition strategy based on the priority order represented by each transition priority; and determining the first preset song transition strategy that successfully matches the feature as the target song transition strategy.

[0006] In some optional implementations, audio control parameters for transitioning from the first song to the second song are generated based on the target song transition strategy. This includes: determining an initial transition time point based on the target song transition strategy, a first audio feature, and a second audio feature; acquiring the first beat information of the first song and the second beat information of the second song; searching for a first candidate beat point corresponding to the first song based on the first beat information, and searching for a second candidate beat point corresponding to the second song based on the second beat information, centered on the initial transition time point; determining a first time deviation between the first candidate beat point and the initial transition time point, and a second time deviation between the second candidate beat point and the initial transition time point; and determining a target beat point with the smallest time deviation and meeting a preset tolerance condition from the first and second candidate beat points based on the first and second time deviations, and defining the target beat point as an alignment anchor point. The audio control parameters include the alignment anchor point, and the preset tolerance condition means that the minimum time deviation does not exceed a preset tolerance value.

[0007] In some optional implementations, the first candidate heavy beat point corresponding to the first song is obtained by searching based on the first beat information, centered on the initial transition time point. This includes: searching for heavy beat points in a first direction where the playback time decreases, centered on the initial transition time point; if there is a heavy beat point in the first direction that meets the preset strong beat condition, then the heavy beat point in the first direction that meets the preset strong beat condition is determined as the first candidate heavy beat point; if there is no heavy beat point in the first direction that meets the preset strong beat condition, then the heavy beat point is searched in a second direction where the playback time increases, and the heavy beat point in the second direction that meets the preset strong beat condition is determined as the first candidate heavy beat point; wherein, the preset strong beat condition means that the heavy beat point is marked as a strong beat in the beat information of the corresponding song.

[0008] In some optional implementations, the second candidate beat point corresponding to the second song is obtained by searching based on the second beat information, with the initial transition time point as the center. This includes: searching for beat points in a third direction where the playback time increases, with the initial transition time point as the center, based on the second beat information; if there is a beat point in the third direction that meets the preset strong beat conditions, then the beat point in the third direction that meets the preset strong beat conditions is determined as the second candidate beat point; if there is no beat point in the third direction that meets the preset strong beat conditions, then the beat point is searched in a fourth direction where the playback time decreases, and the beat point in the fourth direction that meets the preset strong beat conditions is determined as the second candidate beat point.

[0009] In some optional implementations, determining the initial transition time point based on the target song transition strategy, the first audio feature, and the second audio feature includes: determining the end transition point of the first song under the target transition strategy based on the first audio feature, and determining the start transition point of the second song under the target transition strategy based on the second audio feature; determining the initial transition time point based on the end transition point and the start transition point; wherein, the audio control parameters include the end transition point and the start transition point.

[0010] In some alternative implementations, if the minimum time deviation exceeds a preset tolerance value, audio control parameters are generated based on the initial transition time point.

[0011] In some optional implementations, the audio control parameters include an end transition point corresponding to the first song and an alignment anchor point associated with the first and second songs; under the audio control parameters, transitioning the first song to the second song includes: monitoring the playback progress of the first song; when the playback progress reaches the end transition point, using the audio control parameters, mixing the first audio segment of the first song starting from the end transition point and the second audio segment of the second song starting from the alignment anchor point to transition the first song to the second song.

[0012] In some optional implementations, the first audio feature includes multiple end-type structure points of the first song, and the second audio feature includes multiple start-type structure points of the second song; feature matching of the first audio feature and the second audio feature with each preset song transition strategy includes: for any preset song transition strategy, determining whether there is a target end-type structure point required by the preset song transition strategy among the multiple end-type structure points of the first song, and whether there is a target start-type structure point required by the preset song transition strategy among the multiple start-type structure points of the second song.

[0013] In some optional implementations, the ending structural point is determined by at least one of the following methods: the end time of a human voice that is no more than a preset transition window duration away from the end of the audio and whose ending energy exceeds a first preset energy threshold is determined as the human voice ending point; the end time of a drum beat that is no more than a preset transition window duration away from the end of the audio and whose ending energy exceeds a second preset energy threshold is determined as the drum beat ending point; when structural segment information exists and the start point of the last segment is no earlier than the determined human voice ending point or drum beat ending point, the start point of the last segment is determined as the structural segment ending point; within a preset number of measures, the end position of a measure whose distance from the end of the audio meets the duration limit and is no earlier than the determined human voice ending point or drum beat ending point is determined as the music measure ending point; wherein, the ending structural point includes at least one of the following: human voice ending point, drum beat ending point, structural segment ending point, and music measure ending point.

[0014] In some optional implementations, the starting structural point is determined by at least one of the following methods: determining the start time of a human voice that is within a preset transition window duration and whose starting energy exceeds a first preset energy threshold as the human voice starting point; determining the start time of a drum beat that is within a preset transition window duration and whose starting energy exceeds a second preset energy threshold as the drum beat starting point; determining the end point of the first segment as the structural segment starting point when structural segment information exists and the end point of the first segment is earlier than the determined human voice starting point or drum beat starting point; determining the starting point of a measure that meets the duration limit and precedes the determined human voice starting point or drum beat starting point as the music measure starting point within a preset number of measures; wherein, the starting structural point includes at least one of the human voice starting point, drum beat starting point, structural segment starting point, and music measure starting point.

[0015] Secondly, this application provides a song transition device, comprising: an acquisition module for acquiring a first audio feature of a currently playing first song and a second audio feature of a second song to be played; a determination module for determining a target song transition strategy from a variety of preset song transition strategies based on the first audio feature and the second audio feature; a generation module for generating audio control parameters for transitioning from the first song to the second song based on the target song transition strategy; and a transition module for transitioning the first song to the second song under the audio control parameters.

[0016] Thirdly, this application provides an electronic device, including: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to perform the song transition method described in the first aspect or any corresponding embodiment.

[0017] Fourthly, this application provides a computer-readable storage medium storing computer instructions for causing a computer to execute the song transition method of the first aspect or any corresponding embodiment described above.

[0018] Fifthly, this application provides a computer program product, including computer instructions for causing a computer to execute the song transition method described in the first aspect or any corresponding embodiment thereof.

[0019] The song transition method provided in this application, by acquiring the audio features of the first and second songs and intelligently selecting a transition strategy based on these features, fundamentally abandons the fixed and singular fade-in / fade-out mode of traditional players, achieving a leap from indiscriminate transition to content-aware transition. By analyzing the musical attributes of the songs themselves, the most suitable transition method can be dynamically selected for different song combinations, thereby significantly improving the continuity and naturalness of the listening experience and providing users with a smoother and more natural continuous music playback experience. Attached Figure Description

[0020] To more clearly illustrate the technical solutions in the specific embodiments or related technologies of this application, the drawings used in the description of the specific embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0021] Figure 1 This is a schematic flowchart of a first method for song transition according to an embodiment of this application; Figure 2 This is a schematic diagram of a second process for a song transition method according to an embodiment of this application; Figure 3 This is a schematic diagram of the third process of the song transition method according to the embodiments of this application; Figure 4 This is a structural block diagram of a song transition device according to an embodiment of this application; Figure 5 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of this application. Detailed Implementation

[0022] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0023] It should be noted that the information (including but not limited to user input information, such as information entered by the user into input boxes), data (including but not limited to data used for analysis, stored data, and displayed data, such as context code, all code of the current project, the service pressure corresponding to operations performed on all code of the current project, and the code development status of the current project), and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with relevant laws, regulations, and standards. For example, the context code, operations performed on all code of the current project, the corresponding service pressure, and the code development status involved in this application were all obtained with full authorization.

[0024] With the widespread adoption of digital music services, users have higher expectations for the listening experience during song playback, especially in continuous playback scenarios. Achieving smooth and natural transitions between songs has become a key technical aspect for improving user experience. Currently, common song transition solutions in the industry mainly fall into two categories: The first category is pre-mixed audio solutions (such as DJ Mix). This solution relies on professionals pre-editing, speed-changing, tuning, and mixing multiple songs in an audio workstation to generate a complete long audio file for playback. Although its transition effects are quite professional, it also has significant drawbacks: on the one hand, the transition process is completely fixed, and users cannot freely adjust the song order or skip tracks during playback; otherwise, the pre-mixed transition effects will fail. On the other hand, this solution is costly and time-consuming to produce, making it difficult to scale up to adapt to massive music libraries and personalized playlists, and it cannot respond to real-time user interactions.

[0025] The second type is a simple end-side fade-in / fade-out (Crossfade) scheme, which involves a linear or non-linear cross-fade transition between the volumes of two songs based on a fixed timeline. While this scheme offers some flexibility, its transition effect is often abrupt and can easily cause auditory discomfort. For example, because it doesn't consider the rhythm and section structure of the songs (such as the position of verses and choruses), it often results in rhythmic chaos due to misaligned drum beats or a "clashing" effect caused by overlapping vocals. Furthermore, this scheme has a simplistic strategy and cannot adaptively adjust to differences in song style and mood. The logic is usually fixed in each terminal client, leading to high algorithm update and maintenance costs and long cycles.

[0026] The song transition method provided in this application abandons the pre-mixing mode that relies on fixed audio files. By acquiring the audio features of two songs in real time and determining the transition strategy accordingly, the transition effect can be dynamically generated, thus responding to the user's real-time adjustments to the playback order and solving the shortcomings of pre-mixing schemes in terms of inflexibility and non-real-time performance. Simultaneously, it selects from multiple preset song transition strategies based on first and second audio features. This means that the transition strategy is no longer single and fixed, but can adaptively match according to the characteristics of the song itself, thereby avoiding, in principle, the problem of poor listening experience caused by a single or mismatched strategy. Furthermore, since strategy decision-making and parameter generation rely on definable rules and data, this core logic has the potential for centralized maintenance and updates, providing a foundation for overcoming the difficulties of maintenance and fragmentation.

[0027] According to an embodiment of this application, a song transition method embodiment is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Also, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0028] This embodiment provides a song transition method that can be used in electronic devices such as mobile phones, computers, and in-vehicle infotainment systems. Figure 1 This is a flowchart of a song transition method according to an embodiment of this application, such as... Figure 1 As shown, the process includes the following steps: Step S101: Obtain the first audio feature of the currently playing first song and the second audio feature of the second song to be played.

[0029] The first song refers to the currently playing song. During the transition, the first song is the one that needs to fade in or exit. The first audio feature refers to the music feature data of the first song, which may include, for example, beats per minute (BPM), key, cue points, beat grid, energy value, etc. These features are uniformly managed by the song feature engine, which has a multi-level caching mechanism (memory and disk) and supports silent updates of feature data via version number when the network is restored, ensuring data consistency. In the absence of network and caching, it can degrade to a normal fade-in / fade-out mode. The second song refers to the song to be played next. During the transition, it is the one that needs to fade in or enter. The second audio feature refers to the music feature data of the second song, the content of which is the same as the first audio feature. Specifically, the first and second audio features are obtained through a dedicated data management layer. This process does not rely on a single source, but works collaboratively through multiple means to ensure efficiency and reliability. Electronic devices will first try to obtain cached feature data from the local device for fast response. If the required data is not available locally, it will request high-precision feature information from the network server. Even in extreme situations where there is no network connection and no cache, electronic devices still have the ability to perform basic feature estimation on the device side or initiate degradation schemes to ensure the availability of functions. This feature data is a digital description of the musical attributes of a song, including key information such as rhythm, structure, and tonality, providing a unified "raw material" for subsequent intelligent decision-making.

[0030] Step S102: Based on the first audio feature and the second audio feature, determine the target song transition strategy from a variety of preset song transition strategies.

[0031] Pre-defined song transition strategies refer to various pre-defined mixing logics, such as: rhythm continuity strategies (BPM alignment), vocal continuity strategies (avoiding vocal overlap), ambient continuity strategies (utilizing ambient passages), and accompaniment continuity strategies. Specifically, multiple song transition strategies can be preset, each corresponding to different mixing logic and applicable conditions. During decision-making, the electronic device comprehensively compares and logically analyzes the feature data of the first and second songs, evaluating whether the applicable conditions of each preset song transition strategy are met. This analysis process aims to find the optimal mixing logic for the current specific song pairing that achieves a smooth and natural transition. Through its built-in decision logic, the electronic device selects the strategy that best matches the musical characteristics of the song pair from all available strategies as the final transition scheme, thereby ensuring that the transition effect is optimized under given audio feature conditions.

[0032] Step S103: Based on the target song transition strategy, generate audio control parameters for the transition from the first song to the second song.

[0033] The target song transition strategy refers to the final transition strategy selected based on the audio characteristics of the first and second songs. For example, if the BPMs of the two songs are similar, a rhythmic continuity strategy can be chosen as the target strategy. Audio control parameters refer to the specific execution parameters required to implement the selected transition strategy. These may include: time anchors (such as `mix_start_time`), volume curves (fade out of the first song, fade in of the second song), EQ adjustment curves, speed adjustment curves (if BPM alignment is required), diagnostic information (for retrospective analysis), etc. These parameters are encapsulated in a standardized AutoMixOutput object to decouple the decision layer from the rendering layer. For example, the AutoMixOutput object, as the output of the decision engine, may specifically contain the following core fields: mix_strategy: The type of strategy hit (such as rhythm continuation, general fade-in / fade-out, etc.); mix_params: The specific set of execution parameters, including but not limited to the fade-out time of the first song, the fade-in time of the second song, the volume curve, the equalizer curve, the tempo curve, etc. Diagnostics: Diagnostic information records the path and reasons for this decision in a structured form, such as why the current strategy was hit, or why a higher priority strategy was not matched, for subsequent effect backtracking and algorithm optimization.

[0034] Specifically, after determining the target transition strategy, the electronic device translates this strategy into a series of executable, precise control instructions—audio control parameters. These parameters constitute a detailed "transition script," which not only specifies when the transition should begin but also finely plans how the two songs should change throughout the transition period to achieve a smooth transition. For example, the parameters define in detail the fade-out curve of the first song's volume, the fade-in curve of the second song's volume, whether a temporary speed adjustment is needed for a song to align with the rhythm, and whether specific frequency bands need to be boosted or attenuated. All these parameters are encapsulated into a standardized instruction set, thereby decoupling the high-level mixing intentions from the low-level audio execution.

[0035] Step S104: Under the audio control parameters, transition from the first song to the second song.

[0036] A sophisticated collaborative control mechanism ensures the accurate and real-time execution of the "transition script." Specifically, this mechanism comprises two core components: first, state management, which clearly defines the entire lifecycle of the transition from preparation, initiation, execution to termination, ensuring it remains in the correct logical state under any playback scenario; and second, timeline scheduling, which closely monitors the song's playback progress and triggers corresponding audio control actions at preset, precise time points. When the playback progress reaches the moment specified by the control parameters, the scheduling system in the electronic device drives the underlying audio processing engine to apply specific operations such as volume changes and speed adjustments in real time. The audio rendering engine employs dual player instances and an atomic effects processing mechanism, providing cross-platform fine-grained control interfaces for volume, equalizer (EQ), and speed adjustment. During speed adjustments, it prioritizes speed alignment before transitioning volume and EQ changes. Through this linkage between state and time, the electronic device can precisely control the ebb and flow of two songs, ultimately achieving a seamless and smooth intelligent transition effect.

[0037] The song transition method provided in this application, by acquiring the audio features of the first and second songs and intelligently selecting a transition strategy based on these features, fundamentally abandons the fixed and singular fade-in / fade-out mode of traditional players, achieving a leap from indiscriminate transition to content-aware transition. By analyzing the musical attributes of the songs themselves, the most suitable transition method can be dynamically selected for different song combinations, thereby significantly improving the continuity and naturalness of the listening experience and providing users with a smoother and more natural continuous music playback experience.

[0038] This embodiment provides a song transition method that can be used in electronic devices such as mobile phones, computers, and in-vehicle infotainment systems. Figure 2 This is a flowchart of a song transition method according to an embodiment of this application, such as... Figure 2As shown, the process includes the following steps: Step S201: Obtain the first audio feature of the currently playing first song and the second audio feature of the second song to be played. For details, please refer to [link to relevant documentation]. Figure 1 Step S101 of the illustrated embodiment will not be described again here.

[0039] Step S202: Based on the first audio feature and the second audio feature, determine the target song transition strategy from a variety of preset song transition strategies.

[0040] Specifically, step S202 includes: Step S2021: Obtain the transition priority corresponding to each preset song transition strategy.

[0041] The strategy selection is executed by the transition strategy decision engine, which supports dynamic script distribution and hot updates, and can adjust the matching rules and priority logic without relying on client releases.

[0042] Transition priority refers to the level or numerical value assigned to each preset song transition strategy, indicating its degree of priority. Specifically, the electronic device maintains a predefined set of strategy configurations, where each preset song transition strategy is associated with a specific transition priority attribute. This attribute is an identifier used for comparison and ranking, defining the order in which different strategies are considered and applied in the electronic device's decision-making logic. At the start of the decision-making process, the electronic device obtains all available strategies and their corresponding priority identifiers directly by reading this internal configuration or from an updatable rule file. This acquisition process is part of the electronic device's initialization, ensuring that the decision-making logic operates within a clear and consistent comparison framework. The priority itself is pre-assigned during the design phase based on factors such as the theoretical merits of each strategy's transition effect, the stringency of its song feature matching conditions, and its universality as a fallback guarantee.

[0043] In some optional implementations, the preset song transition strategies may include: rhythm continuity strategy (BPM alignment), vocal continuity strategy (avoiding vocal overlap), ambient continuity strategy (utilizing ambient sections), accompaniment continuity strategy, silence transition strategy, general fade-in / fade-out strategy (fallback strategy), and direct song skipping strategy (abnormal fallback). Their priority logic is as follows: P0 - Rhythm Strategy: Optimal effect. Requires the first and second songs to have similar BPMs and compatible Outro / Intro beat sequences. For example, the quantitative condition for similar BPMs is: the BPM difference between the two songs is within 5%, or the difference between twice the BPM of one song and the BPM of another song is within 5%. By adjusting the tempo, the BPMs of the two songs are unified, achieving seamless transitions with synchronized drum beats. P1 - Vocal Strategy: Suboptimal effect. When the rhythms are misaligned but there are clear vocal segments (a clear vocal segment refers to a valid vocal cue point), the vocal entry point is used to connect the segments and avoid vocal "clashing". P2 - Atmosphere Strategy: Utilize atmospheric sections of the song (such as sustained notes and ambient sounds) for a longer period of dissolve; P3 - Accompaniment Strategy: Using pure accompaniment sections for transitions; P4 - Silence Strategy: When two songs have vastly different styles or require a pause, a silence is inserted in the middle. P5 - Universal Fade-in / Fade-out Strategy (Crossfade): A fallback strategy that uses a standard volume cross curve; P6 - Do Nothing Strategy: Directly switch songs in abnormal situations.

[0044] In the above implementation, the setting of transition priorities and the mechanism of feature matching according to priority order ensure that the most optimal transition scheme in terms of musical sound (such as the P0 rhythm continuation strategy) is tried first. Only when a higher-level strategy is unavailable due to the song's features not meeting the applicable conditions will it automatically downgrade to a suboptimal or basic scheme (such as the P1 vocal continuation strategy or even the P5 universal fade-in / fade-out strategy). This design achieves an effective balance between pursuing the best listening experience and ensuring functional robustness—it avoids the complete failure of the transition function due to overly stringent conditions, and it also avoids easily giving up any opportunity to improve the listening experience. Secondly, the rule that the first successful match is determined as the target strategy allows the decision-making process to terminate immediately after finding a strategy that meets the conditions, without having to perform meaningless calculations and evaluations on all subsequent low-priority strategies. This not only avoids unnecessary computational overhead but also significantly improves decision-making efficiency.

[0045] Step S2022: Based on the priority order of each transition priority representation, the first audio feature and the second audio feature are respectively matched with each preset song transition strategy.

[0046] Priority order refers to a fixed attempt or matching order from high to low (i.e., from the optimal strategy to the fallback strategy) based on the transition priority of each transition strategy. Specifically, after clarifying the priority order of all strategies, feature matching checks will be performed on each preset song transition strategy in descending order. For the currently checked strategy, its unique, pre-defined matching conditions are extracted (for example, for the rhythm continuation strategy, the conditions will include the BPM difference between the two songs being within a certain range and both having specific structural segments). Subsequently, the audio feature data of the first and second songs (such as BPM values, CUE point positions, etc.) are substituted into these conditions for logical calculation and comparison. This matching process is a systematic evaluation: it analyzes whether the feature combination of the two songs fully satisfies all applicable conditions of the strategy. If a strategy fails to match, diagnostic information is directly recorded and the next priority strategy is tried, without retaining intermediate states. The electronic device will independently perform such an evaluation for each strategy and decide whether to adopt the current strategy or continue to perform the same feature matching operation on the next lower priority strategy according to the established order based on the evaluation result (satisfied or not satisfied).

[0047] In some optional implementations, the first audio feature includes multiple end-type structure points of the first song, and the second audio feature includes multiple start-type structure points of the second song; feature matching of the first audio feature and the second audio feature with each preset song transition strategy includes: for any preset song transition strategy, determining whether there is a target end-type structure point required by the preset song transition strategy among the multiple end-type structure points of the first song, and whether there is a target start-type structure point required by the preset song transition strategy among the multiple start-type structure points of the second song.

[0048] Multiple ending structure points refer to the set of candidate time points at the end of the first song (the song that is about to end), used to trigger a fade-out or remix. Multiple starting structure points refer to the set of candidate time points at the beginning of the second song (the song that is about to start), used to trigger a fade-in or remix.

[0049] The target start structure point and target end structure point refer to the specific point selected from a set of multiple start structure points and multiple end structure points that meet the specific conditions of the strategy when matching a specific transition strategy (such as P0 rhythm continuation). For example, when performing rhythm continuation, the goal is to find the drum start point and the drum end point; these two points are the target structure points under the current strategy.

[0050] Specifically, for any pre-defined song transition strategy, the specific structural point type upon which the strategy depends is first extracted as a matching condition. For example, a rhythmic continuation strategy requires the first song to have a valid drum end point or a musical measure end point, while the second song requires a valid drum start point or a musical measure start point. Then, the search is iterated through the set of end-type structural points contained in the first audio feature corresponding to the first song and the set of start-type structural points contained in the second audio feature corresponding to the second song. For the first song, it is checked whether its set of end-type structural points contains a point type matching the current strategy's requirements, and whether the point's value is valid (non-negative and within a reasonable transition window). For the second song, it is similarly checked whether its set of start-type structural points contains a corresponding valid point type. Only when both the set of end-type structural points of the first song and the set of start-type structural points of the second song contain the target point type required by the strategy is the strategy considered a successful match. If either set lacks a target structural point of the required type, the match fails, and the decision engine will automatically move to the next strategy according to a pre-defined priority order until the first strategy that satisfies all conditions is found.

[0051] In the above implementation, the first and second audio features are concretized into multiple categories of end-type and start-type structure points, respectively, thereby transforming the abstract feature matching process into a judgment of the existence of specific musical structure markers. For any preset song transition strategy, it is only necessary to check whether the first song has the target end-type structure point on which the strategy depends, and whether the second song has the corresponding target start-type structure point, to quickly determine the applicability of the strategy in the current song combination. This matching method based on structured markers enables the decision engine to accurately identify the specific musical segments within each song that can be used for transition, rather than relying on vague global features for fuzzy judgment, significantly improving the targeting and accuracy of strategy selection, and also providing clear candidate boundaries for the precise positioning of subsequent transition time points.

[0052] In some optional implementations, the ending structural point is determined by at least one of the following methods: the end time of a human voice that is no more than a preset transition window duration away from the end of the audio and whose ending energy exceeds a first preset energy threshold is determined as the human voice ending point; the end time of a drum beat that is no more than a preset transition window duration away from the end of the audio and whose ending energy exceeds a second preset energy threshold is determined as the drum beat ending point; when structural segment information exists and the start point of the last segment is no earlier than the determined human voice ending point or drum beat ending point, the start point of the last segment is determined as the structural segment ending point; within a preset number of measures, the end position of a measure whose distance from the end of the audio meets the duration limit and is no earlier than the determined human voice ending point or drum beat ending point is determined as the music measure ending point; wherein, the ending structural point includes at least one of the following: human voice ending point, drum beat ending point, structural segment ending point, and music measure ending point.

[0053] The vocal ending point refers to the moment when the vocals begin at the start of the second song. The condition is that it is within the first 15 seconds and the energy exceeds the threshold.

[0054] The drum end point refers to the moment when the drumbeat at the beginning of the second song starts. The condition is that it is within the first 15 seconds and the energy level is sufficient.

[0055] The end point of a structural paragraph refers to the end of the first verse of the second song. It must be earlier than the determined start point of the vocals or drumbeats.

[0056] The end point of a musical measure refers to the starting point of a measure at the beginning of the second song that meets the preset transition duration limit. This starting point is determined by: within the preset number of measures, traversing and searching towards the end of the song in descending order of measure count, selecting the first candidate measure starting point that meets the duration requirement and precedes the already determined vocal / drum start point.

[0057] Specifically, the ending structure point describes the key moment at the end of the first song that can be cut out for mixing. Its determination process is also completed by the song feature engine based on song feature data and spectrum analysis. For vocal ending points, the validity of the vocal ending time field is first checked, and the remaining distance between this time point and the total duration of the entire song is calculated to ensure that the remaining distance does not exceed the preset transition window duration limit (e.g., 15 seconds). At the same time, it is checked whether the vocal energy spectrum at the end of the song exceeds the vocal energy threshold. When all these conditions are met, the vocal ending moment is identified as a vocal ending point.

[0058] The determination of the drum beat ending point follows a similar logic. Based on the drum beat ending time field and the drum sound energy spectrum, after confirming that the remaining time between the drum beat ending time and the end of the song is within the allowed window and the final drum sound energy meets the standard, the drum beat ending time is identified as the drum beat ending point.

[0059] The determination of the end point of a structural paragraph depends on the structural segmentation information. If structural paragraph data exists, the starting point of the last structural paragraph is obtained as a candidate, and it is verified whether the starting point is not earlier than the determined end point of the vocal or drum beat (i.e., the starting point of the structural paragraph is located after the end of the vocal or drum beat). If this positional relationship is satisfied, the starting point of the last segment is established as the end point of the structural paragraph.

[0060] The determination of the end point of a music measure is carried out by traversing the search within a preset number of measures, starting from the largest number of measures and moving down to the smallest number of measures in descending order of priority. For each candidate number of measures, the corresponding measure duration is calculated, and the total song duration minus the measure duration is used as the target time point. Then, it is checked whether the target time point is not earlier than the already determined end point of vocals or drum beats (i.e., the end position of the measure is after the end of vocals or drum beats). The first target time point that simultaneously meets the duration limit and positional relationship conditions is determined as the end point of the music measure. If no point that meets the conditions is found after traversing the search, it will fall back to the already determined end point of vocals or directly use the total song duration minus the upper limit of the transition window as the default value.

[0061] In some optional implementations, the starting structural point is determined by at least one of the following methods: determining the start time of a human voice that is within a preset transition window duration and whose starting energy exceeds a first preset energy threshold as the human voice starting point; determining the start time of a drum beat that is within a preset transition window duration and whose starting energy exceeds a second preset energy threshold as the drum beat starting point; determining the end point of the first segment as the structural segment starting point when structural segment information exists and the end point of the first segment is earlier than the determined human voice starting point or drum beat starting point; determining the starting point of a measure that meets the duration limit and precedes the determined human voice starting point or drum beat starting point as the music measure starting point within a preset number of measures; wherein, the starting structural point includes at least one of the human voice starting point, drum beat starting point, structural segment starting point, and music measure starting point.

[0062] The vocal start point refers to the moment when the vocals at the end of the first song cease. The conditions are that the time remaining until the end of the song does not exceed a preset number of seconds (e.g., 15 seconds), and the energy at that moment exceeds a threshold.

[0063] The drum start point refers to the moment when the drumbeat at the end of the first song ends. The conditions are that the time remaining until the end does not exceed a preset number of seconds (e.g., 15 seconds) and the energy level is sufficient.

[0064] The structural section start point refers to the beginning of a section at the end of the first song. It must not be earlier than the determined end point of the vocals or drum beats.

[0065] The starting point of a musical measure refers to the beginning of a measure at the end of the first song that meets the preset transition duration limit. Within the preset number of measures to search, the search proceeds in descending order of measure count towards the beginning of the song, selecting the first candidate measure starting point that meets the duration requirement and is no earlier than the already determined end point of vocals / drum beats.

[0066] Specifically, the starting structure point is used to describe the key moment in the beginning of the second song that can be used for mixing. Its determination process is based on the song feature engine's comprehensive judgment of the input audio metadata and spectrum analysis results. The determination of the vocal starting point depends on the pre-analyzed vocal starting time field, checking whether the time value is valid (non-negative) and whether the time point is within the upper limit of the preset transition window duration (e.g., 15 seconds) from the beginning of the song. At the same time, it combines the vocal energy spectrum index of the beginning of the song. Only when the energy value exceeds the preset vocal energy threshold is the vocal starting moment officially recognized as the vocal starting point.

[0067] The logic for determining the starting point of a drum beat is similar to that for the starting point of a human voice, but it is based on the drum beat start time field and the drum sound energy spectrum index. Only when the start time of the drum beat is within the allowed window duration and the drum sound energy exceeds the corresponding drum sound energy threshold, is that moment recognized as the starting point of the drum beat.

[0068] The starting point of the structural segment is determined based on the structural segmentation information of the song. If the list of structural segments is not empty, the end time of the first structural segment is extracted as a candidate, and it is checked whether the end time is earlier than the determined start point of the vocals or the start point of the drumbeat (i.e. the structural segment ends before the vocals or drumbeats appear). If this leading relationship is satisfied, the end point of the first segment is established as the starting point of the structural segment.

[0069] The determination of the starting point of a music measure employs an adaptive search mechanism. Within a preset range of measure counts, the search proceeds by traversing from the largest to the smallest measure counts in descending order of priority. For each candidate measure count, the corresponding measure duration is calculated, and it is checked whether the starting point of the measure is within the transition window duration limit and whether its time position precedes the determined vocal or drum start point. The first measure start point that meets all the conditions is determined as the starting point of the music measure. If no music measure start point that meets the conditions is found after the search is completed, the search will backtrack to the determined vocal start point (if it exists) or directly use the upper limit of the transition window as a fallback value.

[0070] For example, the determination of start and end structural points (i.e., CUE points) is based on the song's tempo, duration, vocal / drum start and end times, structural segmentation list, and spectral energy threshold. Eight types of CUE points are independently determined using preset rules, providing precise segment anchors for transitions. The list of eight types of CUE points is shown below: Start: cue_vocal_start, cue_drum_start, cue_struct_start, cue_music_start; End: cue_vocal_end, cue_drum_end, cue_struct_end, cue_music_end.

[0071] Specifically, for the starting category: the vocals and drum beats must be within 15 seconds and the starting energy must exceed their respective thresholds; the structure should be the end of the first segment and earlier than the vocals; the music should traverse from 8 to 1 measures to satisfy the length and leading relationship, and if not, it should fall back to the vocals or 15 seconds. For ending categories: vocals and drum beats must be no more than 15 seconds from the end and the energy at the end must exceed the threshold; the structure is taken from the starting point of the last segment, which must not be earlier than the already determined vocals / drum beats; the music searches forward from measure 8→1, and if not found, it goes back to 15 seconds before the vocals or the end.

[0072] The eight CUE points are determined independently according to the above branches, and finally form the set of start and end anchor points of the transition window.

[0073] In the above embodiments, by refining and extracting ending structural points in multiple dimensions and types, a comprehensive characterization of the musical structure of the ending section of the first song is achieved. This application comprehensively considers various musical elements such as the end of vocals, the end of drumbeats, the termination of structural paragraphs, and the boundaries of musical measures, and combines distance constraints and energy thresholds with the end of the audio for validity verification, thereby capturing multiple candidate positions most suitable for the ending of the first song from different musical levels. This multi-level structural point extraction mechanism provides subsequent transition strategies with richer musical semantic information to choose from during matching, significantly improving the flexibility and accuracy of ending transition point positioning, ensuring that the first song can fade out at the most natural structural nodes such as vocal fading, rhythmic decay, or paragraph termination, effectively avoiding the abruptness of abrupt cutoffs. In addition, by defining and extracting starting structural points in a symmetrical and corresponding manner, a deep analysis of the musical structure of the beginning section of the second song is achieved. This application detects various musical anchor points within the intro area of ​​the second song from multiple dimensions, including the start of vocals, the entry of drumbeats, the beginning of structural sections, and the start of musical measures. Using preset transition window durations and starting energy thresholds as constraints, it can accurately pinpoint various musical anchor points with initiation significance within the intro area. This comprehensive mechanism for acquiring starting structural points provides a rich set of candidate entry points for transition strategies, allowing for flexible selection of the most suitable starting position based on the needs of different transition strategies. This ensures that the second song can enter at the most natural moment, such as when vocals emerge, rhythm is established, or sections unfold, laying a precise starting foundation for a seamless transition between the two songs.

[0074] Step S2023: The first preset song transition strategy that successfully matches the feature is determined as the target song transition strategy.

[0075] During feature matching according to priority, the results of each match are continuously monitored. Once a successful feature matching evaluation for a certain strategy is detected (i.e., the features of the two songs fully meet all applicable conditions of the strategy), the electronic device immediately terminates the matching check process for all subsequent low-priority strategies. At this point, the first strategy confirmed as a successful match is selected as the target song transition strategy for this transition. This mechanism ensures that the electronic device always automatically selects the theoretically optimal (i.e., highest priority) and practically usable transition method under the current song feature conditions. The selection of this strategy marks the completion of the decision-making phase, and the electronic device will generate subsequent audio control parameters based on this strategy, without considering other usable but lower-priority alternative strategies.

[0076] The song transition method provided in this application constructs an orderly and efficient strategy selection mechanism by assigning clear transition priorities to each preset song transition strategy and strictly following the order represented by these priorities for feature matching. This mechanism ensures that the decision-making process is not blind traversal or random selection, but rather prioritizes strategies with better listening effects and more natural transitions, only downgrading subsequent strategies for evaluation when the current strategy fails to meet the matching conditions due to audio features. This chain-like matching logic, which prioritizes the optimal strategy, can quickly and deterministically lock in the best transition scheme achievable under the current combination of music features, ensuring the upper limit of transition effect quality while avoiding resource overhead caused by ineffective computation, and making the decision path clear, controllable, and easy to trace.

[0077] In some alternative implementations, before matching audio features with policies, the electronic device first performs preprocessing: verifying the integrity of the first and second audio features, including checking if the BPM value is valid, if the song duration is sufficient to support the transition, and if key CUE data exists. If data is missing or invalid, the electronic device attempts to estimate the BPM in real time on the device side (e.g., using a simple algorithm). If estimation is not feasible, the corresponding song is marked as having unavailable features and considered in subsequent policy matching, potentially triggering a degradation process. Preprocessing ensures the quality of the data input to the decision engine, laying the foundation for accurate matching of subsequent policies.

[0078] Step S203: Based on the target song transition strategy, generate audio control parameters for the transition from the first song to the second song. For details, please refer to [link to details]. Figure 1 Step S103 of the illustrated embodiment will not be described again here.

[0079] Step S204: Under the audio control parameters, transition from the first song to the second song. See details below. Figure 1 Step S104 of the illustrated embodiment will not be described again here.

[0080] This embodiment provides a song transition method that can be used in electronic devices such as mobile phones, computers, and in-vehicle infotainment systems. Figure 3 This is a flowchart of a song transition method according to an embodiment of this application, such as... Figure 3 As shown, the process includes the following steps: Step S301: Obtain the first audio feature of the currently playing first song and the second audio feature of the second song to be played. For details, please refer to [link to relevant documentation]. Figure 2 Step S201 of the illustrated embodiment will not be described again here.

[0081] Step S302: Based on the first audio feature and the second audio feature, determine the target song transition strategy from a variety of preset song transition strategies. For details, please refer to [link to relevant documentation]. Figure 2 Step S202 of the illustrated embodiment will not be described again here.

[0082] Step S303: Based on the target song transition strategy, generate audio control parameters for the transition from the first song to the second song.

[0083] Specifically, step S303 includes: Step S3031: Determine the initial transition time point based on the target song transition strategy, the first audio feature, and the second audio feature.

[0084] The initial transition time point refers to a rough, unrhythm-aligned transition trigger reference time determined based on the target song's transition strategy, the first audio feature, and the second audio feature. It is usually determined based on the song's structural logic, but may not be optimal in terms of rhythmic precision. Specifically, based on the selected target song's transition strategy, and combining the first and second audio features of both songs, logical calculations are performed to determine the initial transition time point. The core of this process is that different transition strategies have different preferences and requirements for the connecting sections of the songs. For example, if the strategy is a rhythmic continuation strategy, it will focus on finding sections in both songs with distinct beat characteristics and structurally suitable for mixing (such as the outro of the first song and the intro of the second song); if the strategy is a vocal continuation strategy, it will prioritize locating clear vocal start and end points. The electronic device analyzes the corresponding segment markers (CUE points), energy distribution, and other data in the first and second audio features, and selects a preliminary connection position for each song according to the rules defined by the strategy. Ultimately, the initial transition point is usually determined as the start of the first song's transition section, which serves as the benchmark for subsequent rhythmic fine-tuning and alignment.

[0085] Specifically, different strategies define their own rules for calculating transition positions. For example, for the rhythm continuation strategy, the first step is to check whether the BPM of the two songs meets the quantitative condition of "closeness"—that is, the difference in BPM between the first and second songs does not exceed a preset percentage (e.g., 5%), or the difference between twice the BPM of the first song and the BPM of the second song does not exceed a preset percentage (and vice versa). If the BPM condition is met, the next step is to check whether the first song has a valid end point of a musical measure or a drumbeat, and whether the second song has a valid start point of a musical measure or a drumbeat. If all of the above CUE points exist, the time positions of these CUE points are used as preliminary transition position candidates, and an appropriate transition overlap duration is calculated based on the BPM difference. Finally, the end transition point of the first song is anchored near the strong beat in its coda, and the time position of this end transition point is determined as the initial transition time point.

[0086] For example, regarding vocal transition strategies, when rhythmic transition strategies are unavailable due to insufficient BPM or missing necessary cue points, the transition is based on vocal segments. This strategy requires the first song to have a valid vocal ending point and the second song to have a valid vocal beginning point, or at least one of them to have a clear vocal segment. After confirming that the cue point is available, the vocal ending point of the first song is used as the reference position for its fade-out start, and the vocal beginning point of the second song is used as the reference position for its fade-in start. The overlap duration of the two vocal segments is controlled to not exceed a preset threshold to avoid a "vocal clash" sound. At this point, the initial transition time is set to the time position of the vocal ending point of the first song.

[0087] In some optional implementations, step S3031 above includes: Step a1: Determine the end transition point of the first song under the target transition strategy based on the first audio feature, and determine the start transition point of the second song under the target transition strategy based on the second audio feature.

[0088] The end transition point refers to the selected time point for the first song (the currently playing song) where fade-out or transition processing begins. The start transition point refers to the selected time point for the second song (the song to be played) where fade-in or transition processing begins. Specifically, for the first song (the currently playing song), its end transition point is determined based on its first audio characteristics and the guidance of the target transition strategy. This point is the starting position in the first song where the fade-out or end processing logic begins. The electronic device scans the segment information in the first audio characteristics (such as CUE points: cue_music_end, cue_vocal_end, etc.) and evaluates whether these segments meet the applicable conditions of the current strategy (e.g., whether it is a pure accompaniment segment, an ambient segment, or a segment with a clear beat). Similarly, for the second song (the song to be played), the electronic device determines its start transition point based on its second audio characteristics. This point is the logical starting position in the second song where the mix begins, and the electronic device looks for characteristic positions such as the intro or vocal intro. The determination of these two points is a rule-based feature matching process, which aims to find the audio segment boundaries for each song that best fit the current strategy intent for connection.

[0089] Step a2: Determine the initial transition time point based on the end transition point and the start transition point.

[0090] The audio control parameters include the end transition point and the start transition point.

[0091] After determining the end and start transition points for each of the two songs, these two points need to be integrated into a unified, executable transition trigger reference time. Typically, the initial transition time point is directly set to the end transition point of the first song. This is because, in the actual playback process, the triggering of transition behaviors (such as initiating a fade-out or starting preloading of the second song) is marked by the first song's playback progress reaching its end transition point. The start transition point of the second song, on the other hand, guides the audio engine to pick up the audio stream of the second song from that point for mixing preparation. Therefore, determining the initial transition time point based on the end and start transition points essentially establishes the structural end time of the first song as the time base for the entire transition event, while the start point of the second song serves as a related parameter, together constituting the time window definition for the transition.

[0092] In the above embodiments, during the generation of audio control parameters, the determination of the initial transition time point is broken down into two independent yet interconnected sub-steps—namely, firstly, the end transition point of the first song and the start transition point of the second song are locked based on the first and second audio features, respectively, and then the initial transition time point is calculated based on these features—achieving a deep binding between the transition boundary and the musical structure features. This application enables the setting of the transition window to no longer rely on a fixed preset duration or general global parameters, but rather to be personalized based on the actual usable musical segments of each song under a specific transition strategy. This feature-driven, dual-end positioning and collaborative calculation method ensures that the first song can naturally end at the most suitable segment and that the second song can smoothly start playing from the most appropriate entry point, thereby providing a more accurate and musically logical time reference for subsequent beat alignment and parameter curve generation.

[0093] Step S3032: Obtain the first beat information of the first song and the second beat information of the second song.

[0094] The first and second beat information refer to the beat grid data of the first and second songs, respectively. This data contains the precise time position information of each beat in the song, especially the position of the accented beat. This is the foundational data for rhythm alignment. Specifically, the first beat information of the first song and the second beat information of the second song are important components of the song's audio feature data. Specifically, they refer to the beat grid data of each song, which contains the precise millisecond-level position of each beat (especially the accented beat) on the song's timeline. This information is not generated in real-time during decision-making. Instead, the backend audio analysis service pre-detects and calculates the song's beats and stores them in the song feature database along with other features such as BPM and CUE points. When the device system needs this information, the song feature engine retrieves it from the local cache or from the server via the network. This retrieval process is part of data retrieval, ensuring that electronic devices can directly use high-precision beat timing data, providing crucial input for subsequent accurate accented beat alignment algorithms.

[0095] Step S3033: Centered on the initial transition time point, search for the first candidate heavy beat point corresponding to the first song based on the first beat information, and search for the second candidate heavy beat point corresponding to the second song based on the second beat information.

[0096] The downbeat refers to the first beat of a musical measure, typically the moment with the strongest rhythm and most energy. In mixing, aligning the downbeats of the first and second songs is crucial for achieving a seamless rhythmic transition. The first and second candidate downbeats refer to candidate downbeat points near the initial transition time, found through a search strategy for the first and second songs respectively, that meet the musical structure requirements. They are alternative solutions for fine-tuned rhythmic alignment. Specifically, for the currently playing first song, the algorithm uses the initial transition time as a reference center and searches the beat grid data of the first song for the closest downbeat point that meets the musical structure requirements. For the second song to be played, the algorithm uses the same initial transition time as a reference center and searches the beat grid data of the second song for a suitable downbeat point. Through this search mechanism defined separately for each song, a candidate downbeat point closest to the initial transition time can be determined for each of the two songs—the first and second candidate downbeats—providing a basis for subsequent precise alignment.

[0097] In some optional implementations, the above search can be a bidirectional search mechanism, covering the entire beat sequence. For the first song, the search prioritizes looking backward (in the direction of decreasing playback time) for the nearest stressed beat; for the second song, the search prioritizes looking backward (in the direction of increasing playback time) for the nearest stressed beat. If no stressed beat matching the phase requirement is found in the main direction, the search switches to the other direction. Specifically, the search constructs a dynamic search window centered on the initial transition time point (T_approx), covering the entire beat sequence without additional window limitations. The search starting point is located at the known beat point in the beat grid closest to T_approx, and then scanning is performed according to a preset main direction priority: for the first song, the main direction is towards decreasing playback time (looking backward to find the nearest stressed beat T_prev of the previous song); for the second song, the main direction is towards increasing playback time (looking backward to find the nearest stressed beat T_next of the next song). If no stressed beat matching the preset conditions is found in the main direction, the search switches to the other direction.

[0098] In some optional implementations, the first candidate hard beat point corresponding to the first song is searched based on the first beat information, with the initial transition time point as the center, including: Step b1: Using the initial transition time point as the center, search for a replay point in the first direction where the playback time decreases based on the first beat information.

[0099] The first direction refers to the search in the direction of decreasing playback time (i.e., towards the earlier part of the song). Specifically, this is the main direction search for the first song in a bidirectional search. The electronic device starts from the initial transition time point (T_approx) and checks the recorded beats in the beat grid one by one along the timeline in the reverse playback direction (i.e., towards the past time period of the song). The algorithm sequentially determines whether each beat is a downbeat, usually by checking its phase within a measure (e.g., whether it is the first beat in 4 / 4 time). Simultaneously, it verifies whether the beat meets other preset musical structure conditions. The search continues until the first beat that simultaneously satisfies both the downbeat and the preset phase / structure requirements is found, or the entire beat sequence is searched. Once such a satisfying downbeat is found in the main direction, the search immediately stops, and that point is considered a valid discovery in that direction. This direction-priority selection aligns with the auditory logic that fade-out sections in a song typically need to begin with a stable strong beat.

[0100] Step b2: If there is a retake point in the first direction that meets the preset strong shot conditions, then the retake point in the first direction that meets the preset strong shot conditions is determined as the first candidate retake point.

[0101] The preset strong beat condition refers to the filtering rules set when searching for candidate strong beats. Specifically, when searching based on beat information, each beat in the beat grid is traversed, and it is checked whether the beat is marked as a strong beat in the beat information of the corresponding song. The beat grid is a sequence of timestamps, each timestamp corresponding to a beat and carrying a flag indicating whether the beat is the first beat of the measure (i.e., the strong beat). Only those beats that are explicitly marked as strong beats in the beat grid are considered to meet the preset strong beat condition and can be included in the range of candidate strong beats. For example, for a 4 / 4 time signature song, the strong beat is the first beat of each measure; for other time signatures (such as 3 / 4 time signature, 6 / 8 time signature), the strong beat also refers to the first beat of the measure.

[0102] Specifically, if a beat point meeting the preset strong beat criteria is successfully found in the search along the first direction (backwards) where playback time decreases, the electronic device will directly adopt this point as the first candidate strong beat point and will not initiate a search along the second direction (where playback time increases). This is because the principle of prioritizing the main direction gives higher priority to the forward search, the goal of which is to find the nearest strong beat in the first song before the transition point. Using this strong beat as the "exit" of the rhythmic section of the first song makes the transition more natural. Once determined, the precise timestamp of this beat point is recorded for subsequent deviation calculations and final anchor point determination.

[0103] Step b3: If there is no repeating point that meets the preset strong beat conditions in the first direction, then search for repeating points in the second direction where the playback time increases, and determine the repeating points that meet the preset strong beat conditions in the second direction as the first candidate repeating points.

[0104] Among them, the preset strong beat condition refers to the point in time that is marked as a strong beat in the tempo information of the corresponding song.

[0105] The second direction refers to the first song. If the search in the first direction fails, the search proceeds in the direction of increasing playback time (i.e., towards the later part of the song, a later time). Specifically, if no matching strong beat is found within the entire search range in the first direction (backwards) (for example, T_approx might happen to be in a long non-strong beat section), the electronic device executes an alternative search path. At this point, the search switches to the second direction of increasing playback time (i.e., towards the later time segment of the song). Starting from T_approx, the search proceeds backwards (towards future time) in the beat grid along the playback direction. The search logic is the same as in the first direction: each beat is checked individually, and the first beat that satisfies the preset strong beat condition is found. Once such a beat is found in the second direction, it is identified as the first candidate strong beat. This mechanism is an important fault tolerance and guarantee strategy, ensuring that regardless of where the initial transition point falls in the song, the electronic device can find at least one usable rhythmic anchor point for the first song, avoiding the complete failure of the rhythm alignment function due to search failures.

[0106] In the above implementation, during the search for the first candidate strong beat point, a mechanism is established that prioritizes backtracking towards the first direction where the playback time decreases, centered on the initial transition time. This makes the anchor point location for the first song more consistent with the natural auditory habit of tracing back to the strong beat at the end of the music. Even if no valid strong beat point meeting the strong beat condition is found in the first direction, the search can automatically switch to the second direction where the playback time increases, thus forming a bidirectional, hierarchical strong beat point detection logic. This backtracking-then-forward search strategy ensures accurate capture of the best strong beat position within the transition window in most musical structures, and effectively addresses the irregularities or edge cases that may exist in the beat markers through a bidirectional fault tolerance mechanism. This significantly improves the robustness and hit rate of strong beat point extraction, laying a reliable data foundation for achieving seamless rhythm alignment.

[0107] In some optional implementations, the second candidate heavy beat point corresponding to the second song is searched based on the second beat information, with the initial transition time point as the center, including: Step c1: Using the initial transition time point as the center, search for replay points in a third direction based on the second beat information and the increase in playback time.

[0108] The third direction pointer, for the second song, searches in the direction of increasing playback time (i.e., towards the later part of the song, a later time). Specifically, contrary to the first song, for the upcoming second song, its main direction is preset to the direction of increasing playback time (i.e., searching backward, looking for a later beat). The electronic device also uses a shared T_approx as the central reference point (note that for the second song, this point is its theoretical starting alignment time), and searches forward along the time axis (towards the later part of the song) from this point in its second beat information (beat grid). The beats of the second song are checked one by one, with the goal of finding the first beat that meets the preset strong beat condition. This direction-priority logic aligns with the auditory habit that fade-in sections in transitions usually need to begin with a clear strong beat, allowing the rhythm of the second song to be presented clearly and powerfully.

[0109] Step c2: If there is a retake point that meets the preset strong shot conditions in the direction of the third party, then the retake point that meets the preset strong shot conditions in the direction of the third party is determined as the second candidate retake point.

[0110] If a strong beat that meets the preset strong beat criteria is successfully found in the search along a third direction (the main direction of the second song) that extends the playback time, this point is immediately established as the second candidate strong beat point (T_next), and the search for other directions in the second song is terminated. This means that the electronic device prioritizes the nearest strong beat in the second song after the transition point as the "entry point" of its rhythmic segment. The timestamp of this point will be recorded for comparison with candidate points in the first song and for final evaluation.

[0111] Step c3: If there is no repetition point that meets the preset strong beat conditions in the third direction, then search for repetition points in the fourth direction where the playback time decreases, and determine the repetition points that meet the preset strong beat conditions in the fourth direction as the second candidate repetition points.

[0112] The fourth direction refers to the second song. If the third-direction search fails, the search proceeds in the direction of decreasing playback time (i.e., towards the beginning of the song, an earlier time). Specifically, symmetrical to the search logic of the first song, if the search fails in the main direction (the third-direction, backward) of the second song, a backup search is initiated, shifting to the fourth direction of decreasing playback time (i.e., backtracking to an earlier point in the second song's time). Starting from T_approx, the search proceeds backward along the playback time axis of the second song, looking for the first stressed beat that meets the preset strong beat conditions. Once found, this point is designated as the second candidate stressed beat point. This ensures that even if the second song lacks suitable subsequent strong beats near the preset starting point, the electronic device can find a rhythmic anchor point from its preceding sections, thus ensuring that a candidate alignment point is always provided for the second song. This maintains the completeness of the bidirectional search and provides the necessary comparative basis for subsequently selecting the optimal alignment anchor point with the smallest deviation.

[0113] In the above implementation, for the second song to be played, a mechanism is established that prioritizes forward searching in a third direction where the playback time increases, centered on the initial transition time. This ensures that the strong beat location strategy matches the song's role as the "entering" element, better meeting the auditory requirement of finding the first strong beat to establish a rhythmic baseline when the music starts playing. If no valid strong beat is found in the third direction, the search automatically switches to a fourth direction where the playback time decreases, forming a bidirectional fault-tolerant detection logic for the second song. This orderly search strategy of forward and backward ensures that even with irregular intro structures or boundary deviations in beat markings, reliable strong beat anchors can still be stably extracted, greatly improving the alignment accuracy and connection stability between the start time of the second song and the overall transition rhythm framework.

[0114] Step S3034: Determine the first time deviation between the first candidate replay point and the initial transition time point, and the second time deviation between the second candidate replay point and the initial transition time point.

[0115] After identifying candidate beat points for the first and second songs respectively, a quantitative deviation assessment is needed to provide a basis for the final alignment decision. Specifically, this process is accomplished through simple mathematical calculations: the absolute value of the difference between the timestamp of the first candidate beat point and the timestamp of the initial transition point is calculated; this result is defined as the first time deviation. Similarly, the absolute value of the difference between the timestamp of the second candidate beat point and the timestamp of the same initial transition point is calculated; this result is defined as the second time deviation. These two deviation values ​​precisely measure the offset distance of each candidate point from the theoretical transition center point on the time axis. Their core purpose is to provide a comparable, objective numerical indicator for subsequently selecting the anchor point with the smallest deviation, ensuring that the alignment operation is as accurate as possible in time.

[0116] That is, the first time deviation (D_prev) and the second time deviation (D_next) are the absolute values ​​of the time difference between the first candidate repetition point (T_prev) and the initial transition time point (T_approx), and the second candidate repetition point (T_next) and the initial transition time point (T_approx), respectively, which are used to quantitatively evaluate the closeness between the candidate point and the preset transition point.

[0117] Step S3035: Based on the first time deviation and the second time deviation, determine the target repetition point with the smallest time deviation and that meets the preset tolerance condition from the first candidate repetition point and the second candidate repetition point, and determine the target repetition point as the alignment anchor point.

[0118] The audio control parameters include alignment anchor points, and the preset tolerance condition refers to the minimum time deviation not exceeding the preset tolerance value.

[0119] Minimum time deviation refers to the smaller of the first and second time deviations. The preset tolerance value is the upper limit of the allowed minimum time deviation, for example, 50ms. The preset tolerance condition refers to the constraint on the minimum time deviation when finally selecting the alignment anchor point. Specifically, a preset tolerance value is set, and the preset tolerance condition requires that the minimum time deviation must be less than or equal to this preset tolerance value. The preset tolerance value can be configured and adjusted based on actual audio processing accuracy and listening test results.

[0120] The alignment anchor point refers to the precise point in time at which the rhythms of the two songs (i.e., the first and second songs) are finally aligned. It replaces the initial transition point and is the core timing instruction in the audio control parameters, ensuring that the drum beats or strong beats of the two songs are precisely synchronized at this moment.

[0121] Specifically, after calculating the time deviations of the two candidate beat points, the optimal solution adjudication stage begins. First, the numerical values ​​of the first and second time deviations are directly compared, and the candidate beat point with the smaller deviation is selected. However, having only the smallest deviation is not sufficient; it must also pass a preset tolerance condition check. The electronic equipment further checks whether this candidate point with the smallest deviation meets the preset tolerance condition, that is, whether the smallest deviation does not exceed the preset tolerance value. Only candidate beat points that simultaneously meet both the conditions of having the smallest time deviation and not exceeding the preset tolerance value are ultimately determined as the target beat point. Subsequently, this target beat point is officially determined as the alignment anchor point for synchronizing the rhythms of the two songs. This adjudication mechanism ensures that the selected transition point is not only closest to the preset position in time, but also that its time adjustment range is within an audibly acceptable range, thereby guaranteeing the tightness of the rhythmic transition and the natural smoothness of the sound.

[0122] The song transition method provided in this application, based on the initial determination of the transition time point, further introduces a refined calibration mechanism based on the song's own beat information. By acquiring the beat information of each of the two songs and performing bidirectional search and positioning of candidate heavy beat points centered on the initial transition time point, the macro-level transition decision can be precisely mapped to the most rhythmic heavy beat position within the musical measure. By calculating and comparing the time deviation between each candidate heavy beat point and the initial transition time point, and selecting the final alignment anchor point based on the minimum deviation and compliance with preset tolerance conditions, micro-level alignment optimization from approximate time point to precise beat point is achieved. This process effectively avoids rhythmic inconsistencies or abrupt sounds caused by the transition point falling on a non-heavy beat position. At the same time, the tolerance protection mechanism also prevents unnatural pauses that may be introduced by over-alignment, thereby ensuring a seamless and smooth transition from the first song to the second song in terms of rhythmic connection.

[0123] In some optional implementations, step S303 further includes: if there is no beat point that meets the preset tolerance condition among the first candidate beat point and the second candidate beat point, then generate audio control parameters based on the initial transition time point.

[0124] After calculating the first and second time deviations, their magnitudes are compared to select the candidate beat point with the smallest time deviation. Next, it is determined whether this minimum deviation exceeds a preset tolerance value (e.g., 50ms). If the minimum deviation exceeds the preset tolerance value, it means that even the candidate beat point with the smallest deviation has too large a time difference from the initial transition time point, and forced alignment may result in a jarring or abrupt sound. In this case, the electronic device abandons the current rhythm alignment operation, does not generate an alignment anchor point, and directly generates subsequent audio control parameters based on the initial transition time point. If the minimum deviation does not exceed the preset tolerance value, the candidate beat point is determined as the target beat point and output as the alignment anchor point.

[0125] In the above implementation, while pursuing high-precision beat alignment, an effective fault tolerance and degradation protection mechanism is introduced. When the minimum time deviation between the searched target beat point and the initial transition time point exceeds the preset reasonable tolerance range, the forced alignment operation can be intelligently abandoned, and audio control parameters can be directly generated based on the initial transition time point. This design avoids excessive audio stretching, compression, or abrupt jumps caused by forced alignment due to excessively large beat point distances, and prevents auditory distortion caused by pursuing theoretical precision. It ensures that even in edge cases such as abnormal beat markings or extremely irregular musical structures, the transition behavior remains stable and natural, thus achieving a pragmatic and reliable balance between ensuring the ideal goal of alignment accuracy and maintaining the bottom line of auditory comfort.

[0126] Step S304: Under the audio control parameters, transition from the first song to the second song.

[0127] Specifically, the audio control parameters include the end transition point corresponding to the first song and the alignment anchor points associated with the first and second songs; the above step S304 includes: Step d1: Monitor the playback progress of the first song.

[0128] Playback progress monitoring is handled by the timeline scheduler module within the transition strategy scheduling engine. This scheduler works closely with the audio player core through an event callback mechanism. Specifically, during playback, the audio player continuously reports the current playback progress (i.e., the precise time position currently being played) at a high frequency (e.g., every millisecond or every rendered frame). The timeline scheduler registers and listens for these progress update events. Whenever a new progress event is received, the scheduler immediately retrieves the current timestamp and compares it with key trigger points pre-registered on the timeline (such as end transition points, alignment anchors, or initial transition times). This event-driven approach is more efficient and accurate than polling, ensuring that electronic devices can perceive changes in playback status in real time with low latency, providing the necessary prerequisites for triggering mixing operations at precise moments.

[0129] Step d2: When the playback progress reaches the end transition point, use audio control parameters to mix the first audio segment of the first song starting from the end transition point and the second audio segment of the second song starting from the alignment anchor point, so as to transition the first song to the second song.

[0130] When the timeline scheduler detects that the playback progress of the first song has reached the pre-calculated end transition point, it immediately sends a trigger signal. The electronic devices then execute a series of complex audio rendering operations based on the generated audio control parameters. First, the audio rendering engine is instructed to apply the fade-out effects defined in the parameters to the subsequent audio segments (the first audio segment) starting from the end transition point of the first song, such as lowering the volume according to a specified curve or adjusting the equalizer. At the same time, the instruction engine starts playing the audio segment (the second audio segment) of the second song from its alignment anchor point (or the corresponding starting point if no alignment anchor point is used), and applies the fade-in effects defined in the parameters, such as increasing the volume according to a curve, performing speed alignment, or equalizer mixing. The audio rendering engine applies the parameter array (such as the volume value array and the corresponding time array) output by the decision engine to the playback streams of the two songs in real time by calling fine-grained interfaces such as setVolume (volume envelope), setTempo (real-time speed adjustment), and setEQ (multi-band equalization). Throughout the transition window, the two audio segments are played synchronously in the underlying mixer and receive real-time parametric control, thereby achieving a smooth interweaving and connection of volume, spectrum, and rhythm, ultimately completing a seamless and coherent intelligent transition from the first song to the second song.

[0131] In some optional implementations, the transition from the first song to the second song is primarily executed by a transition strategy scheduling engine. This engine, through the collaborative work of a state machine controller and a timeline scheduler, ensures precise synchronization between the transition logic and playback progress. The state machine controller maintains a lifecycle comprising six states: Idle, Prepare, RuleCalc, PreRun, Running, and Completed. Transitions between these states are triggered by specific events, as detailed below: Idle→Prepare: Triggered by the loading of the next song. The action performed is to obtain song features and invoke the decision engine. Prepare→RuleCalc: Triggered after entering the Prepare state. The action performed is for the decision engine to calculate the transition strategy. RuleCalc→PreRun: The trigger condition is the completion of policy calculation. The execution action is to register the calculated time anchor (such as mix_start_time) to the timeline scheduler; PreRun→Running: The trigger condition is that the playback progress reaches the timeline trigger point. The action executed is to trigger the underlying audio rendering operation; Running→Completed: The trigger condition is the end of the transition. The action performed is for the system to return to the Idle state. Exception handling: In any state other than Completed, if the user switches songs or an abnormal interruption occurs, the state machine is directly reset back to the Idle state.

[0132] The timeline scheduler is responsible for monitoring the real-time playback progress of the audio player with high precision through event callbacks. When the playback progress reaches the trigger point registered in the PreRun state, the timeline scheduler immediately sends a signal to the state machine controller, driving the state flow to the Running state, and parses the time parameters in AutoMixOutput to precisely trigger the preloading of the second song, the volume / EQ change of the first song, and the start of the second song, thereby achieving seamless linkage between logic control and time triggering.

[0133] In the above implementation, by specifying the audio control parameters as the end transition point of the first song and the alignment anchor point associated with the two songs, and introducing a real-time monitoring mechanism for the playback progress of the first song, precise synchronization between the transition trigger moment and the audio segment processing is achieved. Mixing processing for the ending segment of the first song and the beginning segment of the second song is only initiated when the playback progress reaches the end transition point. This ensures that the timing of the transition operation is precisely linked to the musical structure of the first song, avoiding auditory discomfort caused by premature interruption or late intervention. Simultaneously, using the alignment anchor point as the reference position for the second song's participation in the mixing ensures that the second song can begin to integrate from the pre-selected optimal rhythm point, thereby achieving a musically orderly connection within the overlapping area of ​​the two audio segments, effectively guaranteeing the timing accuracy and auditory continuity of the transition process.

[0134] In some alternative implementations, when the playback progress reaches the time specified by the control parameters, the scheduling system in the electronic device drives the underlying audio processing engine to apply specific operations such as volume changes and speed adjustments in real time. This engine, as the executor, is typically implemented using a computer programming language to provide high-performance low-level audio processing capabilities across platforms, and performs specific rendering through fine-grained control interfaces such as setVolume (curve), setTempo (value), setEQ (band), seek (position), etc.

[0135] In some optional implementations, the real-time application mechanism of audio control parameters (such as volume, EQ, speed curve, etc.) in the rendering engine is as follows: The AutoMixOutput object output by the decision engine contains parameter arrays (such as volume value arrays) and their corresponding time point arrays defined for the two songs respectively. The rendering engine listens to the playback progress callbacks of the two songs. At each callback, based on the current playback time, it quickly finds and selects the parameter value corresponding to the nearest (or next) time point in the time array, and immediately calls interfaces such as setVolume to apply the value to the corresponding song audio stream, thereby achieving precise synchronization between parameter changes and playback progress.

[0136] In the processing involving speed changes, to ensure the priority of rhythm alignment, the rendering engine will first complete the speed change processing of the two songs (calling setTempo) to make their playback speed consistent, and then perform transition effects such as volume fade-in and fade-out and EQ adjustment, so as to achieve a smooth auditory transition while ensuring a tight rhythm.

[0137] This embodiment also provides a song transition device for implementing the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, hardware implementations, or a combination of software and hardware, are also possible and contemplated.

[0138] This embodiment provides a song transition device, such as... Figure 4 As shown, it includes: The acquisition module 401 is used to acquire the first audio features of the currently playing first song and the second audio features of the second song to be played. The determination module 402 is used to determine the target song transition strategy from a variety of preset song transition strategies based on the first audio feature and the second audio feature. The generation module 403 is used to generate audio control parameters for transitioning from the first song to the second song based on the target song transition strategy. Transition module 404 is used to transition from the first song to the second song under audio control parameters.

[0139] In some alternative implementations, the determining module 402 includes: The first acquisition submodule is used to acquire the transition priority corresponding to each preset song transition strategy; The matching submodule is used to perform feature matching between the first audio feature and the second audio feature and each preset song transition strategy based on the priority order of each transition priority representation. The first determination submodule is used to determine the first preset song transition strategy that successfully matches a feature as the target song transition strategy.

[0140] In some alternative implementations, the generation module 403 includes: The second determining submodule is used to determine the initial transition time point based on the target song transition strategy, the first audio feature, and the second audio feature. The second acquisition submodule is used to acquire the first beat information of the first song and the second beat information of the second song; The search submodule is used to search for the first candidate heavy beat point corresponding to the first song based on the first beat information, with the initial transition time point as the center, and to search for the second candidate heavy beat point corresponding to the second song based on the second beat information. The third determining submodule is used to determine the first time deviation between the first candidate repetition point and the initial transition time point, and the second time deviation between the second candidate repetition point and the initial transition time point; The fourth determination submodule is used to determine the target repetition point with the smallest time deviation and that meets the preset tolerance condition from the first candidate repetition point and the second candidate repetition point based on the first time deviation and the second time deviation, and to determine the target repetition point as the alignment anchor point. The audio control parameters include alignment anchor points, and the preset tolerance condition refers to the minimum time deviation not exceeding the preset tolerance value.

[0141] In some alternative implementations, the search submodule includes: The first search unit is used to search for a re-beat point in the first direction where the playback time decreases, with the initial transition time point as the center and based on the first beat information. The first determining unit is used to determine the repetition point in the first direction that meets the preset strong repetition conditions as the first candidate repetition point if there is a repetition point in the first direction that meets the preset strong repetition conditions. The second search unit is used to search for a repetition point in the second direction where the playback time increases if there is no repetition point that meets the preset strong beat conditions in the first direction, and to determine the repetition point that meets the preset strong beat conditions in the second direction as the first candidate repetition point. Among them, the preset strong beat condition refers to the point in time that is marked as a strong beat in the tempo information of the corresponding song.

[0142] In some optional implementations, the search submodule further includes: The third search unit is used to search for re-beat points in a third direction based on the second beat information, with the initial transition time point as the center, and to increase the playback time in a third direction. The second determining unit is used to determine the retake point that meets the preset strong retake conditions in the direction of the third party as the second candidate retake point if there is a retake point in the direction of the third party that meets the preset strong retake conditions. The fourth search unit is used to search for a repetition point in the fourth direction where the playback time decreases if there is no repetition point that meets the preset strong beat conditions in the third direction, and to determine the repetition point that meets the preset strong beat conditions in the fourth direction as the second candidate repetition point.

[0143] In some optional implementations, the second determining submodule includes: The third determining unit is used to determine the end transition point of the first song under the target transition strategy based on the first audio features, and to determine the start transition point of the second song under the target transition strategy based on the second audio features. The fourth determining unit is used to determine the initial transition time point based on the end transition point and the start transition point; The audio control parameters include the end transition point and the start transition point.

[0144] In some alternative implementations, the generation module 403 further includes: The generation submodule is used to generate audio control parameters based on the initial transition time point if the minimum time deviation exceeds the preset tolerance value.

[0145] In some optional implementations, the audio control parameters include the end transition point corresponding to the first song and alignment anchor points associated with the first and second songs; the transition module 404 includes: The listening submodule is used to monitor the playback progress of the first song; The mixing submodule is used to mix the first audio segment of the first song starting from the end transition point and the second audio segment of the second song starting from the alignment anchor point when the playback progress reaches the end transition point, so as to transition the first song to the second song.

[0146] In some optional implementations, the first audio feature includes multiple end-of-song structure points, and the second audio feature includes multiple start-of-song structure points; the matching submodule includes: The judgment unit is used to determine, for any preset song transition strategy, whether there is a target end-type structure point required by the preset song transition strategy among the multiple end-type structure points of the first song, and whether there is a target start-type structure point required by the preset song transition strategy among the multiple start-type structure points of the second song.

[0147] In some alternative implementations, the end-of-class structure point is determined by at least one of the following methods: The moment when a human voice ends, with the distance from the end of the audio not exceeding the preset transition window duration and the ending energy exceeding the first preset energy threshold, is defined as the human voice end point. The drum beat ending point is defined as the drum beat ending point when the distance between the drum beat and the end of the audio is no more than the duration of the preset transition window and the ending energy exceeds the second preset energy threshold. When structural paragraph information exists and the starting point of the last paragraph is not earlier than the already determined end point of the vocal or drum beats, the starting point of the last paragraph is determined as the end point of the structural paragraph. Within the preset number of measures, the end position of the measure that meets the duration limit and is no earlier than the already determined end point of the vocals or the end point of the drumbeat is determined as the end point of the music measure; Among them, the ending structural points include at least one of the following: vocal ending point, drum ending point, structural paragraph ending point, and musical measure ending point.

[0148] In some alternative implementations, the starting class structure point is determined by at least one of the following methods: The moment when a human voice begins within a preset transition window and whose initial energy exceeds a first preset energy threshold is defined as the human voice starting point. The drum start time is defined as the drum start point when the drumbeat is within the preset transition window duration and the initial energy exceeds the second preset energy threshold. When structural paragraph information exists and the end point of the first paragraph is earlier than the already determined starting point of the vocals or the starting point of the drumbeat, the end point of the first paragraph is determined as the starting point of the structural paragraph. Within the preset number of measures, the starting point of the measure that meets the duration limit and precedes the already determined starting point of the vocals or drum beats will be determined as the starting point of the music measure. Among them, the starting structural points include at least one of the following: vocal starting point, drum starting point, structural paragraph starting point, and musical measure starting point.

[0149] The song transition device provided in this application embodiment can execute the song transition method provided in any embodiment of this application, and has the corresponding functional modules and beneficial effects of the method execution. Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.

[0150] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.

[0151] The following is a detailed reference. Figure 5The diagram illustrates a structural schematic suitable for implementing the electronic device described in the embodiments of this application. The electronic device may include a processor (e.g., a central processing unit, graphics processor, etc.) 501, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 502 or a program loaded from memory 508 into random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of the electronic device. The processor 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0152] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; memory devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 5 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown, and more or fewer devices may be implemented or have instead.

[0153] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a memory 508, or installed from a ROM 502. When the computer program is executed by the processor 501, it performs the functions defined in the song transition method of embodiments of this application.

[0154] Figure 5 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0155] This application also provides a computer-readable storage medium. The methods described above according to this application can be implemented in hardware or firmware, or implemented as recordable on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code. When the software or computer code is accessed and executed by the computer, processor, or hardware, the song transition method shown in the above embodiments is implemented.

[0156] A portion of this application can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to this application through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.

[0157] Although embodiments of this application have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of this application, and all such modifications and variations fall within the scope defined by the appended claims.

Claims

1. A song transition method, characterized in that, The method includes: Obtain the first audio feature of the currently playing first song, and the second audio feature of the second song to be played; Based on the first audio feature and the second audio feature, the target song transition strategy is determined from a variety of preset song transition strategies; Based on the target song transition strategy, audio control parameters are generated for the transition from the first song to the second song; Under the aforementioned audio control parameters, the first song is transitioned to the second song.

2. The method according to claim 1, characterized in that, The step of determining the target song transition strategy from multiple preset song transition strategies based on the first audio feature and the second audio feature includes: Obtain the transition priority corresponding to each of the preset song transition strategies; Based on the priority order of each transition priority representation, the first audio feature and the second audio feature are respectively matched with each preset song transition strategy; The preset song transition strategy that is the first feature to be successfully matched is determined as the target song transition strategy.

3. The method according to claim 1, characterized in that, The step of generating audio control parameters for transitioning from the first song to the second song based on the target song transition strategy includes: The initial transition time point is determined based on the target song transition strategy, the first audio feature, and the second audio feature; Obtain the first beat information of the first song and the second beat information of the second song; Centered on the initial transition time point, the first candidate heavy beat point corresponding to the first song is obtained by searching based on the first beat information, and the second candidate heavy beat point corresponding to the second song is obtained by searching based on the second beat information; Determine the first time deviation between the first candidate repetition point and the initial transition time point, and the second time deviation between the second candidate repetition point and the initial transition time point; Based on the first time deviation and the second time deviation, the target repetition point with the smallest time deviation and meeting the preset tolerance condition is determined from the first candidate repetition point and the second candidate repetition point, and the target repetition point is determined as the alignment anchor point. The audio control parameters include the alignment anchor point, and the preset tolerance condition refers to the minimum time deviation not exceeding the preset tolerance value.

4. The method according to claim 3, characterized in that, The step of searching for the first candidate heavy beat point corresponding to the first song based on the first beat information, centered on the initial transition time point, includes: Centered on the initial transition time point, a replay point is searched in the first direction where the playback time decreases, based on the first beat information; If there is a repetition point in the first direction that meets the preset strong shot conditions, then the repetition point in the first direction that meets the preset strong shot conditions is determined as the first candidate repetition point. If there is no repetition point that meets the preset strong beat condition in the first direction, then search for repetition points in the second direction where the playback time increases, and determine the repetition point that meets the preset strong beat condition in the second direction as the first candidate repetition point; The preset strong beat condition refers to the fact that the downbeat point is marked as a strong beat in the rhythm information of the corresponding song.

5. The method according to claim 3 or 4, characterized in that, Centered on the initial transition time point, the second candidate heavy beat point corresponding to the second song is searched based on the second beat information, including: Centered on the initial transition time point, search for replay points in a third direction based on the second beat information as the playback time increases; If there is a retake point that meets the preset strong shot conditions in the direction of the third party, then the retake point that meets the preset strong shot conditions in the direction of the third party is determined as the second candidate retake point. If no repetition point that meets the preset strong beat condition is found in the third direction, then a repetition point is searched in the fourth direction where the playback time decreases, and the repetition point that meets the preset strong beat condition in the fourth direction is determined as the second candidate repetition point.

6. The method according to claim 3, characterized in that, The step of determining the initial transition time point based on the target song transition strategy, the first audio feature, and the second audio feature includes: The end transition point of the first song under the target transition strategy is determined based on the first audio feature, and the start transition point of the second song under the target transition strategy is determined based on the second audio feature. The initial transition time point is determined based on the end transition point and the start transition point; The audio control parameters include the end transition point and the start transition point.

7. The method according to claim 3, characterized in that, The method further includes: If the minimum time deviation exceeds the preset tolerance value, the audio control parameters are generated based on the initial transition time point.

8. The method according to claim 1, characterized in that, The audio control parameters include the end transition point corresponding to the first song and the alignment anchor point associated with the first song and the second song; The step of transitioning from the first song to the second song under the audio control parameters includes: Monitor the playback progress of the first song; When the playback progress reaches the end transition point, the audio control parameters are used to mix the first audio segment of the first song starting from the end transition point and the second audio segment of the second song starting from the alignment anchor point, so as to transition the first song to the second song.

9. The method according to claim 2, characterized in that, The first audio feature includes multiple end-type structural points of the first song, and the second audio feature includes multiple start-type structural points of the second song. The first audio feature and the second audio feature are matched with each of the preset song transition strategies, including: For any of the preset song transition strategies, determine whether there is a target end-of-song transition strategy required by the preset song transition strategy among the multiple end-of-song structure points of the first song, and whether there is a target start-of-song transition strategy required by the preset song transition strategy among the multiple start-of-song structure points of the second song.

10. The method according to claim 9, characterized in that, The termination structure point is determined by at least one of the following methods: The moment when a human voice ends, with the distance from the end of the audio not exceeding the preset transition window duration and the ending energy exceeding the first preset energy threshold, is defined as the human voice end point. The drum beat ending point is defined as the drum beat ending point when the distance between the drum beat and the end of the audio is no more than the preset transition window duration and the ending energy exceeds the second preset energy threshold. When structural paragraph information exists and the starting point of the last paragraph is not earlier than the already determined end point of the vocal or drum beat, the starting point of the last paragraph is determined as the end point of the structural paragraph. Within the preset number of measures, the end position of the measure that meets the duration limit and is no earlier than the already determined end point of the vocals or the end point of the drumbeat is determined as the end point of the music measure; The term "ending structural point" includes at least one of the following: the ending point of the vocal part, the ending point of the drumbeat, the ending point of the structural paragraph, and the ending point of the musical measure.

11. The method according to claim 9, characterized in that, The initial class structure point is determined by at least one of the following methods: The moment when a human voice begins within a preset transition window and whose initial energy exceeds a first preset energy threshold is defined as the human voice starting point. The drumbeat start time is determined as the drumbeat start point when the drumbeat starts within the preset transition window duration and the starting energy exceeds the second preset energy threshold. When structural paragraph information exists and the end point of the first paragraph is earlier than the already determined starting point of the human voice or the starting point of the drumbeat, the end point of the first paragraph is determined as the starting point of the structural paragraph. Within the preset number of measures, the starting point of the measure that meets the duration limit and precedes the already determined starting point of the vocals or drum beats will be determined as the starting point of the music measure. The starting structural point includes at least one of the following: the vocal starting point, the drum starting point, the structural paragraph starting point, and the musical measure starting point.

12. A song transition device, characterized in that, The device includes: The acquisition module is used to acquire the first audio features of the currently playing first song and the second audio features of the second song to be played. The determining module is used to determine the target song transition strategy from a variety of preset song transition strategies based on the first audio feature and the second audio feature; The generation module is used to generate audio control parameters for transitioning from the first song to the second song based on the target song transition strategy. A transition module is used to transition the first song to the second song under the audio control parameters.

13. A computer device, characterized in that, include: A memory and a processor are communicatively connected, the memory storing computer instructions, and the processor executing the computer instructions to perform the song transition method according to any one of claims 1 to 11.

14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to perform the song transition method according to any one of claims 1 to 11.

15. A computer program product, characterized in that, Includes computer instructions for causing a computer to perform the song transition method according to any one of claims 1 to 11.