Audio playing method and system, electronic equipment and storage medium

By collecting real-time transmission performance status data of audio source nodes through the playback terminal and dynamically selecting alternative nodes by the scheduling server, the latency and buffering problems in existing audio playback systems are solved, and the continuity and stability of audio playback are achieved.

CN121151402APending Publication Date: 2025-12-16LINKPLAY TECHNOLOGY INC NANJING
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511119273.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-12-16

AI Technical Summary

Technical Problem

Existing audio playback systems rely on static configuration or proximity strategies when selecting audio sources, which can easily lead to problems such as latency, frequent buffering, and unstable sound quality at the playback terminal.

Method used

By collecting real-time transmission performance status data of audio source nodes through the playback terminal, calculating performance levels using the scheduling server, dynamically selecting alternative audio source nodes, and seamlessly switching when performance degrades, the continuity and stability of audio playback are achieved.

Benefits of technology

It improves the continuity and stability of audio playback, avoids delays and buffering, and ensures stable and smooth sound quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121151402A_ABST
    Figure CN121151402A_ABST
Patent Text Reader

Abstract

The invention provides an audio playing method and system, electronic equipment and a storage medium. The method is applied to the audio playing system, comprises a playing terminal and a scheduling server, and comprises the following steps: the playing terminal transmits transmission performance state data of a current sound source node to the scheduling server in a process of playing audio data transmitted by the current sound source node; the scheduling server performs transmission performance grade calculation on the transmission performance state data to obtain a current performance grade; the scheduling server determines an alternative sound source node according to the current performance grade, and transmits attribute information and transmission performance state data of the alternative sound source node to the playing terminal; the playing terminal determines a target sound source node according to the attribute information and the transmission performance state data in response to triggering of a preset performance grade degradation event; and the playing terminal is switched to obtain the audio data from the target sound source node, and switching playing of the audio source is carried out. According to the invention, the continuity and stability of audio playing can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of audio processing technology, and in particular to an audio playback method, system, electronic device, and storage medium. Background Technology

[0002] Currently, most audio playback systems, such as streaming media players, smart speakers, and audio services in mobile terminals, typically use Content Delivery Networks (CDNs) or edge nodes to distribute content.

[0003] However, this implementation method usually relies on static configuration or a simple proximity strategy when selecting audio sources, which makes the playback terminal prone to delays, frequent buffering, and unstable sound quality. There is an urgent need for a technical solution that can improve the continuity and stability of audio playback. Summary of the Invention

[0004] In view of this, the purpose of this disclosure is to provide an audio playback method, system, electronic device, and storage medium to improve the continuity and stability of audio playback.

[0005] In a first aspect, embodiments of this disclosure provide an audio playback method applied to an audio playback system, the audio playback system including a playback terminal and a scheduling server. The method includes: during the playback of audio data transmitted by a current audio source node, the playback terminal transmits the transmission performance status data of the current audio source node to the scheduling server; the scheduling server calculates the transmission performance level of the transmission performance status data to obtain the current performance level of the current audio source node; the scheduling server determines at least one candidate audio source node based on the current performance level, and transmits the attribute information and transmission performance status data of the candidate audio source node to the playback terminal; in response to the triggering of a preset performance level degradation event, the playback terminal determines a target audio source node from the candidate audio source nodes based on the attribute information and transmission performance status data of the candidate audio source nodes; the playback terminal switches to acquiring audio data from the target audio source node, and performs audio source switching playback based on the acquired audio data.

[0006] Secondly, embodiments of this disclosure provide an audio playback system, the system including a playback terminal and a scheduling server, wherein the playback terminal is configured to: transmit the transmission performance status data of the current audio source node to the scheduling server during the playback of audio data transmitted by the current audio source node; the scheduling server is configured to: calculate the transmission performance level of the transmission performance status data to obtain the current performance level of the current audio source node; the scheduling server is configured to: determine at least one candidate audio source node based on the current performance level, and transmit the attribute information and transmission performance status data of the candidate audio source node to the playback terminal, wherein the candidate audio source node has a transmission performance level at least higher than the current performance level; the playback terminal is configured to: in response to the triggering of a preset performance level degradation event, determine a target audio source node from the candidate audio source nodes based on the attribute information and transmission performance status data of the candidate audio source nodes; the playback terminal is configured to: switch to acquiring audio data from the target audio source node, and switch playback of the audio source based on the acquired audio data.

[0007] Thirdly, embodiments of this disclosure provide an electronic device, including a processor and a memory, wherein the memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the above-described audio playback method.

[0008] Fourthly, embodiments of this disclosure provide a computer-readable storage medium storing computer-executable instructions. When the computer-executable instructions are invoked and executed by a processor, the computer-executable instructions cause the processor to implement the aforementioned audio playback method.

[0009] The embodiments disclosed herein bring the following beneficial effects:

[0010] The aforementioned audio playback method, system, electronic device, and storage medium measure the transmission performance of each audio source node by collecting node transmission performance status data from the playback terminal. The scheduling server uniformly calculates and generates alternative audio source nodes for the playback terminal, realizing a performance feedback linkage mechanism with the upper-layer audio application layer protocol. When the playback performance of the playback terminal degrades, a fast and seamless switch can be performed, thereby avoiding delays, buffering, and reduced sound quality, thus improving the continuity and stability of audio playback.

[0011] Other features and advantages of this disclosure will be set forth in the following description and will be apparent in part from the description or may be learned by practicing the disclosure. The objects and other advantages of this disclosure are realized and obtained through the structures particularly pointed out in the description, claims and drawings.

[0012] To make the above-mentioned objects, features and advantages of this disclosure more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0013] To more clearly illustrate the technical solutions in the specific embodiments of this disclosure or the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0014] Figure 1 This is a flowchart of one embodiment of the audio playback method in this disclosure;

[0015] Figure 2 A schematic diagram of an audio playback system provided in an embodiment of this disclosure;

[0016] Figure 3 This is a schematic diagram of an electronic device provided in an embodiment of the present disclosure. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0018] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in this disclosure, claims, and accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” or “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0019] For ease of understanding, the specific process of the embodiments of this disclosure is described below. Please refer to [link / reference]. Figure 1 This disclosure applies to an audio playback system, which includes a playback terminal and a scheduling server. One embodiment of the audio playback method in this disclosure includes:

[0020] Step S10: During the playback of audio data transmitted by the current audio source node, the playback terminal transmits the transmission performance status data of the current audio source node to the scheduling server.

[0021] A playback terminal is a device used to play audio in an audio playback system. Specifically, in one implementation, the playback terminal receives audio data transmitted from an audio source node, caches the received audio data through a preset caching mechanism, and plays the cached audio data, allowing the playback terminal to download and play audio data simultaneously. The playback terminal can be any device capable of receiving, decoding, and playing audio data, such as a smart speaker, a mobile app, or a streaming media player; no specific limitations are specified here.

[0022] Based on the above, during the process of playing audio data transmitted by the current audio source node, the playback terminal actually plays the audio data pre-transmitted by the current audio source node in the cache according to the preset transmission frame rate. The playback terminal downloads audio data from the current audio source node according to the above transmission frame rate and saves it in the cache space. The transmission frame rate can be dynamically calculated, and the cache space is managed by the above-mentioned caching mechanism. The specifics are not limited here.

[0023] The audio source node is a server or content delivery network (CDN) node that stores the audio data to be played by the playback terminal. All audio nodes in this embodiment can be server nodes in a self-built server network, edge nodes provided by CDN services, or network nodes in any other distribution mode. No specific limitation is made here.

[0024] The current audio source node can also be the target audio source node determined by the audio playback method provided in this embodiment when the preset performance level degradation event was triggered last time. Steps S10-S30 provided in this embodiment can be steps that are repeatedly executed at a preset frequency. Steps S10-S30 are executed in each execution cycle. The preset frequency can be the same as the above transmission frame rate, or it can be dynamically calculated or preset in other ways. The specific frequency is not limited here.

[0025] Specifically, when the playback terminal transmits the audio data transmitted by the current audio source node to the scheduling server, it can transmit the transmission performance status data of the current audio source node within the current period to the scheduling server while playing the audio data transmitted by the current audio source node in the current period. This allows the transmission performance status data of the current audio source node to be updated periodically (near real-time) and reported to the scheduling server for decision-making on alternative audio source nodes.

[0026] The scheduling server can be any server node in the self-built server network or content delivery network mentioned above, or it can be a designated server independent of the aforementioned network; the specifics are not limited here. In this embodiment, the scheduling server uniformly receives node transmission performance status data uploaded by the playback terminal, performs unified analysis and decision-making, and provides candidate audio source nodes to the terminal device. This enables the upper-layer audio application layer of the system to also link with the lower-layer decision layer for performance feedback, thereby greatly improving the continuity and stability of audio playback.

[0027] Transmission performance status data refers to data that reflects the network transmission quality between the current audio source node and the playback terminal. It may include, but is not limited to, round-trip time (RTT), packet loss rate, bitstream stability, audio buffering times and duration, first frame audio startup time, etc. It can be used to evaluate the performance of the audio source, and no specific limitation is made here.

[0028] Step S20: The scheduling server calculates the transmission performance level of the transmission performance status data to obtain the current performance level of the current audio source node;

[0029] In this embodiment, after the scheduling server receives the transmission performance status data of the current audio source node, it can calculate the transmission performance level of the current audio source node through a preset evaluation algorithm or model, thereby determining the current performance level of the current audio source node. This level can be used to represent the transmission performance level of the current audio source node in the current period, so that the performance of the current audio source node can be quantitatively represented by the level, thereby generating comparability.

[0030] As an example, and not a limitation, when calculating transmission performance levels, each parameter in the transmission performance status data can be used as a transmission performance status indicator. According to the grading standard corresponding to each transmission performance status indicator, the value of each transmission performance status indicator in the transmission performance status data of the current audio source node can be divided into the corresponding level to obtain the level corresponding to each transmission performance status indicator. Then, through a preset fusion algorithm, such as a weighted superposition algorithm, the levels corresponding to all transmission performance status indicators are fused to obtain the current performance level of the current audio source node. Each transmission performance status indicator can correspond to a weight.

[0031] In one implementation, the weight corresponding to each transmission performance status indicator can be preset or dynamically calculated. In dynamic calculation, factors such as the importance score of the indicator based on user feedback and the immediate impact on the current playback quality can also be incorporated. Specifically, the weight corresponding to each transmission performance status indicator can be determined by the following algorithm:

[0032] Wi(t)=Wi(t-1)·α+(1-α)·[β·li(t)+(1-β)·Fi(t)]

[0033] Where Wi(t) represents the weight of the i-th transmission performance status indicator at time t, Wi(t-1) represents the weight at the previous time, α represents a smoothing factor (0.7-0.9) used to control the retention of historical weights, β represents an importance balancing factor (0.6-0.8) used to balance immediate impact and user feedback, li(t) represents the immediate impact of the i-th transmission performance status indicator on the current playback quality, and Fi(t) represents the importance score of the indicator in user feedback. This implementation method achieves dynamic adaptive adjustment of weights, which can improve the accuracy of performance evaluation and the system's adaptability to different network environments.

[0034] In one implementation, the grading standard corresponding to each transmission performance status indicator can be fixed or dynamic. Specifically, it can be adaptively adjusted according to the network environment in which the playback terminal is located. For example, in a Wi-Fi environment, the threshold for RTT at a certain level can be 100ms, while in a 4G environment, the threshold for that level can be relaxed to 150ms to improve the accuracy of judgment and avoid frequent level switching.

[0035] Step S30: The scheduling server determines at least one alternative audio source node based on the current performance level, and transmits the attribute information and transmission performance status data of the alternative audio source node to the playback terminal.

[0036] In this embodiment, the scheduling server selects alternative audio source nodes that meet preset conditions from the aforementioned server network or CDN network based on the current performance level of the current audio source node. These alternative audio source nodes are then sent to the playback terminal along with their attribute information and transmission performance status data. The playback terminal can then select the alternative audio source node when needed, thereby maintaining a high-quality audio data source and improving the continuity and stability of audio playback.

[0037] In one implementation, the alternative audio source node can be an audio source node whose transmission performance level is higher than the current performance level of the current audio source node, or an audio source node with the same transmission performance level as the current audio source node, or an audio source node determined by other methods / conditions, which are not limited here.

[0038] After determining the candidate audio source nodes, the scheduling server sends the candidate audio source node's attribute information and transmission performance status data to the terminal. The attribute information of the candidate audio source node may include information required for the playback terminal to connect or determine whether to connect to the corresponding candidate audio source node, such as identifier, IP address, geographical location, carrying capacity, current load rate, etc., which are not limited here.

[0039] When determining candidate audio source nodes, the geographical locations of the audio source nodes and the playback terminal can be compared. Audio source nodes that are closer to the playback terminal are more suitable as candidate audio source nodes. Furthermore, factors such as the carrying capacity of the audio source nodes and their current load rate can be considered to calculate a score indicating suitability as a candidate audio source node. The candidate audio source nodes are then determined based on these scores. The score calculation can use the following algorithm:

[0040]

[0041] Where S(j) represents the score of the j-th audio source node, Wi represents the weight of the ith transmission performance status index, n equals the total number of transmission performance status indices, P(i,j) represents the normalized performance score of audio source node j on transmission performance status index i, d(j) represents the geographical distance between audio source node j and the playback terminal, Dmax represents the preset maximum effective distance, L(j) represents the current load rate of audio source node j, and Lmax represents the maximum carrying capacity of the audio source node. This implementation comprehensively considers performance, geographical location, and load balancing to improve the scientific nature of candidate node selection and the stability after switching.

[0042] The scheduling server maintains a real-time mapping database between Domain Name System (DNS) server IPs and audio source nodes, recording their respective access quality metrics (RTT, stability, TTL, etc.), including node ID, IP, performance level, weight, and last update time. This information is obtained by periodically collecting connection quality data from each DNS to the audio source node, actual user access performance feedback, and data records from third-party monitoring services. When determining candidate audio source nodes, the server can quickly find audio source nodes that meet the access quality metrics criteria from the aforementioned mapping database, which are then used as candidate audio source nodes for the current audio source node.

[0043] Step S40: In response to the triggering of the preset performance level degradation event, the playback terminal determines the target audio source node from the candidate audio source nodes based on the attribute information and transmission performance status data of the candidate audio source nodes.

[0044] A preset performance level degradation event refers to an event that causes the transmission performance level of the current audio source node to decrease from its current performance level. The preset performance level degradation event may not be an instantaneous event, but may be an event triggered after a certain period of accumulated monitoring and the achievement of certain conditions.

[0045] For example, "the current audio source node's RTT growth rate reaches 100% in the current period" can be considered as a preset performance level degradation event that occurs instantaneously. On the other hand, "the playback terminal experiences two consecutive buffering events with a total duration of more than 5 seconds", "RTT growth exceeds 50% and the duration exceeds 30 seconds", "packet loss rate exceeds 5% and the duration exceeds 10 seconds", and "bitrate fluctuation coefficient is greater than 0.3" can be considered as preset performance level degradation events that are triggered after accumulating a certain monitoring time and meeting certain conditions. The specifics are not limited here.

[0046] After the preset performance level degradation event is triggered, the playback device selects the target audio source node that best matches the playback terminal and is most suitable to replace the current audio source node from the candidate audio source nodes based on the attribute information and transmission performance status data of the candidate audio source nodes. This target audio source node is then selected as the audio source node to be switched to by the playback terminal, so that the playback terminal can always be connected to the optimal audio source node to maintain the continuity and stability of audio playback.

[0047] In this embodiment, after the performance degradation event is triggered, the playback terminal does not directly request the default DNS resolution. Instead, it uses the locally cached list of candidate audio source nodes or requests a new list from the scheduling server. The playback terminal can prioritize the candidate audio source nodes according to its own performance-aware DNS response sorting and caching strategy, and select the node with the higher performance level as the target audio source node.

[0048] When determining the target audio source node, one can use the geographical location information in the attribute information, the connection success rate determined by the transmission performance status data, the average RTT, etc., or the user satisfaction rating of the candidate audio source node, etc., but the specifics are not limited here.

[0049] In one implementation, a preset performance degradation event can be triggered by calculating the performance degradation trend score of the current audio source node. When the performance degradation trend score is greater than a preset score threshold, the preset performance degradation event is triggered, causing the playback terminal to respond to the triggering of the preset performance degradation event and determine the target audio source node from the candidate audio source nodes based on the attribute information and transmission performance status data of the candidate audio source nodes. The performance degradation trend score Td can be calculated using the following algorithm:

[0050]

[0051] Where m equals the total number of transmission performance status indicators, γ i X represents the degradation sensitivity coefficient of the i-th transmission performance status index. i (t) represents the value of transmission performance status index i at the current moment, X i (tk) represents the value of transmission performance status index i k cycles ago, σ i θ represents the historical standard deviation of the transmission performance status index i. i This represents the preset degradation threshold for transmission performance status indicator i. This implementation can achieve early warning and predictive switching of performance degradation, avoiding noticeable playback interruptions or quality degradation perceived by the user.

[0052] Step S50: The playback terminal switches to obtain audio data from the target audio source node, and switches the audio source for playback based on the obtained audio data.

[0053] After the target audio source node is determined, the playback terminal can establish a connection with the target audio source node and start receiving audio data streams from the target audio source node. The audio data streams received by the playback terminal can be seamlessly integrated into the playback queue to avoid playback interruption and achieve a smooth transition and seamless switching from the old audio source to the new audio source.

[0054] In one implementation, if the playback terminal fails to switch audio sources, it rolls back to the most stable historical node for audio data transmission. Switching failures can be determined after a certain number of attempts. For example, the client sets a failure threshold of two attempts; if attempts to switch to nodes D and E both fail, it rolls back to node B. A cooldown period can also be set for failed nodes. During the cooldown period, all playback terminals will not actively select that node, but it can still be used as a last resort if all other nodes fail. After the cooldown period, the node can be re-added to the candidate pool as a backup audio source node.

[0055] The audio playback method provided by the above embodiments measures the transmission performance of each audio source node by collecting node transmission performance status data actually by the playback terminal. The scheduling server uniformly calculates and generates alternative audio source nodes for the playback terminal, realizing a performance feedback linkage mechanism with the upper-layer audio application layer protocol. When the playback performance of the playback terminal degrades, it can quickly and seamlessly switch, thereby avoiding the occurrence of delays, buffering, and reduced sound quality, thus improving the continuity and stability of audio playback.

[0056] Next, we will explain the specific methods for audio playback.

[0057] In one implementation, the step of transmitting the transmission performance status data of the current audio source node to the scheduling server during the playback of audio data transmitted by the current audio source node includes: During the playback of audio data transmitted by the current audio source node, the playback terminal periodically collects raw network metrics of the network layer based on the connection between the playback terminal and the current audio source node, wherein the raw network metrics include round-trip propagation delay and / or packet loss rate; periodically collects playback status metrics of the playback terminal through the player of the playback terminal, wherein the playback status metrics include buffer event parameters of the player and / or first frame audio playback time parameters; combines the raw network metrics and playback status metrics of the same period to obtain the transmission performance status data of the current audio source node within the same period; and periodically transmits the transmission performance status data of the current audio source node within the most recent period to the scheduling server.

[0058] In this embodiment, the transmission performance status data of the current audio source node may include multiple transmission performance status indicators. The original network indicators of the network layer and the playback status indicators of the application layer can each be used as one or more transmission performance status indicators. This embodiment adopts a layered data acquisition method to establish a performance feedback linkage mechanism between the lower layer and the upper layer audio application layer protocol, so that the network quality assessment is more in line with the actual situation.

[0059] The playback terminal may include an Agent module. The Agent module periodically obtains basic network communication quality data, i.e. raw network metrics, between the current audio source node and the network protocol layer through detection or monitoring. The lightweight Agent module built into the playback terminal can continuously listen to and record all network requests and responses with the current audio source node.

[0060] Round-trip propagation delay (RTT) refers to the time required for a data packet to travel from the sender (playback terminal) to the receiver (audio source node) and back to the sender. It directly reflects the real-time latency of network transmission and is a key indicator for measuring network performance. Specifically, TRT can be measured by the time difference between the transmission control protocol (TCP) handshake and the sending of UDP probe packets and the return of acknowledgment (ACK) packets.

[0061] In one implementation, the Agent module can measure RTT through the three-way handshake process during TCP connection establishment, or periodically send small UDP probe packets to the audio source node and record the time difference between sending and receiving an ACK. For example, after connecting to audio source node A, the playback terminal records all network requests and responses generated during audio playback within 30 seconds and obtains the average RTT through TCP / UDP log analysis.

[0062] Packet loss rate refers to the proportion of data packets lost during network transmission out of the total number of data packets sent. A higher packet loss rate results in more incomplete audio data received by the playback terminal, affecting playback quality. Specifically, the packet loss rate can be obtained by statistically analyzing the ratio of sent packets to acknowledgment packets, using TCP ACK or UDP heartbeat detection; details will not be elaborated here.

[0063] In one implementation, the Agent module can analyze TCP / UDP protocol stack logs to calculate packet loss rate. For example, it can record the ratio of TCP retransmissions to the total number of packets, or for UDP streams, it can calculate the packet loss rate by periodically sending heartbeat packets and checking the number of received acknowledgment packets. For instance, a smart speaker sends a heartbeat packet to a backup audio source every 30 seconds and collects the packet loss rate of the returned ACK packets.

[0064] The playback status indicators at the application layer can be collected by the player used by the playback terminal to play the above audio data. Among them, the buffering event parameters can include the number of times the player pauses playback due to insufficient data during playback and the duration of each buffering. The first frame audio playback time parameter refers to the time required from the playback terminal triggering a playback request to the first audio frame being successfully decoded and played.

[0065] Understandably, playback status metrics directly reflect the user's actual playback experience. By combining buffering events and / or first frame time, the playback quality perceived by the user can be accurately assessed, making up for the shortcomings of pure network metrics and enabling the system to more comprehensively understand and optimize the user experience.

[0066] After periodically collecting raw network metrics and playback status metrics, the collected data can be periodically transmitted to the scheduling server. Specifically, raw network metrics and playback status metrics for the same period can be packaged and transmitted to the scheduling server, for example, these transmission performance status data can be reported to the scheduling server every 30 seconds.

[0067] In one implementation, the transmission performance status data includes multiple transmission performance status indicators; the step of the scheduling server calculating the transmission performance level of the transmission performance status data to obtain the current performance level of the current audio source node includes: the scheduling server determining the transmission performance level threshold corresponding to each transmission performance status indicator based on the network connection status attribute of the playback terminal, wherein the transmission performance level threshold includes the numerical range corresponding to multiple transmission performance levels respectively; comparing each transmission performance status indicator in the transmission performance status data with the numerical range corresponding to each transmission performance level respectively to determine the first performance level to which each transmission performance status indicator of the current audio source node belongs; and determining the current performance level of the current audio source node based on the first performance level.

[0068] In this embodiment, a mechanism is introduced to dynamically adjust the threshold based on the network connection status of the playback terminal to obtain a more accurate performance level. This makes the evaluation of the network transmission quality of the audio source node more consistent with the actual network environment of the playback terminal, and makes the level evaluation more adaptable.

[0069] The network connection status attributes of the playback terminal can include the connection type (such as Wi-Fi, 4G / 5G cellular network), network bandwidth, signal strength, geographical location, etc. These attributes can be actively collected by the playback terminal and uploaded to the scheduling server. The playback terminal can also query these attributes through the system interface. Details will not be elaborated here.

[0070] In this embodiment, the transmission performance status data includes multiple transmission performance status indicators. Each transmission performance status indicator corresponds to a level classification method, which is a transmission performance level threshold. Each level classification method includes multiple pre-divided transmission performance levels and their corresponding numerical ranges, which are used to evaluate the transmission performance level to which the corresponding transmission performance status indicator belongs.

[0071] For example, when RTT is used as one of the transmission performance status indicators in the transmission performance status data, the threshold values ​​for its transmission performance level, in descending order of transmission performance level, can be: [0,100], [101,200], [201,250], [251,+∞). By comparing the RTT value in the transmission performance status data of the current audio source node with the above value range using this level classification method, the first performance level of RTT can be determined.

[0072] Suppose that the RTT value in the transmission performance status data of the current audio source node is 150, which belongs to [101, 200], that is, the second value in the above range. Then, the first performance level of the RTT transmission performance status index is the performance level corresponding to [101, 200], which is the second best level. The specific level is not limited here.

[0073] Specifically, the transmission performance level threshold corresponding to the above RTT can be dynamically changed according to the network connection status. For example, the optimal level threshold can be less than 90ms in a Wi-Fi environment, while it can be relaxed to 150ms in a 4G environment. That is to say, when the connection type of the playback terminal is Wi-Fi, the above value range of [0,100] can be adjusted to [0,90], and the value range of other levels is adjusted accordingly. If the connection type of the playback terminal is 4G, the above value range of [0,100] can be adjusted to [0,150], and the value range of other levels is adjusted accordingly. The specifics are not limited here.

[0074] After determining the first performance level, the comprehensive performance level of the current sound source node can be determined by a preset comprehensive evaluation algorithm and used as the current performance level. Alternatively, the average value of all first performance levels can be used as the current performance level, or the lowest or highest value among all first performance levels can be used as the current performance level. Other methods can also be used to determine the current performance level, which will not be elaborated here.

[0075] In one implementation, the step of determining the current performance level of the current audio source node based on the first performance level includes: the scheduling server weighting and superimposing the first performance level to which each transmission performance status indicator of the current audio source node belongs according to the preset weights corresponding to each transmission performance status indicator, to obtain the current performance level of the current audio source node.

[0076] In this embodiment, different transmission performance status indicators can correspond to different preset weights, thereby realizing the difference in importance of different performance indicators in the overall evaluation, making the determination of performance levels more precise and reasonable, and thus making more optimized scheduling decisions.

[0077] The preset weights corresponding to each transmission performance status indicator can be pre-set based on experience, user feedback, or machine learning models. For example, packet loss rate and buffering events can be given higher weights because they have a more direct and negative impact on the user's playback experience.

[0078] In one implementation, the step of the scheduling server determining at least one candidate audio source node based on the current performance level includes: the scheduling server obtaining the location information of the playback terminal and the preset audio source nodes, determining at least one first audio source node from the preset audio source nodes by matching the location information; and determining the first audio source node whose latest performance level is greater than or equal to the current performance level of the current audio source node as a candidate audio source node based on the latest performance level of the first audio source node.

[0079] In this embodiment, the scheduling server first selects nodes geographically close to the playback terminal from all preset audio source nodes based on the playback terminal's location information. Then, it further selects nodes whose performance levels meet the requirements as the final candidate audio source nodes. This embodiment introduces a geographically based initial screening mechanism and a performance-level-based secondary screening mechanism to ensure that the selected candidate audio source nodes are not only geographically close but also have excellent performance.

[0080] The location information of the playback terminal refers to the geographical location data of the playback terminal. It can be obtained through the Global Positioning System (GPS), Wi-Fi positioning, IP address lookup, or base station positioning. It should be noted that the accuracy of the location information of the playback terminal does not need to be too high. It can be accurate to the street or administrative district. The specifics are not limited here.

[0081] The preset location information of audio source nodes refers to the geographical location information of all available audio source servers or CDN nodes pre-configured by the audio playback system, usually the physical address of the server room or data center where they are deployed. This preset location information can be pre-stored in the scheduling server in the form of a database or data table, and the scheduling server can obtain it directly.

[0082] When performing location information matching, the first audio source node can be selected by calculating the straight-line distance or network routing distance between the playback terminal and the preset audio source node. Nodes with a distance less than a certain threshold or the N nodes closest to the playback terminal can be selected as the first audio source node. Alternatively, nodes on the path with fewer than a certain threshold of hops or lower than a certain threshold of network latency can be selected as the first audio source node based on the preset network topology map. Other methods can also be used to match the first audio source node, but these will not be elaborated here.

[0083] After determining the primary audio source node, the latest performance level of all primary audio source nodes stored on the scheduling server can be obtained. The primary audio source node with a performance level no lower than the current audio source node is determined as the alternative audio source node to ensure that the playback experience is at least not degraded after switching, and is usually improved.

[0084] In one implementation, the step of determining a target audio source node from candidate audio source nodes based on the attribute information and transmission performance status data of candidate audio source nodes in response to the triggering of a preset performance level degradation event includes: the playback terminal predicting the performance level change trend of the playback terminal in response to the numerical change of any transmission performance status index in the transmission performance status data, and obtaining a prediction result; if the prediction result indicates that the performance level of the playback terminal has degraded, then the preset performance level degradation event is triggered, and the target audio source node is determined from candidate audio source nodes based on the attribute information and transmission performance status data of candidate audio source nodes.

[0085] In this embodiment, when the playback terminal detects a change in the value of any transmission performance status indicator, it triggers a prediction of the performance level change trend of the playback terminal and obtains a prediction result. The prediction result is used to indicate whether the performance level of the playback terminal has degraded, stabilized, or improved. Degradation means that the performance level has deteriorated, and improvement means that the performance level has improved. Through trend prediction, switching can be performed before the user perceives obvious problems, thereby further improving the smoothness and continuity of the playback experience and turning passive response into active avoidance.

[0086] In one implementation, the playback terminal can trigger the prediction of the performance level change trend only when the degree of change of any transmission performance status indicator exceeds the corresponding degree of change threshold, so as to improve the prediction performance.

[0087] When predicting the performance level change trend of a playback terminal, the mean change rate, variance / standard deviation, linear regression, standard score (z-score), or interquartile range (IQR) of each transmission performance status index can be calculated based on the transmission performance status data of the most recent multiple periods to obtain the prediction result.

[0088] For example, the average rate of change of a certain indicator over a recent period can be calculated as the prediction result. Suppose that if the average RTT increases by more than 20ms / cycle and the average packet loss rate increases by more than 0.5% / cycle over three consecutive cycles (assuming 90 seconds), the prediction result obtained can indicate that the performance is declining significantly. The specifics are not limited here.

[0089] Based on the prediction results obtained through different methods, the corresponding judgment method is used to determine whether the performance level of the playback terminal has degraded. If degradation occurs, a preset performance level degradation event is triggered. Based on the attribute information and transmission performance status data of the candidate audio source nodes, the target audio source node is determined from the candidate audio source nodes. If no degradation occurs, there is no need to trigger the preset performance level degradation event. Alternatively, a preset performance level evolution event can be triggered to report the latest performance level of the current audio source node to the scheduling server or to execute other preset evolution events. The details are not elaborated here.

[0090] In one implementation, before the step of switching the playback terminal to obtain audio data from the target audio source node and switching the audio source based on the obtained audio data, the method further includes: if the playback terminal responds to a situation where the performance level of any candidate audio source node increases by a factor greater than a preset threshold and the increased performance level is better than the current audio source level, then the corresponding candidate audio source node is determined as the target audio source node.

[0091] In addition to switching audio source nodes when the current audio source node experiences significant performance degradation, the playback terminal can also proactively switch to a better-performing node when a candidate audio source node is found to have significantly improved performance. In this embodiment, the playback terminal not only considers switching when the current node's performance declines, but also continuously monitors the performance of candidate audio source nodes and proactively switches when a better option is found, thereby achieving continuous performance optimization.

[0092] The scheduling server can identify and send the judgment results regarding whether the performance level improvement of the candidate audio source node is greater than the preset threshold and whether the improved performance level is better than the current audio source level. The playback terminal only needs to respond to the judgment results sent by the scheduling server, determine whether to identify the corresponding candidate audio source node as the target audio source node, and execute the following steps: switch to obtain audio data from the target audio source node, and switch the audio source for playback based on the obtained audio data.

[0093] For the corresponding method embodiments described above, see [link to relevant documentation]. Figure 2The diagram illustrates an audio playback system, comprising a playback terminal 22 and a scheduling server 24. The playback terminal 22 is configured to: transmit the transmission performance status data of the current audio source node to the scheduling server 24 during playback of audio data transmitted from the current audio source node; the scheduling server 24 is configured to: calculate the transmission performance level of the transmission performance status data to obtain the current performance level of the current audio source node; the scheduling server 24 is configured to: determine at least one candidate audio source node based on the current performance level, and transmit the attribute information and transmission performance status data of the candidate audio source node to the playback terminal 22, wherein the candidate audio source node has a transmission performance level at least higher than the current performance level; the playback terminal 22 is configured to: in response to the triggering of a preset performance level degradation event, determine a target audio source node from the candidate audio source nodes based on the attribute information and transmission performance status data; and the playback terminal 22 is configured to: switch to acquiring audio data from the target audio source node, and switch audio source playback based on the acquired audio data.

[0094] The aforementioned audio playback system measures the transmission performance of each audio source node by collecting actual node transmission performance status data from the playback terminal. The scheduling server uniformly calculates and generates alternative audio source nodes for the playback terminal, realizing a performance feedback linkage mechanism with the upper-layer audio application layer protocol. When the playback performance of the playback terminal degrades, it can quickly and seamlessly switch over, thereby avoiding delays, buffering, and reduced sound quality, thus improving the continuity and stability of audio playback.

[0095] Optionally, when the playback terminal 22 transmits the transmission performance status data of the current audio source node to the scheduling server 24 during the playback of audio data transmitted by the current audio source node, the playback terminal 22 is configured to: periodically collect raw network metrics of the network layer based on the connection between the playback terminal 22 and the current audio source node during the playback of audio data transmitted by the current audio source node, wherein the raw network metrics include round-trip propagation delay and / or packet loss rate; periodically collect playback status metrics of the playback terminal 22 through the player of the playback terminal 22, wherein the playback status metrics include buffer event parameters of the player and / or first frame audio playback time parameters; combine the raw network metrics and the playback status metrics of the same period to obtain the transmission performance status data of the current audio source node within the same period; and periodically transmit the transmission performance status data of the current audio source node within the most recent period to the scheduling server 24.

[0096] Optionally, the transmission performance status data includes multiple transmission performance status indicators; the scheduling server 24 is used to: calculate the transmission performance level of the transmission performance status data to obtain the current performance level of the current audio source node, including: the scheduling server 24 is used to: determine the transmission performance level threshold corresponding to each transmission performance status indicator according to the network connection status attribute of the playback terminal 22, wherein the transmission performance level threshold includes the numerical range corresponding to multiple transmission performance levels respectively; compare each transmission performance status indicator in the transmission performance status data with the numerical range corresponding to each transmission performance level respectively to determine the first performance level to which each transmission performance status indicator of the current audio source node belongs; and determine the current performance level of the current audio source node according to the first performance level.

[0097] Optionally, when the scheduling server 24 determines the current performance level of the current audio source node based on the first performance level, the scheduling server 24 performs weighted summation of the first performance level to which each transmission performance status indicator of the current audio source node belongs, according to the preset weights corresponding to each transmission performance status indicator, to obtain the current performance level of the current audio source node.

[0098] Optionally, when the scheduling server 24 determines at least one candidate audio source node based on the current performance level, it includes: the scheduling server 24 obtaining the location information of the playback terminal 22 and the preset audio source nodes, and determining at least one first audio source node from the preset audio source nodes by matching the location information; and determining the first audio source node whose latest performance level is greater than or equal to the current performance level of the current audio source node as a candidate audio source node based on the latest performance level of the first audio source node.

[0099] Optionally, when the playback terminal 22 responds to the triggering of a preset performance level degradation event and determines the target audio source node from the candidate audio source nodes based on the attribute information and transmission performance status data of the candidate audio source nodes, the following steps are included: the playback terminal 22 responds to the numerical change of any transmission performance status index in the transmission performance status data, predicts the performance level change trend of the playback terminal 22, and obtains a prediction result; if the prediction result indicates that the performance level of the playback terminal 22 has degraded, then the preset performance level degradation event is triggered, and the target audio source node is determined from the candidate audio source nodes based on the attribute information and transmission performance status data of the candidate audio source nodes.

[0100] Optionally, before the playback terminal 22 switches to obtain audio data from the target audio source node and performs audio source switching playback based on the obtained audio data, the playback terminal 22 is further configured to: determine the corresponding candidate audio source node as the target audio source node in response to the fact that the performance level improvement of any candidate audio source node is greater than a preset threshold and the improved performance level is better than the current audio source level.

[0101] This embodiment also provides an electronic device, including a processor and a memory. The memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the above-described audio playback method. This electronic device can be a server or a terminal device.

[0102] See Figure 3 As shown, the electronic device includes a processor 100 and a memory 101. The memory 101 stores machine-executable instructions that can be executed by the processor 100. The processor 100 executes the machine-executable instructions to implement the above-described audio playback method.

[0103] Furthermore, Figure 3 The electronic device shown also includes a bus 102 and a communication interface 103, with the processor 100, the communication interface 103 and the memory 101 connected via the bus 102.

[0104] The memory 101 may include high-speed random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Communication between this system network element and at least one other network element is achieved through at least one communication interface 103 (which can be wired or wireless), such as the Internet, wide area network, local area network, or metropolitan area network. The bus 102 may be an ISA bus, PCI bus, or EISA bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 3 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.

[0105] The processor 100 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of the processor 100 or by instructions in software form. The processor 100 may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this disclosure. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this disclosure can be directly manifested as execution by a hardware decoding processor, or execution by a combination of hardware and software modules in the decoding processor. The software module can reside in a readily available storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory 101. The processor 100 reads information from memory 101 and, in conjunction with its hardware, completes the steps of the method described in the foregoing embodiments, for example:

[0106] This method, applied to an audio playback system including a playback terminal and a scheduling server, includes the following steps: During playback of audio data transmitted from the current audio source node, the playback terminal transmits the transmission performance status data of the current audio source node to the scheduling server; the scheduling server calculates the transmission performance level of the transmission performance status data to obtain the current performance level of the current audio source node; based on the current performance level, the scheduling server determines at least one candidate audio source node and transmits the attribute information and transmission performance status data of the candidate audio source node to the playback terminal; in response to the triggering of a preset performance level degradation event, the playback terminal determines a target audio source node from the candidate audio source nodes based on the attribute information and transmission performance status data of the candidate audio source nodes; the playback terminal switches to acquire audio data from the target audio source node and performs audio source switching playback based on the acquired audio data.

[0107] In this method, the transmission performance of each audio source node is measured by the actual node transmission performance status data collected by the playback terminal. The scheduling server uniformly calculates and generates alternative audio source nodes for the playback terminal, realizing a performance feedback linkage mechanism with the upper-layer audio application layer protocol. When the playback performance of the playback terminal degrades, a fast and seamless switch can be performed, thereby avoiding delays, buffering, and reduced sound quality, thus improving the continuity and stability of audio playback.

[0108] Optionally, the step of transmitting the transmission performance status data of the current audio source node to the scheduling server during the playback of audio data transmitted by the current audio source node includes: During the playback of audio data transmitted by the current audio source node, the playback terminal periodically collects raw network metrics of the network layer based on the connection between the playback terminal and the current audio source node, wherein the raw network metrics include round-trip propagation delay and / or packet loss rate; periodically collects playback status metrics of the playback terminal through the player of the playback terminal, wherein the playback status metrics include buffer event parameters of the player and / or first frame audio playback time parameters; combines the raw network metrics and playback status metrics of the same period to obtain the transmission performance status data of the current audio source node within the same period; and periodically transmits the transmission performance status data of the current audio source node within the most recent period to the scheduling server.

[0109] Optionally, the transmission performance status data includes multiple transmission performance status indicators; the step of the scheduling server calculating the transmission performance level of the transmission performance status data to obtain the current performance level of the current audio source node includes: the scheduling server determining the transmission performance level threshold corresponding to each transmission performance status indicator based on the network connection status attribute of the playback terminal, wherein the transmission performance level threshold includes the numerical range corresponding to multiple transmission performance levels; comparing each transmission performance status indicator in the transmission performance status data with the numerical range corresponding to each transmission performance level to determine the first performance level to which each transmission performance status indicator of the current audio source node belongs; and determining the current performance level of the current audio source node based on the first performance level.

[0110] Optionally, the step of determining the current performance level of the current audio source node based on the first performance level includes: the scheduling server weighting and superimposing the first performance level to which each transmission performance status indicator of the current audio source node belongs according to the preset weight corresponding to each transmission performance status indicator, to obtain the current performance level of the current audio source node.

[0111] Optionally, the step of the scheduling server determining at least one candidate audio source node based on the current performance level includes: the scheduling server obtaining the location information of the playback terminal and the preset audio source nodes, determining at least one first audio source node from the preset audio source nodes by matching the location information; and determining the first audio source node whose latest performance level is greater than or equal to the current performance level of the current audio source node as a candidate audio source node based on the latest performance level of the first audio source node.

[0112] Optionally, the step of determining the target audio source node from the candidate audio source nodes based on the attribute information and transmission performance status data of the candidate audio source nodes in response to the triggering of a preset performance level degradation event includes: the playback terminal predicting the performance level change trend of the playback terminal in response to the numerical change of any transmission performance status index in the transmission performance status data, and obtaining a prediction result; if the prediction result indicates that the performance level of the playback terminal has degraded, then the preset performance level degradation event is triggered, and the target audio source node is determined from the candidate audio source nodes based on the attribute information and transmission performance status data of the candidate audio source nodes.

[0113] Optionally, before the step of switching the playback terminal to obtain audio data from the target audio source node and switching the audio source based on the obtained audio data, the method further includes: if the playback terminal responds to the fact that the performance level of any candidate audio source node increases by more than a preset threshold and the improved performance level is better than the current audio source level, then the corresponding candidate audio source node is determined as the target audio source node.

[0114] This embodiment also provides a computer-readable storage medium storing computer-executable instructions. When these computer-executable instructions are invoked and executed by a processor, they cause the processor to implement the aforementioned audio playback method, for example:

[0115] This method, applied to an audio playback system including a playback terminal and a scheduling server, includes the following steps: During playback of audio data transmitted from the current audio source node, the playback terminal transmits the transmission performance status data of the current audio source node to the scheduling server; the scheduling server calculates the transmission performance level of the transmission performance status data to obtain the current performance level of the current audio source node; based on the current performance level, the scheduling server determines at least one candidate audio source node and transmits the attribute information and transmission performance status data of the candidate audio source node to the playback terminal; in response to the triggering of a preset performance level degradation event, the playback terminal determines a target audio source node from the candidate audio source nodes based on the attribute information and transmission performance status data of the candidate audio source nodes; the playback terminal switches to acquire audio data from the target audio source node and performs audio source switching playback based on the acquired audio data.

[0116] In this method, the transmission performance of each audio source node is measured by the actual node transmission performance status data collected by the playback terminal. The scheduling server uniformly calculates and generates alternative audio source nodes for the playback terminal, realizing a performance feedback linkage mechanism with the upper-layer audio application layer protocol. When the playback performance of the playback terminal degrades, a fast and seamless switch can be performed, thereby avoiding delays, buffering, and reduced sound quality, thus improving the continuity and stability of audio playback.

[0117] Optionally, the step of transmitting the transmission performance status data of the current audio source node to the scheduling server during the playback of audio data transmitted by the current audio source node includes: During the playback of audio data transmitted by the current audio source node, the playback terminal periodically collects raw network metrics of the network layer based on the connection between the playback terminal and the current audio source node, wherein the raw network metrics include round-trip propagation delay and / or packet loss rate; periodically collects playback status metrics of the playback terminal through the player of the playback terminal, wherein the playback status metrics include buffer event parameters of the player and / or first frame audio playback time parameters; combines the raw network metrics and playback status metrics of the same period to obtain the transmission performance status data of the current audio source node within the same period; and periodically transmits the transmission performance status data of the current audio source node within the most recent period to the scheduling server.

[0118] Optionally, the transmission performance status data includes multiple transmission performance status indicators; the step of the scheduling server calculating the transmission performance level of the transmission performance status data to obtain the current performance level of the current audio source node includes: the scheduling server determining the transmission performance level threshold corresponding to each transmission performance status indicator based on the network connection status attribute of the playback terminal, wherein the transmission performance level threshold includes the numerical range corresponding to multiple transmission performance levels; comparing each transmission performance status indicator in the transmission performance status data with the numerical range corresponding to each transmission performance level to determine the first performance level to which each transmission performance status indicator of the current audio source node belongs; and determining the current performance level of the current audio source node based on the first performance level.

[0119] Optionally, the step of determining the current performance level of the current audio source node based on the first performance level includes: the scheduling server weighting and superimposing the first performance level to which each transmission performance status indicator of the current audio source node belongs according to the preset weight corresponding to each transmission performance status indicator, to obtain the current performance level of the current audio source node.

[0120] Optionally, the step of the scheduling server determining at least one candidate audio source node based on the current performance level includes: the scheduling server obtaining the location information of the playback terminal and the preset audio source nodes, determining at least one first audio source node from the preset audio source nodes by matching the location information; and determining the first audio source node whose latest performance level is greater than or equal to the current performance level of the current audio source node as a candidate audio source node based on the latest performance level of the first audio source node.

[0121] Optionally, the step of determining the target audio source node from the candidate audio source nodes based on the attribute information and transmission performance status data of the candidate audio source nodes in response to the triggering of a preset performance level degradation event includes: the playback terminal predicting the performance level change trend of the playback terminal in response to the numerical change of any transmission performance status index in the transmission performance status data, and obtaining a prediction result; if the prediction result indicates that the performance level of the playback terminal has degraded, then the preset performance level degradation event is triggered, and the target audio source node is determined from the candidate audio source nodes based on the attribute information and transmission performance status data of the candidate audio source nodes.

[0122] Optionally, before the step of switching the playback terminal to obtain audio data from the target audio source node and switching the audio source based on the obtained audio data, the method further includes: if the playback terminal responds to the fact that the performance level of any candidate audio source node increases by more than a preset threshold and the improved performance level is better than the current audio source level, then the corresponding candidate audio source node is determined as the target audio source node.

[0123] The computer program products for audio playback methods, systems, electronic devices, and storage media provided in this disclosure include computer-readable storage media storing program code. The instructions included in the program code can be used to execute the methods described in the preceding method embodiments. For specific implementation details, please refer to the method embodiments, which will not be repeated here.

[0124] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the system and apparatus described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0125] Furthermore, in the description of the embodiments of this disclosure, unless otherwise expressly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this disclosure based on the specific circumstances.

[0126] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0127] In the description of this disclosure, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing this disclosure and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this disclosure. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0128] Finally, it should be noted that the above embodiments are merely specific implementations of this disclosure, used to illustrate the technical solutions of this disclosure, and not to limit it. The protection scope of this disclosure is not limited thereto. Although this disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this disclosure. Such modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this disclosure, and should all be covered within the protection scope of this disclosure. Therefore, the protection scope of this disclosure should be determined by the protection scope of the claims.

Claims

1. An audio playback method, applied to an audio playback system, characterized in that, The audio playback system includes a playback terminal and a scheduling server, and the method includes: During the playback of audio data transmitted by the current audio source node, the playback terminal transmits the transmission performance status data of the current audio source node to the scheduling server. The scheduling server calculates the transmission performance level of the transmission performance status data to obtain the current performance level of the current sound source node; The scheduling server determines at least one candidate audio source node based on the current performance level, and transmits the attribute information and transmission performance status data of the candidate audio source node to the playback terminal. In response to the triggering of a preset performance level degradation event, the playback terminal determines the target audio source node from the candidate audio source nodes based on the attribute information and transmission performance status data of the candidate audio source nodes. The playback terminal switches to acquiring audio data from the target audio source node, and switches the audio source for playback based on the acquired audio data.

2. The method according to claim 1, characterized in that, The step of transmitting the transmission performance status data of the current audio source node to the scheduling server during the playback of audio data transmitted by the current audio source node includes: During the process of playing audio data transmitted by the current audio source node, the playback terminal periodically collects raw network metrics of the network layer based on the connection between the playback terminal and the current audio source node. The raw network metrics include round-trip propagation delay and / or packet loss rate. The playback status indicators of the playback terminal are periodically collected through the player of the playback terminal, wherein the playback status indicators include the buffer event parameters of the player and / or the first frame audio playback time parameters. By combining the original network metrics and the playback status metrics of the same period, the transmission performance status data of the current audio source node within the same period can be obtained. The transmission performance status data of the current audio source node in the most recent period is periodically transmitted to the scheduling server.

3. The method according to claim 1, characterized in that, The transmission performance status data includes multiple transmission performance status indicators; The step of the scheduling server calculating the transmission performance level of the transmission performance status data to obtain the current performance level of the current audio source node includes: The scheduling server determines the transmission performance level threshold corresponding to each transmission performance status indicator based on the network connection status attribute of the playback terminal. The transmission performance level threshold includes a numerical range corresponding to multiple transmission performance levels. The transmission performance status indicators in the transmission performance status data are compared with the numerical ranges corresponding to each transmission performance level to determine the first performance level to which each transmission performance status indicator of the current audio source node belongs. Based on the first performance level, determine the current performance level of the current sound source node.

4. The method according to claim 3, characterized in that, The step of determining the current performance level of the current sound source node based on the first performance level includes: The scheduling server performs a weighted summation of the first performance level of each transmission performance status indicator of the current audio source node according to the preset weights corresponding to each transmission performance status indicator, so as to obtain the current performance level of the current audio source node.

5. The method according to claim 1, characterized in that, The step of the scheduling server determining at least one candidate audio source node based on the current performance level includes: The scheduling server obtains the location information of the playback terminal and the preset sound source node, and determines at least one first sound source node from the preset sound source nodes by matching the location information. Based on the latest performance level of the first audio source node, the first audio source node whose latest performance level is greater than or equal to the current performance level of the current audio source node is determined as a candidate audio source node.

6. The method according to claim 1, characterized in that, The step of the playback terminal determining the target audio source node from the candidate audio source nodes based on the attribute information and transmission performance status data of the candidate audio source nodes in response to the triggering of a preset performance level degradation event includes: The playback terminal responds to changes in the values ​​of any transmission performance status index in the transmission performance status data, predicts the trend of performance level changes of the playback terminal, and obtains the prediction result. If the prediction result indicates that the performance level of the playback terminal has degraded, a preset performance level degradation event is triggered, and the target audio source node is determined from the candidate audio source nodes based on the attribute information and transmission performance status data of the candidate audio source nodes.

7. The method according to claim 1, characterized in that, Before the step of switching the playback terminal to obtain audio data from the target audio source node and switching playback based on the obtained audio data, the method further includes: If the playback terminal responds to a situation where the performance level of any candidate audio source node increases by a factor greater than a preset threshold and the improved performance level is better than the current audio source level, then the corresponding candidate audio source node is determined as the target audio source node.

8. An audio playback system, characterized in that, The system includes a playback terminal and a scheduling server, wherein... The playback terminal is used to: transmit the transmission performance status data of the current audio source node to the scheduling server during the playback of audio data transmitted by the current audio source node; The scheduling server is used to: calculate the transmission performance level of the transmission performance status data to obtain the current performance level of the current sound source node; The scheduling server is used to: determine at least one candidate audio source node according to the current performance level, and transmit the attribute information and transmission performance status data of the candidate audio source node to the playback terminal, wherein the candidate audio source node has a transmission performance level that is at least higher than the current performance level; The playback terminal is used to: in response to the triggering of a preset performance level degradation event, determine the target audio source node from the candidate audio source nodes based on the attribute information and transmission performance status data of the candidate audio source nodes; The playback terminal is used to: switch to obtain audio data from the target audio source node, and switch audio source playback based on the obtained audio data.

9. An electronic device, characterized in that, The device includes a processor and a memory, the memory storing machine-executable instructions that can be executed by the processor, the processor executing the machine-executable instructions to implement the audio playback method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when invoked and executed by a processor, cause the processor to implement the audio playback method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Multimedia content monitoring system, method and device based on content distributing network

    CN101420458A

  • Dynamic following broadcast device and method

    CN101820503A

  • Method of CDN to actively select high quality nodes in advance to conduct optimizing content distribution service

    CN102984279A

  • Video stream acquisition method and device, equipment and storage medium

    CN119562093A

  • Video stream code rate adjustment method and apparatus, computer device, and storage medium

    WO2024051426A1