System, method and device for balancing network transmission bandwidth and medium
By constructing a closed-loop adjustment mechanism and utilizing state detection and feedback mechanisms to optimize audio transmission parameters, the stability problem of AoIP technology under network fluctuations is solved, thereby improving the robustness of audio transmission and user experience.
Patent Information
- Application Number
- CN202511475469.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-15
- Publication Date
- 2025-11-18
AI Technical Summary
Existing AoIP technology suffers from insufficient audio transmission stability during network fluctuations, leading to audio dropouts. It also lacks comprehensive and coordinated optimization of network conditions and receiver buffering, impacting user experience.
A closed-loop adjustment mechanism integrating performance adjustment strategies and triple state feedback at the receiver is constructed. The state detection module monitors network transmission indicators and clock synchronization information to generate a network quality score. Based on the feedback data, audio transmission parameters are adjusted to achieve adaptive optimization.
It significantly improves the robustness of audio transmission and the user's listening experience, ensures the continuity and sound quality of audio playback, and enhances system reliability in weak network environments.
Smart Images

Figure CN120980041A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of audio transmission, and in particular to a system, method, apparatus, and medium for balancing network transmission bandwidth. Background Technology
[0002] In the field of audio transmission, Audio over Internet Protocol (AoIP) technology has been widely used. It replaces analog cables with standard network equipment to realize the digital and networked transmission of audio signals, significantly simplifying system cabling and providing routing flexibility.
[0003] However, existing AoIP technology faces key challenges in practical applications: dynamic network fluctuations (bandwidth, latency, jitter, packet loss) lead to insufficient stability in audio transmission, especially when the network deteriorates, bandwidth resource contention can cause audio data transmission interruptions (i.e., audio dropouts), severely impairing the user experience. Current network performance tuning strategies are relatively simple and rigid, lacking comprehensive and coordinated optimization of network conditions and receiver buffers, resulting in insufficient adaptability.
[0004] Therefore, there is an urgent need for a system, method, device, and medium that can balance network transmission bandwidth, intelligently balance bandwidth usage, sound quality, and latency while ensuring the continuity of audio playback, thereby improving the integrity, smoothness, and user listening experience of audio transmission. Summary of the Invention
[0005] This specification provides one or more embodiments of a system for balancing network transmission bandwidth. The system includes: a status detection module configured to: monitor network transmission metrics and clock synchronization information; generate a network quality score based on the network transmission metrics and clock synchronization information; a metric detection module consisting of a sender and a receiver, wherein: the sender is configured to send audio data; the receiver is configured to generate feedback data based on the received audio data and send the feedback data back to the sender, the feedback data including buffer change trends, the number of packet losses due to buffer overflow, and clock offset; and a performance adjustment module configured to: determine a performance adjustment strategy based on the network quality score and the feedback data; and adjust audio transmission parameters based on the performance adjustment strategy.
[0006] This specification provides one or more embodiments of a method for balancing network transmission bandwidth. The method includes: monitoring network transmission metrics and clock synchronization information; generating a network quality score based on the network transmission metrics and clock synchronization information; generating feedback data based on received audio data, the feedback data including buffer change trends, the number of packet losses caused by buffer overflow, and clock offset; determining a performance adjustment strategy based on the network quality score and feedback data; and adjusting audio transmission parameters based on the performance adjustment strategy.
[0007] One or more embodiments of this specification also provide an apparatus for balancing network transmission bandwidth, the apparatus including at least one processor and at least one memory; the at least one memory is used to store computer instructions; the at least one processor is used to execute at least a portion of the computer instructions to implement a method for balancing network transmission bandwidth.
[0008] One or more embodiments of this specification also provide a computer-readable storage medium having executable instructions stored thereon that, when executed by a processor, enable the aforementioned method for balancing network transmission bandwidth. Attached Figure Description
[0009] This specification will be further described by way of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. These embodiments are not limiting; in these embodiments, the same reference numerals denote the same structures, wherein: Figure 1 This is a schematic diagram illustrating an application scenario of a method for balancing network transmission bandwidth according to some embodiments of this specification; Figure 2 This is an exemplary block diagram of a network transmission bandwidth balancing system according to some embodiments of this specification; Figure 3 This is an exemplary flowchart illustrating a method for balancing network transmission bandwidth according to some embodiments of this specification; Figure 4 This is a schematic diagram illustrating the generation of feedback data according to some embodiments of this specification; Figure 5 This is a schematic diagram illustrating a performance tuning strategy according to some embodiments of this specification; Figure 6 This is a schematic diagram illustrating another performance tuning strategy according to other embodiments of this specification. Detailed Implementation
[0010] The accompanying drawings used in the description of the embodiments will be briefly introduced below. The drawings do not represent all embodiments.
[0011] The terms “system,” “device,” “unit,” and / or “module” as used herein are one method of distinguishing different components, elements, parts, sections, or assemblies at different levels. However, if other terms can achieve the same purpose, they may be replaced by other expressions.
[0012] Unless the context clearly indicates an exception, words such as "a," "an," "a kind," and / or "the" do not specifically refer to the singular and may also include the plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of explicitly identified steps and elements, which do not constitute an exclusive list, and the method or apparatus may also include other steps or elements.
[0013] In demanding AoIP scenarios such as broadcasting and live sound reinforcement, existing technologies often cause audio interruptions or degraded sound quality and generate noise because they cannot dynamically adapt to network fluctuations (bandwidth, latency, packet loss). Furthermore, they neglect clock synchronization stability, which impairs system reliability.
[0014] This application provides a method, system, apparatus, and medium for balancing network transmission bandwidth. By constructing a closed-loop adjustment mechanism that integrates performance adjustment strategies and triple state feedback at the receiving end, it achieves progressive and adaptive optimization of audio transmission strategies, significantly improving the robustness of audio transmission and user listening experience in weak network environments.
[0015] Figure 1 This is a schematic diagram illustrating an application scenario of a method for balancing network transmission bandwidth according to some embodiments of this specification. In some embodiments, such as Figure 1 As shown, the application scenario 100 of the method for balancing network transmission bandwidth may include a processor 110, an audio receiving device 120, an audio transmitting device 130, a network 140, and a storage device 150.
[0016] The application scenarios of the method for balancing network transmission bandwidth can include, but are not limited to, various scenarios such as conference rooms, cinemas, shopping malls, live broadcast rooms, and concerts.
[0017] In some embodiments, the processor 110 may process data and / or information obtained from components or other external devices in the application scenario 100 of the method for balancing network transmission bandwidth. The processor may execute program instructions based on this data, information, and / or processing results to perform one or more functions described in this application. For example, monitoring network transmission metrics and clock synchronization information, generating performance tuning strategies, receiving audio data to generate feedback data, and determining performance tuning strategies, etc.
[0018] In some embodiments, processor 110 may be a local processor or an external processor of a system that balances network transmission bandwidth. In some embodiments, processor 110 may be a computer, a user console, a single processor, or a processor group, etc. The processor group may be centralized or distributed. In some embodiments, processor 110 may be implemented on a cloud platform. For example, the cloud platform may include one or any combination of private cloud, public cloud, hybrid cloud, etc.
[0019] Audio receiving device 120 refers to a device responsible for receiving, decoding, and playing audio. In some embodiments, audio receiving device 120 may integrate a microphone array, a WIFI module, a Bluetooth module, and a network module, etc. For example, audio receiving device 120 may be a microphone 120-1, a recording device 120-2, or any combination thereof.
[0020] In some embodiments, the audio receiving device 120 may be used to receive audio data from the audio transmitting device 130, generate feedback data based on the audio data, and send the feedback data back to the audio transmitting device 130.
[0021] Audio transmitting device 130 refers to a device used to transmit audio data. In some embodiments, audio transmitting device 130 may integrate audio output devices (such as speaker 130-1, power amplifier 130-2, speaker 130-3), WIFI module, Bluetooth module, and network module, etc.
[0022] In some embodiments, the audio transmitting device 130 can be used to transmit audio data.
[0023] In some embodiments, the audio receiving device 120 can serve as the receiving end of the indicator detection module, and the audio transmitting device 130 can serve as the transmitting end of the indicator detection module. For a description of the indicator detection module, the receiving end, and the transmitting end, please refer to [link to relevant documentation]. Figure 2 Related descriptions.
[0024] Network 140 includes any suitable network capable of facilitating information and / or data exchange in application scenario 100 of the method for balancing network transmission bandwidth. In some embodiments, one or more components of application scenario 100 of the method for balancing network transmission bandwidth (e.g., processor 110, audio receiving device 120, audio transmitting device 130, and storage device 150, etc.) can exchange information and / or data through network 140.
[0025] Network 140 can be any one or more of wired or wireless networks. For example, network 140 may include Bluetooth, Wi-Fi, cable network, fiber optic network, telecommunications network, cable connection, etc., or any combination thereof.
[0026] In some embodiments, storage device 150 may be used to store data and / or instructions. For example, storage device 150 may store network transmission metrics, clock synchronization information, audio data, and performance tuning strategies.
[0027] Storage device 150 may include one or more storage components, each of which may be a separate device or part of another device. In some embodiments, storage device 150 may include random access memory (RAM), read-only memory (ROM), mass storage (e.g., disk, optical disk, solid-state drive, etc.), or any combination thereof.
[0028] In some embodiments, storage device 150 may also be implemented on a cloud platform. By way of example only, a cloud platform may include a private cloud, a public cloud, a hybrid cloud, or any combination thereof.
[0029] In some embodiments, the storage device 150 can communicate with one or more components in the application scenario 100 of the balanced network transmission bandwidth method via the network 140.
[0030] For more information on the network transmission metrics, clock synchronization information, performance tuning strategies, audio data, feedback data, and other parameters mentioned above, please refer to [link to relevant documentation]. Figure 3 The relevant description in the document.
[0031] Figure 2 This is an exemplary block diagram of a system for balancing network transmission bandwidth, as shown in some embodiments of this specification. Figure 2 As shown, the system 200 for balancing network transmission bandwidth includes a status detection module 210, an indicator detection module 220, and a performance adjustment module 230.
[0032] In some embodiments, the status detection module 210, the index detection module 220, and the performance adjustment module 230 in the system 200 for balancing network transmission bandwidth can be integrated into the processor.
[0033] The state detection module is used to monitor and evaluate network quality scores in real time. In some embodiments, the state detection module 210 is configured to monitor network transmission metrics and clock synchronization information; and generate a network quality score based on the network transmission metrics and clock synchronization information.
[0034] The indicator detection module is used to detect the transmission channel conditions for transmitting audio data. In some embodiments, the indicator detection module 220 consists of a transmitter 221 and a receiver 222.
[0035] The transmitting end, also known as the Transmit (TX) end, is one of the core devices deployed in the system 200 that balances network transmission bandwidth. It can transmit audio data to various receiving ends 222 via the Real-time Transport Protocol (RTP). In some embodiments, the transmitting end 221 is the signal source node of the system 200 that balances network transmission bandwidth during audio transmission, and performs digital encapsulation and multiplexing of the audio signal via RTP.
[0036] In some embodiments, the transmitter 221 is configured to transmit audio data.
[0037] The receiving end, also known as the Receive (RX) end, is responsible for receiving and analyzing audio data.
[0038] In some embodiments, the receiver 222 is configured to generate feedback data based on the received audio data and send the feedback data back to the sender 221; the feedback data includes buffer change trends, the number of packet losses caused by buffer overflow, and clock offset.
[0039] In some embodiments, the receiver 222 is further configured to determine feedback data based on the received audio data through a triple feedback mechanism, wherein the triple feedback mechanism includes: analyzing the trend of buffer changes; counting the number of packet losses caused by buffer overflow; and monitoring the clock offset.
[0040] In some embodiments, in response to the network quality score meeting a preset condition, the sending end 221 preloads audio data to the receiving end 222, and the preloading duration is determined based on the buffer consumption rate of the receiving end 222.
[0041] In some embodiments, the indicator detection module is further configured to monitor the amount of data in the buffer; trigger a progressive sampling rate conversion to smoothly reduce the sampling rate in response to the amount of data in the buffer falling below a safety threshold of the dynamic tolerance range; and limit the preload amount in response to the amount of data in the buffer falling above a saturation threshold of the dynamic tolerance range.
[0042] In some embodiments, the indicator detection module is further configured to prioritize audio data packets based on the content characteristics of the audio data, dividing them into high-priority packets and low-priority packets; and to prioritize discarding low-priority packets in response to the necessity of packet loss.
[0043] In some embodiments, the indicator detection module is further configured to split the audio data into multiple data units of fixed duration during the audio data transmission phase; and to perform high-precision interpolation compensation on the lost data units based on the data units that were not lost before and after them in response to the loss of data units.
[0044] The performance tuning module is used to adjust audio transmission parameters. In some embodiments, the performance tuning module is configured to determine a performance tuning strategy based on network quality scores and feedback data; and adjust the audio transmission parameters based on the performance tuning strategy.
[0045] In some embodiments, the performance tuning module is further configured to: reduce the number of transmission channels and decrease the sampling rate in response to a network quality score lower than a score threshold; and increase the number of transmission channels and increase the sampling rate in response to a network quality score higher than a score threshold; wherein the amount of change in the number of transmission channels and the amount of change in the sampling rate are determined based on feedback data.
[0046] In some embodiments, the performance adjustment module further includes a hysteresis response mechanism, which includes maintaining the audio transmission parameters unchanged in response to the fact that the difference between the network quality score and the score threshold is continuously less than a preset difference threshold within a preset time period.
[0047] In some embodiments, the performance tuning module is further configured to dynamically adjust the timing of the performance tuning strategy based on the user's preset listening priority or the user's historical behavior data, wherein the historical behavior data includes historical records of the analysis buffer state and historical records of manual intervention.
[0048] In some embodiments, the performance tuning module is further configured to: predict the bandwidth fluctuation trend in the future period based on network bandwidth data using a time series prediction model; and, in response to a positive bandwidth fluctuation trend in the future period, initiate a degradation strategy in the performance tuning strategy in advance, wherein the degradation strategy includes reducing the number of transmission channels and / or reducing the sampling rate.
[0049] In some embodiments, the performance tuning module is further configured to perform multiple downgrade adjustments to packet parameters in response to a network quality score below a score threshold, the downgrade adjustments including reducing packet size and shortening transmission interval; and to perform multiple upgrade adjustments to packet parameters in response to a network quality score above a score threshold, the upgrade adjustments including increasing packet size and extending transmission interval; wherein the packet size and transmission interval are determined based on feedback data, and the packet size adjustment adopts a gradual curve-like change.
[0050] For further explanation of the above content, please refer to [link / reference]. Figures 3-6 and its corresponding embodiments.
[0051] In some embodiments, Figure 2 The status detection module, indicator detection module, and performance adjustment module disclosed herein can be different modules within a single system, or a single module can implement the functions of two or more of the aforementioned modules. For example, the modules can share a single storage module, or each module can have its own separate storage module. Such variations are all within the scope of protection of this specification.
[0052] Figure 3 This is an exemplary flowchart of an exemplary process 300 of a method for balancing network transmission bandwidth according to some embodiments of this specification.
[0053] In some embodiments, process 300 is executed by a processor. For example... Figure 3 As shown, process 300 includes the following steps.
[0054] Step 310: Monitor network transmission metrics and clock synchronization information.
[0055] Network transmission metrics are indicators that reflect network transmission performance. Audio data is transmitted over a network, and the quality of audio transmission is related to network transmission metrics.
[0056] Audio data refers to digital audio signals, which are transmitted by the sending end over the network in the form of real-time streaming media data packets. Each audio data includes multiple audio data packets, which are audio data segments, including corresponding sequence numbers and timestamps. The receiving end arranges the multiple audio data packets received according to their sequence numbers, decodes the digital signals to obtain sound signals, and plays the sound signals precisely according to the timestamps to complete the network transmission of audio data.
[0057] In some embodiments, network transmission metrics include latency, packet loss rate, and bandwidth utilization.
[0058] Latency refers to the transmission time of audio data from the sending end to the receiving end. It reflects the immediacy of data transmission. The higher the latency, the worse the real-time performance of the sound and the lower the audio transmission quality. High latency may cause audio and video to be out of sync or interactive dialogue to stutter.
[0059] Packet loss rate refers to the percentage of data packets lost during audio data transmission out of the total transmitted data packets. A higher packet loss rate results in poorer sound integrity and lower audio transmission quality. Increased packet loss rate disrupts audio continuity, causing intermittent or distorted speech.
[0060] Bandwidth utilization refers to the percentage of network bandwidth currently being used relative to the network's maximum bandwidth. It is used to determine whether the network load is close to saturation. A higher bandwidth utilization indicates greater network congestion, triggering issues such as spiked latency and increased packet loss, and resulting in poorer audio transmission quality.
[0061] In some embodiments, network transmission metrics can be represented in sequence or vector form. For example, network transmission metrics with a latency of 10ms, a packet loss rate of 10%, and a bandwidth utilization rate of 50% can be expressed as a vector (10ms, 10%, 50%).
[0062] In some embodiments, network transmission metrics can be obtained through monitoring by the state detection module 210. For more information about the state detection module 210, please refer to [link to relevant documentation]. Figure 2 And its related descriptions.
[0063] Clock synchronization information refers to a parameter that measures the time coordination performance among multiple devices in a network.
[0064] In some embodiments, clock synchronization information includes clock offset, clock jitter, clock drift, etc.
[0065] Clock skew refers to the absolute time difference between the master clock and the slave clock at the same moment. For example, 5 μs. The master clock is the time reference source, such as the Precision Time Protocol Grandmaster (PTPGrandmaster); the slave clock is the device that needs to synchronize with the master clock (such as a receiver, transmitter, etc.). In some embodiments, the clock skew can be obtained by the processor based on a clock skew monitoring protocol (such as a timestamp exchange protocol).
[0066] Clock jitter refers to the fluctuation in the clock offset between the master and slave clocks during audio transmission, reflecting the short-term stability of master-slave clock synchronization. For example, clock jitter can be the variance or standard deviation of multiple clock offsets at multiple sampling time points within a first preset time period. The first preset time period and multiple sampling time points can be manually preset, for example, a sampling time point every 10ms within 100ms.
[0067] Clock drift refers to the relative frequency deviation between the master and slave clocks, reflecting the long-term stability of clock synchronization. It is calculated from the clock offset within a second preset time period. For example, if the clock offset is 200μs within 2 seconds, then the clock drift is 100ppm. The second preset time period is longer than the first preset time period, and the second preset time period can be manually preset.
[0068] In some embodiments, clock synchronization information can be represented in the form of a sequence or a vector. For example, clock synchronization information with a clock offset of 5 μs, a clock jitter of 15 μs, and a clock drift of 100 ppm can be represented by a vector (5 μs, 15 μs, 100 ppm).
[0069] In some embodiments, clock synchronization information can be obtained through monitoring by the status detection module 210.
[0070] Step 320: Generate a network quality score based on network transmission metrics and clock synchronization information.
[0071] Network quality score is a comprehensive evaluation parameter used to reflect the current state of the network. In some embodiments, the network quality score can be represented by a value from 0 to 10, with higher values indicating a better network quality score and a better current network state.
[0072] In some embodiments, the processor can determine the network quality score based on network transmission metrics and clock synchronization information through various methods. For example, the processor can query a first preset table based on the network transmission metrics and clock synchronization information to determine the network quality score. The first preset table includes the correspondence between network transmission metrics, clock synchronization information, and network quality scores. The first preset table can be manually preset based on historical data. The processor can also determine the network quality score through other feasible methods (such as weighted fusion calculation methods, etc.), which are not limited here.
[0073] Step 330: Generate feedback data based on the received audio data.
[0074] Feedback data refers to the status information of the receiving end when receiving audio data. In some embodiments, feedback data includes buffer change trends, the number of packet losses caused by buffer overflow, and clock offset.
[0075] A buffer is a temporary storage area at the receiving end used to temporarily store audio data in order to ensure playback continuity.
[0076] The buffer change trend refers to the change in the amount of audio data stored in the buffer, including whether the amount of data in the buffer increases or decreases, as well as the specific buffer growth rate or buffer consumption rate.
[0077] Buffer consumption rate refers to the decrease in buffer data volume per unit time due to playback reading, reflecting the audio output rate. Buffer growth rate refers to the increase in buffer data volume per unit time due to network transmission writing, reflecting the network input rate. The buffer change trend reflects the balance between network transmission and playback consumption. A continuously decreasing buffer change trend indicates that the network input rate is less than the audio output rate, meaning the network input rate is insufficient, resulting in insufficient buffer data and interruption of audio playback. A continuously increasing buffer change trend indicates that the network input rate is greater than the audio output rate, meaning the network input rate is too fast, leading to buffer overflow and deliberate packet loss. Buffer overflow refers to the buffer's storage space being filled, while packet loss refers to the dropping of audio data packets.
[0078] In some embodiments, the processor can perform linear fitting analysis on the amount of data in the buffer using a preset buffer trend analysis algorithm (such as OLS linear fitting) to obtain the buffer change trend.
[0079] Packet loss count refers to the number of audio data packets dropped. Packet loss count due to buffer overflow refers to the number of packet losses caused by buffer overflow. In some embodiments, the processor can determine whether the data volume of the buffer has reached the data storage limit of the buffer through a preset monitoring algorithm (such as event-triggered counting logic), and when the data volume of the buffer has reached the storage limit, it is determined that the current packet loss is a buffer overflow packet loss, and the number of packet losses caused by buffer overflow is counted through a preset statistical method (such as an overflow packet loss counter).
[0080] For an explanation of clock offset, please refer to the relevant description above.
[0081] In some embodiments, the feedback data can be represented in the form of a sequence or a vector. For example, when the buffer change trend is +200KB / s, the number of packet losses is 5, and the clock offset is 0.01ms, the feedback data can be represented as a vector (+200KB / s, 5 times, 0.01ms).
[0082] In some embodiments, as described above, based on the received audio data packets, the receiving end generates buffer trend information by real-time monitoring of changes in the amount of data in the buffer, counts the number of buffer overflows and packet losses, and calculates the clock offset by combining timestamp exchange, together forming feedback data.
[0083] In some embodiments, the processor can determine feedback data based on the received audio data using a triple feedback mechanism. This triple feedback mechanism includes: analyzing buffer change trends; counting the number of packet losses caused by buffer overflows; and monitoring clock offsets. For more information on how to generate feedback data using a triple feedback mechanism, please refer to [link to relevant documentation]. Figure 4 And its related descriptions.
[0084] Step 340: Determine performance tuning strategies based on network quality scores and feedback data.
[0085] Performance tuning strategies are strategies for adjusting audio transmission parameters to achieve a balance between audio transmission quality and bandwidth resource utilization.
[0086] Performance tuning strategies can include degradation strategies and upgrade strategies. Degradation strategies include reducing the number of transmission channels and / or lowering the sampling rate to proactively reduce the amount of data and reduce audio quality when network bandwidth deteriorates, in order to prioritize playback continuity. Upgrade strategies include increasing the number of transmission channels and / or increasing the sampling rate to increase the amount of data when network bandwidth improves, maximizing the use of the network window and improving audio quality.
[0087] The processor can determine the corresponding performance adjustment strategies based on the feedback data generated by each receiver; or it can determine the overall performance adjustment strategy based on the feedback data generated by all receivers.
[0088] In some embodiments, the processor may use network quality scores as triggering conditions and feedback data as the basis for strategy selection to determine a performance adjustment strategy according to a second preset table. For example, when the network quality score is lower than a score threshold, the processor may determine a performance adjustment strategy by querying the second preset table based on the feedback data.
[0089] In some embodiments, the rating threshold can be set based on experience. For more information on rating thresholds, please see [link to relevant documentation]. Figure 5 And its related descriptions.
[0090] The second preset table can be manually set based on historical data and / or experience. The second preset table includes different types of feedback data and corresponding performance tuning strategies. For example, when the buffer change trend in the feedback data is continuously decreasing, the performance tuning strategy could be to reduce the sampling rate and increase the preloading duration; when the number of packet losses in the feedback data exceeds a threshold, the performance tuning strategy could be to reduce the number of transmission channels and prioritize the transmission of high-priority packets; when the clock offset in the feedback data exceeds an offset threshold, the performance tuning strategy could be to perform local clock calibration and reduce the number of transmission channels for devices with clock synchronization failures. More information on preloading and high-priority packets can be found in the relevant descriptions below.
[0091] In some embodiments, in response to a network quality score below a score threshold, the processor can reduce the number of transmission channels and decrease the sampling rate; in response to a network quality score above the score threshold, the processor can increase the number of transmission channels and increase the sampling rate; wherein the changes in the number of transmission channels and the changes in the sampling rate are determined based on feedback data. For more information on this section, please refer to [link to relevant documentation]. Figure 5 And its related descriptions.
[0092] In some embodiments, in response to a network quality score below a score threshold, the processor can perform multiple downgrade adjustments to the packet parameters, each downgrade adjustment including reducing the packet size and shortening the transmission interval; in response to a network quality score above a score threshold, the processor can perform multiple upgrade adjustments to the packet parameters, each upgrade adjustment including increasing the packet size and extending the transmission interval; wherein, the packet size and transmission interval are determined based on feedback data, and the adjustment of the packet size adopts a gradual, curve-like change to avoid buffer jumps. For more information on this section, please refer to [link to relevant documentation]. Figure 6 And its related descriptions.
[0093] Step 350: Adjust audio transmission parameters based on performance tuning strategies.
[0094] Audio transmission parameters refer to the controllable operational variables during audio transmission, including the number of transmission channels, sampling rate, and packet parameters.
[0095] The number of transmission channels refers to the number of independent audio tracks that can be played simultaneously. Examples include mono, stereo, and surround sound. The more transmission channels, the larger the amount of audio data transmitted, and the greater the network bandwidth required.
[0096] Sampling rate refers to the amount of audio data collected by the transmitting end per unit of time. For example, 48kHz. The higher the sampling rate, the larger the amount of audio data sampled per unit of time, the greater the network transmission pressure, and the greater the network bandwidth consumption.
[0097] Packet parameters refer to the operational parameters used to segment audio data into audio data packets. In some embodiments, packet parameters may include packet size and transmission interval. Packet size refers to the data size of a single audio data packet. Transmission interval refers to the time interval between transmissions of adjacent audio data packets. The transmission interval is positively correlated with the packet size. The larger the packet size and the larger the transmission interval, the less network transmission pressure.
[0098] In some embodiments, the processor can autonomously adjust audio transmission parameters based on performance tuning strategies, according to corresponding tuning amounts (such as changes in the number of transmission channels and changes in the sampling rate) and tuning methods (such as adjusting the packet size using a gradual, curved change).
[0099] In some embodiments of this specification, by continuously evaluating the network status and adjusting the transmission strategy in real time, a dynamic balance between audio quality and network resources is intelligently maintained. This not only prioritizes sound clarity and playback continuity but also keeps bandwidth consumption within a reasonable range, effectively solving the problem that traditional transmission schemes struggle to balance efficiency and stability in complex network environments.
[0100] Figure 4 This is a schematic diagram illustrating the generation of feedback data according to some embodiments of this specification.
[0101] In some embodiments, such as Figure 4 As shown, the receiving end can determine the feedback data 430 based on the received audio data 410 through a triple feedback mechanism 420. The triple feedback mechanism 420 includes: analyzing the trend of buffer change 421; counting the number of packet losses caused by buffer overflow 422; and monitoring the clock offset 423.
[0102] For more information on audio data, feedback data, buffer change trends, packet loss counts, and clock offsets, please refer to [link / reference]. Figure 3 Related content.
[0103] In some embodiments, the triple feedback mechanism can be an embedded algorithm at the receiver. For example, the triple feedback mechanism may include... Figure 3The buffer trend analysis algorithm, overflow packet loss counter, and clock offset monitoring protocol described herein are used to analyze the buffer change trend, count the number of packet losses caused by buffer overflow, and monitor the clock offset, respectively.
[0104] In some embodiments of the specification, by collaboratively analyzing three sets of core parameters—buffer change trends, packet loss counts, and clock offset—transmission bottlenecks in the audio data transmission process can be accurately located (e.g., predicting network bandwidth bottlenecks by real-time monitoring of buffer trends, quantifying local processing bottlenecks by combining packet loss counts, and calibrating playback synchronization with clock offset). This facilitates precise control of audio transmission parameters when optimizing audio stream data transmission in the future, significantly improving playback continuity and synchronization stability under weak network conditions, and enhancing the targeting and effectiveness of optimization operations.
[0105] In some embodiments, in response to a network quality score meeting a preset condition, the sending end preloads audio data to the receiving end, and the preloading duration is determined based on the buffer consumption rate of the receiving end.
[0106] For more information on network quality scores, audio data, and buffer consumption rates, please see [link to relevant documentation]. Figure 3 And its related descriptions.
[0107] Preset conditions refer to the constraints that the network quality score must meet to trigger the preloading of audio data. In some embodiments, the preset condition may be that the difference between the network quality score and a scoring threshold is greater than a first preset difference threshold, which can be manually preset. For an explanation of the scoring threshold, please refer to [link to relevant documentation]. Figure 5 Related descriptions.
[0108] If the network quality score meets the preset conditions, it means that the current network status is good (such as sufficient network bandwidth, low latency, and controllable packet loss rate). The sending end can preload audio data to the receiving end and use idle bandwidth to build buffer redundancy to counteract potential network fluctuations in the future.
[0109] The preloading duration for preloaded audio data can be determined by consulting a third preset table based on the buffer consumption rate at the receiving end. This third preset table contains the correspondence between buffer consumption rate and preloading duration; the faster the buffer consumption rate and the faster the audio output rate, the longer the preloading duration. The third preset table can be set manually based on experience.
[0110] In some embodiments of this specification, preloading audio data when the network quality score meets preset conditions can enable over-transmission of data during the high-quality network window, avoid bandwidth idleness, maximize bandwidth utilization, and significantly improve playback continuity under sudden weak network conditions. In addition, the preloading duration of the preloaded audio data is determined based on the buffer consumption rate of the receiving end, realizing intelligent bandwidth adaptation, preventing data redundancy, and achieving both high efficiency and robustness.
[0111] In some embodiments, there are three known scenarios: the amount of data in the buffer is below the safety threshold of the dynamic tolerance range, the amount of data in the buffer is above the saturation threshold of the dynamic tolerance range, and the amount of data in the buffer is within the dynamic tolerance range.
[0112] In some embodiments, the indicator detection module can monitor the amount of data in the buffer; in response to the amount of data in the buffer being lower than the safety threshold of the dynamic tolerance range, it can trigger a progressive sampling rate conversion to smoothly reduce the sampling rate; in response to the amount of data in the buffer being higher than the saturation threshold of the dynamic tolerance range, it can limit the amount of preloading.
[0113] The amount of data stored in the buffer refers to the amount of real-time audio data stored in the buffer.
[0114] In some embodiments, the processor can monitor the buffer in real time to obtain the amount of data in the buffer.
[0115] The dynamic tolerance range refers to the safe range of buffer data volume. When the buffer data volume is within the dynamic tolerance range, it means that the buffer data volume can both ensure playback continuity and avoid buffer overflow.
[0116] The dynamic tolerance range has a lower limit of a safety threshold and an upper limit of a saturation threshold. The safety threshold is the minimum amount of data required by the buffer to prevent playback interruption. For example, the safety threshold could be the minimum amount of data required for decoding a single playback. The saturation threshold is the maximum amount of data the buffer is constrained to prevent data overflow. For example, the saturation threshold could be the maximum amount of data that can be handled by sudden network fluctuations. The specific values of the safety threshold and the saturation threshold can be preset manually.
[0117] When the amount of data in the buffer is below the safety threshold, it means there is insufficient data in the buffer, posing a risk of audio dropout. In this case, the processor can trigger a progressive sampling rate conversion to smoothly reduce the sampling rate. When the amount of data in the buffer is above the saturation threshold, it means there is excessive data in the buffer, posing a risk of overflow. In this case, the processor can limit the preload amount.
[0118] Progressive sampling rate conversion refers to smoothly reducing the sampling rate in a decreasing manner. The decreasing method includes, but is not limited to, smoothly decreasing according to a preset curve, in order to avoid abrupt changes in audio.
[0119] In some embodiments, when the amount of data in the buffer exceeds the saturation threshold of the dynamic tolerance range, the processor can control the transmitter to reduce the number of preloaded audio data packets sent based on a reduction amount, thereby limiting the preload amount. The preload amount refers to the amount of preloaded audio data. The reduction amount can be a reduction percentage. For example, a reduction of 5% means that the amount of preloaded audio data is reduced to 95% of its previous value.
[0120] The reduction amount can be negatively correlated with the percentage of remaining space in the buffer. The percentage of remaining space is the ratio of the remaining buffer space to the total buffer capacity, which is the difference between the total buffer capacity and the current amount of data in the buffer. A larger percentage of remaining space means more available buffer capacity, so the reduction amount should be smaller accordingly to prevent excessive suppression of preloading from leading to idle bandwidth resources.
[0121] In some embodiments of the specification, bidirectional adjustment is precisely triggered by dynamic tolerance intervals: during low buffering, the sampling rate is gradually reduced to prevent playback interruption and ensure audio continuity; during high buffering, intelligent preload limiting is used to prevent memory overflow and avoid forced packet loss, thereby achieving an efficient balance between bandwidth and memory resources.
[0122] In some embodiments, there are two scenarios: mandatory packet loss and optional packet loss. When there is a risk of buffer overflow (e.g., the buffer data size exceeds a saturation threshold), packet loss is mandatory. When there is no risk of buffer overflow (e.g., the buffer data size is within a dynamic tolerance range) and the network condition is good (e.g., the network quality score exceeds a score threshold), packet loss is optional.
[0123] In some embodiments, the processor can prioritize audio data packets based on their content characteristics, dividing them into high-priority packets and low-priority packets; in response to the necessity of packet loss, low-priority packets are discarded first. For information on the content of audio data packets, please refer to [link to relevant documentation]. Figure 3 Related descriptions.
[0124] Content characteristics refer to the distinguishing features that differentiate audio content.
[0125] In some embodiments, content characteristics include dominant frequencies and volume characteristics. Dominant frequencies refer to frequencies in the audio data where energy is concentrated or prominent, while volume characteristics refer to the intensity of the sound signal.
[0126] In some embodiments, the processor can detect the dominant frequency and volume characteristics of each audio data packet using a real-time audio analysis algorithm (such as a real-time speech recognition algorithm) to determine the content characteristics of the audio data packet.
[0127] A high-priority packet is an audio data packet with high priority. In some embodiments, a high-priority packet may be an audio data packet that mainly contains key audio elements (such as human voice). Discarding a high-priority packet will result in the loss of key audio elements, which will directly affect the effective dissemination of audio content or the user's listening experience.
[0128] Low-priority packets refer to audio data packets with low priority. In some embodiments, low-priority packets may be audio data packets that primarily contain secondary audio elements (such as harmony and background sounds). Discarding low-priority packets will result in the loss of secondary audio elements, but has a minimal impact on overall audio propagation or sound quality.
[0129] In some embodiments, the processor can determine whether an audio data packet mainly includes key audio elements or secondary audio elements based on the content characteristics of the audio data. Audio data packets mainly containing key audio elements are designated as high-priority packets, and those mainly containing secondary audio elements are designated as low-priority packets. Specifically, the dominant frequency of key audio elements is the human voice frequency (e.g., 200Hz-2500Hz), and the volume characteristic is that the volume exceeds a preset volume threshold (e.g., 20dB). The dominant frequency of secondary audio elements is the background audio frequency (e.g., below 80Hz), and the volume characteristic is that the volume is below a preset volume threshold.
[0130] In some embodiments, in response to the necessity of packet loss, the processor may prioritize dropping low-priority packets.
[0131] In some embodiments, by prioritizing audio data packets, low-priority packets can be discarded first when packet loss is necessary, thereby reducing the risk of losing high-priority packets, prioritizing the transmission of core audio, and balancing bandwidth efficiency and smooth listening experience.
[0132] In some embodiments, there are known scenarios where data units are lost or not. Data unit loss means that some data failed to reach the receiving end during audio data transmission, potentially causing audio playback stuttering or missing information. In this case, data retransmission can be triggered or an interpolation compensation algorithm can be used to complete the data. No data unit loss indicates that the audio data transmission is stable and complete, and the receiving end can directly parse and play the audio data in sequence.
[0133] In some embodiments, the indicator detection module can split the audio data into multiple data units of fixed duration during the audio data transmission stage; in response to the loss of data units, it can perform high-precision interpolation compensation on the lost data units based on the data units that were not lost before and after.
[0134] A data unit refers to an audio segment that audio data is broken down into during transmission. As the smallest unit of audio data during transmission, an audio data packet, serving as a network transmission carrier, can encapsulate one or more data units.
[0135] In some embodiments, the data unit is of fixed duration, which can be set manually based on experience, for example, 5ms. In some embodiments, the fixed duration can also be positively correlated with the network quality score; the higher the network quality score, the longer the fixed duration.
[0136] In some embodiments, before transmitting audio data, the transmitting end can divide the audio data into multiple consecutive audio segments, i.e., multiple data units, according to a fixed duration. The division method includes, but is not limited to, equal temporal division (such as dividing into 5ms frames of fixed duration).
[0137] High-precision interpolation compensation refers to the process of using interpolation techniques to compensate for lost data units with high precision.
[0138] In some embodiments, the processor can reconstruct the lost data unit using features of the audio data of the preceding and following unlost units through an interpolation algorithm. The features of the audio data include, but are not limited to, waveform features, spectral features, and phase features. These features can be obtained based on feature extraction algorithms. Feature extraction algorithms include, but are not limited to, Fast Fourier Transform (FFT) and Linear Predictive Coding (LPC). Interpolation algorithms include, but are not limited to, linear time-domain interpolation, cubic spline curves, and frequency-domain phase reconstruction.
[0139] In some embodiments of this specification, refined transmission is achieved by splitting fixed-duration data units, and high-precision interpolation compensation is performed by combining the audio data characteristics of the preceding and following units, reducing audio damage caused by packet loss, significantly weakening dropouts and abrupt changes in timbre, and effectively reducing the perceptual distortion rate of the human ear.
[0140] Figure 5 This is a schematic diagram illustrating a performance tuning strategy according to some embodiments of this specification.
[0141] In some embodiments, performance tuning strategies may include one or more of the following: increasing the number of transmission channels, decreasing the number of transmission channels, increasing the sampling rate, and decreasing the sampling rate. For details regarding performance tuning strategies, the number of transmission channels, and the sampling rate, please refer to [link to relevant documentation]. Figure 3 Related descriptions.
[0142] In some embodiments, there are three known scenarios: network quality score below a scoring threshold, network quality score above a scoring threshold, and network quality score equal to a scoring threshold.
[0143] A scoring threshold is a critical value used to determine whether network quality scoring should be triggered. The scoring threshold can be set by the processor by default or preset manually based on experience. For example, if the network quality score is a value in the range of 0-10, then the scoring threshold could be 7 points.
[0144] A network quality score equal to the score threshold means that the current network condition is just right to maintain the current audio transmission, and no performance adjustment strategy needs to be activated. A network quality score below the score threshold means that the current network condition is poor and cannot maintain the current audio transmission, and the degradation strategy of the performance adjustment strategy needs to be activated to prioritize the normal transmission of audio (to ensure playback continuity and avoid problems such as dropouts). A network quality score above the score threshold means that the current network condition is good and can maintain the current audio transmission, and the upgrade strategy of the performance adjustment strategy can be activated at this time to improve audio quality.
[0145] like Figure 5 As shown, in response to a network quality score below the score threshold, the number of transmission channels is reduced and the sampling rate is lowered; in response to a network quality score above the score threshold, the number of transmission channels is increased and the sampling rate is raised; wherein, the amount of change in the number of transmission channels and the amount of change in the sampling rate are determined based on feedback data.
[0146] For more information on network quality scores and feedback data, please refer to [link / reference]. Figure 3 And its related descriptions.
[0147] The change in the number of transmission channels includes the direction of change (increase or decrease) and the amount of increase or decrease. The change in the sampling rate includes the direction of change (increase or decrease) and the amount of increase or decrease.
[0148] In some embodiments, the processor can determine the changes in the number of transmission channels and the changes in the sampling rate based on feedback data in various ways. For example, the processor can determine the changes based on feedback data (buffer change trends, number of packet losses due to buffer overflows, and clock offsets) according to... Figure 3 The second preset table described in the document determines the direction of change and processing priority of the number of transmission channels, as well as the direction of change and processing priority of the sampling rate. The higher the processing priority, the earlier the processing order (first priority > second priority > third priority).
[0149] For example, when the network quality score is lower than the score threshold, the second preset table may include the following: when the number of packet losses due to buffer overflow exceeds the threshold, significantly reduce the number of transmission channels while maintaining the sampling rate or slightly reduce the sampling rate with the processing priority set to first priority; when the buffer change trend shows a continuous decrease in data volume, reduce the number of transmission channels and gradually and smoothly reduce the sampling rate with the processing priority set to second priority; when the clock offset exceeds the offset threshold, slightly reduce the number of transmission channels while maintaining the sampling rate with the processing priority set to third priority. A continuous decrease can be a continuous decrease within a preset sampling period or a continuous decrease after M consecutive samples. Both the preset sampling period and the value of M can be manually preset.
[0150] When the network quality score is higher than the score threshold, the second preset table can include the following: If the number of packet losses due to buffer overflow is lower than the number of losses threshold, significantly increase the number of transmission channels while keeping the sampling rate unchanged or slightly increase the sampling rate, and the processing priority is first priority; if the buffer change trend is that the data volume continues to increase, increase the number of transmission channels and gradually and smoothly increase the sampling rate, and the processing priority is second priority; if the clock offset does not exceed the offset threshold, keep the number of transmission channels and the sampling rate unchanged, and the processing priority is third priority. The offset threshold and the number of losses threshold can be preset manually.
[0151] The change in the number of transmission channels is related to the number of packet losses caused by buffer overflow. For example, the change in the number of transmission channels can be the rounded-up value of the product of the base channel step size and the increase in the packet loss rate.
[0152] The basic channel step size refers to the smallest unit of change used to adjust the number of transmission channels. In some embodiments, the basic channel step size can be preset manually based on system performance; for example, the larger the total number of channels in the system, the larger the basic channel step size can be.
[0153] The increase in packet loss rate refers to the difference between the number of packet losses due to buffer overflow in the current statistical period and the number of packet losses due to buffer overflow in the previous statistical period, and the ratio of the difference to the number of packet losses due to buffer overflow in the previous statistical period.
[0154] It's important to note that reducing the number of transmission channels is not achieved by randomly merging channels. Instead, it involves combining the spatial location information of the receivers and prioritizing the merging of transmission channels that are spatially close or have lower transmission priority. For example, multiple receivers for left surround sound can share a single channel. Multiple receivers with lower transmission priority can also share a single channel. The transmission priority of each receiver can be pre-set manually based on its spatial location information.
[0155] The amount of change in the sampling rate is related to the rate at which the buffer is consumed. For example, the amount of change in the sampling rate can be the product of the base sampling rate step size and the rate at which the buffer is consumed.
[0156] The base sampling rate step size refers to the step size used to adjust the sampling rate. In some embodiments, the base sampling rate step size can be set manually based on historical experience. For an explanation of buffer consumption rates, please refer to [link to relevant documentation]. Figure 3 Related descriptions.
[0157] In some embodiments, the amount of variation in the sampling rate is also related to the hardware capabilities of the receiver. For example, if a receiver only supports a maximum sampling rate of 44.1K, the transmitter will directly send 44.1K data without transmitting a higher sampling rate.
[0158] It should be noted that when multiple feedback conflicts occur, the processor can also determine the changes in the number of transmission channels and the sampling rate based on the conflict coverage principle. For example, when the network quality score is lower than the score threshold, the number of packet losses caused by buffer overflow exceeds the threshold, and the buffer change trend shows a continuous decrease in data volume, meaning that the feedback data includes both first and second priority, then the performance adjustment strategy corresponding to the highest priority will be executed first according to the conflict coverage principle. This means significantly reducing the number of transmission channels while keeping the sampling rate unchanged or slightly reducing the sampling rate.
[0159] In some embodiments of this specification, by establishing a dynamic closed-loop adjustment mechanism between performance adjustment strategy, receiver feedback data and the number of transmission channels and sampling rate, it is possible to actively adapt to limited bandwidth by intelligently reducing audio parameters when network conditions are poor, prioritizing smooth playback; when the network recovers, it can automatically increase parameters to restore high-quality sound, ensuring that the adjustment process itself will not be interrupted and will not cause unnecessary damage to the user experience.
[0160] In some embodiments, there are two known cases: the difference between the network quality score and the score threshold is consistently less than a preset difference threshold within a preset time period, and the difference is not less than a preset difference threshold.
[0161] The preset time period refers to the period during which the network quality score fluctuates relative to the score threshold. It can be a short time length that can be preset manually or by a processor, such as 0.5s, 1s, 2s, or other time lengths.
[0162] The difference between the network quality score and the scoring threshold can be the absolute value of the difference between the network quality score and the scoring threshold.
[0163] The preset difference threshold is a threshold used to assess whether the difference between the network quality score and the score threshold is stable. In some embodiments, the preset fluctuation threshold can be preset manually based on historical experience or set by the processor by default.
[0164] If the difference between the network quality score and the score threshold remains consistently no less than a preset difference threshold within a preset time period, it indicates that the network quality score fluctuates significantly relative to the score threshold, requiring timely network performance adjustments. In other words, the processor can adjust according to... Figure 5 and / or Figure 6 The method shown determines the performance tuning strategy. If the difference between the network quality score and the score threshold remains less than a preset difference threshold within a preset time period, it indicates that the network quality score fluctuates only slightly relative to the score threshold. This means the current network state can basically maintain audio transmission, i.e., the audio transmission parameters can remain unchanged, and the performance tuning strategy is not activated for the time being. If the difference between the network quality score and the score threshold remains no less than the preset difference threshold within a preset time period, this can be defined as K consecutive calculations within the preset time period where the difference is less than the preset difference threshold, and the value of K can be manually preset.
[0165] In some embodiments, the performance adjustment module further includes a hysteresis response mechanism, which includes maintaining the audio transmission parameters unchanged in response to the fact that the difference between the network quality score and the score threshold is continuously less than a preset difference threshold within a preset time period.
[0166] For explanations regarding network quality scoring, scoring thresholds, and audio transmission parameters, please refer to [link / reference]. Figure 3 Related descriptions.
[0167] The hysteresis response mechanism is a mechanism to prevent excessively frequent switching of performance tuning strategies. By delaying the triggering of performance tuning strategies, the mechanism executes the strategy only when the network quality score consistently falls below a threshold and the difference between the network quality score and the threshold exceeds a preset difference threshold. This avoids frequent switching due to network fluctuations and ensures the stability and smoothness of audio parameter adjustments.
[0168] In some embodiments described herein, a hysteresis response mechanism is introduced to effectively filter out brief, meaningless transient jitter in the network, ensuring that the system only responds to real and continuous trend changes, and avoiding frequent switching of audio transmission parameters due to transient network fluctuations.
[0169] In some embodiments, the performance tuning module is further configured to dynamically adjust the timing of the performance tuning strategy based on the user's preset listening priority or the user's historical behavior data, wherein the preset listening priority includes the historical records of the analysis buffer state and the historical records of manual intervention.
[0170] A user is a person who interacts with a system that balances network bandwidth. Examples include sound engineers, broadcast technicians, and ordinary listeners.
[0171] Preset listening preferences refer to listening preferences set by the user in advance, such as prioritizing sound quality or smoothness. In some embodiments, the user can input the preset listening preferences into the system through an interactive terminal (such as a mobile phone or computer) that communicates with the processor.
[0172] Historical behavior records can include historical records of the analysis buffer state and historical records of user manual interventions.
[0173] Historical analysis of buffer states refers to the analysis data of the historical state of the receiver's buffers. For example, historical analysis of buffer states can include the amount of buffer data and the buffer consumption rate at multiple historical time points.
[0174] The history of manual intervention refers to the history of user-manual overwriting or correction of system adjustment results. For example, if the system detects that the rate of increase in buffer data volume exceeds the growth threshold and automatically reduces the sampling rate from 96kHz to 48kHz, the user can manually force the sampling rate to be restored to 96kHz.
[0175] The timing for initiating a performance tuning strategy refers to the decision of when to start the performance tuning strategy. Adjusting the timing for initiating a performance tuning strategy may include adjusting the scoring threshold corresponding to the network quality score, adjusting the threshold corresponding to the number of packet losses caused by buffer overflow, etc.
[0176] In some embodiments, the processor can dynamically adjust the timing of performance tuning strategy activation by activating a tuning mechanism based on the user's preset listening priority or the user's historical behavior data.
[0177] For example, the activation adjustment mechanism could include: when the user's preset listening priority is sound quality priority, the processor will automatically delay the activation of performance adjustment strategies, such as lowering the scoring threshold by a preset amount (e.g., from 6 points to 5 points). Lowering the scoring threshold means that only a worse network will trigger degradation (e.g., reducing the number of transmission channels, lowering the sampling rate, etc.) to protect sound quality; conversely, when the user's preset listening priority is smoothness priority, the processor will activate the mechanism earlier, such as raising the scoring threshold by a preset amount. Raising the scoring threshold means earlier degradation to ensure smooth playback. The preset amount can be set manually based on historical experience.
[0178] For example, the activation adjustment mechanism can also include: adjusting the activation timing of the performance adjustment strategy based on historical behavioral data. For instance, if the scoring threshold is 6, but the probability of users manually canceling the system's automatic downgrade operation is greater than 90% when the network quality score is between 6 and 6.5, indicating that users tolerate low-scoring networks and care more about sound quality, then the scoring threshold can be lowered from 6 to 5.5.
[0179] In some embodiments of this specification, the timing of strategy activation is dynamically adjusted based on the user's preset auditory priority or the user's historical behavior data, thereby achieving personalized adaptive adjustment, reducing manual intervention and avoiding rigid "one-size-fits-all" strategies, and improving the system intelligence in complex scenarios.
[0180] In some embodiments, the performance tuning module is further configured to: predict the bandwidth fluctuation trend in the future period based on network bandwidth data using a time series prediction model; and, in response to a positive bandwidth fluctuation trend in the future period, initiate a degradation strategy in the performance tuning strategy in advance, wherein the degradation strategy includes reducing the number of transmission channels and / or reducing the sampling rate.
[0181] Network bandwidth data refers to the total bandwidth utilization across the entire network link. The entire link includes the system balancing network transmission bandwidth and other software (such as browsers) on the same network link. In other words, network bandwidth data is the ratio of the bandwidth used by the system balancing network transmission bandwidth to the total network bandwidth (sum of bandwidth used by the other software). For example, if the total network bandwidth is 100Mbps, the system balancing network transmission bandwidth uses 20Mbps, and other software uses 60Mbps, then the network bandwidth data is 80%.
[0182] In some embodiments, the processor can use a network monitoring agent deployed at the sending end to query the traffic counters of switch or router ports using standard network management protocols to obtain network bandwidth data.
[0183] In some embodiments, network bandwidth data can be a sequence of total end-link bandwidth utilization rates at multiple historical time points within a historical period. A historical period can refer to a historical time span preceding the current moment, and the duration of the historical period and the selection of historical time points can be preset manually.
[0184] Bandwidth fluctuation trend refers to the relative rate of change of the total bandwidth utilization across the entire link. The bandwidth fluctuation trend within a future time period is the ratio of the difference between the total bandwidth utilization at the end of the future time period and the total bandwidth utilization at the beginning of the future time period, to the total bandwidth utilization at the beginning of the future time period, including both positive and negative values. The future time period can be preset manually. For example, if the future time period is 15 seconds after the current time, and the total bandwidth utilization increases from 60% to 72% within 15 seconds, then the bandwidth fluctuation trend is +20%.
[0185] A negative bandwidth fluctuation trend indicates a decrease in the total bandwidth utilization across the entire link, signifying improved bandwidth. This allows for delaying the activation of degradation strategies within the performance tuning strategy or activating upgrade strategies earlier. Conversely, a positive bandwidth fluctuation trend indicates an increase in the total bandwidth utilization across the entire link, signifying bandwidth degradation. This allows for earlier activation of degradation strategies within the performance tuning strategy. Degradation strategies include reducing the number of transmission channels and / or lowering the sampling rate to proactively reduce data volume and audio quality when network bandwidth deteriorates, prioritizing playback continuity. Upgrade strategies include increasing the number of transmission channels and / or increasing the sampling rate to increase data volume when network bandwidth improves, maximizing network window utilization and improving audio quality.
[0186] In some embodiments, the processor can process network bandwidth data and network quality scores based on a time series prediction model to obtain bandwidth fluctuation trends.
[0187] A time series forecasting model is a model used to predict trends in bandwidth fluctuations. In some embodiments, the time series forecasting model is a machine learning model. For example, it may be one or more combinations of Long Short-Term Memory (LSTM), Recurrent Neural Network (RNN), or other custom models.
[0188] In some embodiments, the input to a time series forecasting model may include network bandwidth data and network quality scores, and the output may be a bandwidth fluctuation trend.
[0189] For an explanation of network quality scoring, please refer to [link / reference]. Figure 3 Related descriptions.
[0190] In some embodiments, a time series prediction model can be trained using multiple labeled training samples. For example, a processor can input multiple labeled training samples into an initial time series prediction model, construct a loss function using the labels and the output of the initial time series prediction model, and iteratively update the parameters of the initial time series prediction model based on the loss function using methods such as gradient descent. When a preset condition is met, the model training is complete, resulting in a trained time series prediction model. The preset condition could be loss function convergence, the number of iterations reaching a threshold, etc.
[0191] A training sample may include sample network bandwidth data, sample network quality score, and a label representing the actual bandwidth fluctuation trend corresponding to that training sample.
[0192] Training samples can be obtained based on historical data, and labels can be obtained through manual annotation. For example, training samples and labels can be constructed as follows: obtain network bandwidth data and network quality scores for a first time period from historical data, as sample network bandwidth data and sample network quality scores; obtain the historical actual bandwidth fluctuation trend obtained from actual monitoring within a second time period, as the labels corresponding to the training samples. The second time period is later than the first time period.
[0193] In some embodiments, when the bandwidth fluctuation trend output by the time series prediction model for a future period is positive, i.e., when the processor predicts an impending bandwidth degradation trend, the performance tuning module will initiate the corresponding performance tuning strategy in advance. The performance tuning strategy is determined based on network quality scores and feedback data; for more information, please refer to [link to relevant documentation]. Figure 3 Related descriptions.
[0194] In some embodiments of this specification, a time series prediction model is introduced to predict potential future bandwidth degradation, so as to initiate performance adjustment strategies in advance before bandwidth utilization spikes. This proactive adjustment can seize the time window to achieve a seamless transition, effectively avoiding problems such as audio stuttering and dropouts.
[0195] Figure 6 This is a schematic diagram illustrating another performance tuning strategy according to other embodiments of this specification.
[0196] In some embodiments, such as Figure 6 As shown, the performance adjustment module is further configured to: in response to a network quality score lower than a score threshold, perform multiple downgrade adjustments to packet parameters, including reducing packet size and shortening the transmission interval; in response to a network quality score higher than a score threshold, perform multiple upgrade adjustments to packet parameters, including increasing packet size and extending the transmission interval; wherein, packet size and transmission interval are determined based on feedback data, and the adjustment of packet size adopts a gradual curve-like change.
[0197] For more information on network quality scoring, feedback data, scoring thresholds, packet parameters, packet size, and sending intervals, please refer to [link to relevant documentation]. Figure 3 Related descriptions.
[0198] The gradual change in packet size refers to using a smooth curve (such as an S-shaped function) to transition when adjusting the packet size (e.g., 20ms→18ms→15ms→12ms→10ms) to avoid sudden changes in the amount of data in the buffer caused by abrupt changes in packet size, which could lead to playback stuttering or overflow.
[0199] In some embodiments, the start time of the curve-like gradual change is determined based on the bandwidth fluctuation trend in the future period.
[0200] In some embodiments, the processor may determine the start time of the gradual, curvilinear change based on a preset formula. For example, the preset formula may be: T start = T d - T buffer Among them, T start Indicates the start time of the gradual, curvilinear change; T d This indicates the time when a deterioration trend (i.e., a positive bandwidth fluctuation trend) occurs within the predicted bandwidth fluctuation trend; T buffer This indicates the safety buffer time, which can be preset manually.
[0201] For example, if the network in a bandwidth fluctuation trend will begin to degrade 5 seconds after the current moment, and the safety buffer time is 3 seconds, then the start time of the gradual, curvilinear change could be 2 seconds after the current moment. It should be noted that if the start time T of the gradual, curvilinear change... start If the current time is earlier than or equal to the current time, the packet size will be adjusted immediately according to a gradual, curved change. If the current time reaches T... d If no deterioration trend appears in the bandwidth fluctuation trend (i.e., the bandwidth fluctuation trend remains negative), then the adjustment of packet parameters is stopped.
[0202] In some embodiments, the security buffer time can be positively correlated with the adjustment amount of the packet size and the adjustment amount of the transmission interval. The specific correlation can be preset manually.
[0203] In some embodiments of this specification, when the network quality score is lower than the score threshold, a downgrade adjustment is performed to improve real-time performance and packet loss resistance; when the network quality score is higher than the score threshold, an upgrade adjustment is performed to improve bandwidth utilization. Furthermore, by combining the bandwidth fluctuation trend predicted by the model with a curve-based gradual change adjustment method, and combining feedback data for dynamic calibration, millimeter-level matching between the transmission strategy and the receiver status can be achieved, improving the user's listening experience.
[0204] In some embodiments, the processor can determine the packet size and transmission interval based on feedback data in various ways, including the direction of change (increase or decrease) of the packet size and transmission interval and the specific adjustment amount.
[0205] For example, the processor can determine the direction of packet size change and processing priority, as well as the direction of transmission interval change and processing priority, based on feedback data (buffer change trend, number of packet losses due to buffer overflow, and clock offset) according to a fourth preset table. Among them, the higher the processing priority, the earlier the processing order (first priority > second priority > third priority).
[0206] For example, the fourth preset table may include: when the number of packet losses caused by buffer overflow exceeds the threshold, the packet size and transmission interval are reduced, and the processing priority is the first priority; when the buffer change trend is that the data volume continues to decrease, the packet size remains unchanged and the transmission interval is significantly reduced, and the processing priority is the second priority; when the clock offset exceeds the offset threshold, the packet size and transmission interval are increased, and the processing priority is the third priority.
[0207] The fourth preset table may also include specific adjustments (including increases and decreases) to packet size and transmission interval. The fourth preset table can be preset manually.
[0208] It should be noted that when multiple feedback conflicts occur, the processor can also determine the specific adjustment amount of packet size and transmission interval based on the conflict coverage principle. For example, if the number of packet losses caused by buffer overflow exceeds the threshold, and the buffer changes with a continuous decrease in data volume, meaning that the feedback data includes both first and second priority, then the packet parameter adjustment strategy corresponding to the highest priority will be executed first according to the conflict coverage principle, i.e., reducing the packet size and transmission interval.
[0209] In some embodiments of this specification, the packet size is gradually adjusted by a curve to avoid buffer jumps, and when the network deteriorates, it is downgraded (small packets + high frequency) to reduce the granularity of single packet loss and ensure continuity. When the network is optimized, it is upgraded (large packets + low frequency) to reduce protocol header overhead and improve bandwidth efficiency. Combined with feedback data, the transmission strategy and the receiver status are accurately matched.
[0210] The embodiments in this specification are merely illustrative and not intended to limit the scope of this specification. Various modifications and alterations that can be made by those skilled in the art under the guidance of this specification remain within its scope.
[0211] Furthermore, certain features, structures, or characteristics in one or more embodiments of this specification may be appropriately combined.
[0212] In some embodiments, numbers describing the quantity of components and attributes are used. It should be understood that such numbers used in the description of embodiments are modified in some examples with the terms "approximately," "approximately," or "generally." Unless otherwise stated, "approximately," "approximately," or "generally" indicates that the numbers are allowed to vary by ±20%. Accordingly, in some embodiments, the numerical parameters used in the specification and claims are approximate values, which may be changed depending on the characteristics required by individual embodiments. In some embodiments, numerical parameters should take into account specified significant digits and employ a general method of digit reservation. Although the numerical ranges and parameters used to confirm their breadth of range in some embodiments of this specification are approximate values, in specific embodiments, such values are set as precisely as feasible.
[0213] In the event of any inconsistency or conflict between the descriptions, definitions, and / or terms used in the supplementary materials to this specification and the contents of this specification, the descriptions, definitions, and / or terms used in this specification shall prevail.
Claims
1. A system for balancing network transmission bandwidth, characterized in that, The system includes: The status detection module is configured as follows: Monitor network transmission metrics and clock synchronization information; A network quality score is generated based on the network transmission metrics and the clock synchronization information. The indicator detection module consists of a sender and a receiver, wherein: The transmitting end is configured to transmit audio data; The receiving end is configured to generate feedback data based on the received audio data and send the feedback data back to the sending end; the feedback data includes buffer change trend, number of packet losses caused by buffer overflow, and clock offset; The performance tuning module is configured as follows: Based on the network quality score and the feedback data, a performance tuning strategy is determined; Based on the performance tuning strategy, adjust the audio transmission parameters.
2. The system as described in claim 1, characterized in that, The receiving end is further configured as follows: Based on the received audio data, the feedback data is determined through a triple feedback mechanism, wherein the triple feedback mechanism includes: analyzing the trend of the buffer change; counting the number of packet losses caused by the buffer overflow; and monitoring the clock offset.
3. The system as described in claim 1, characterized in that, The performance tuning module is further configured to: In response to the network quality score falling below the score threshold, the number of transmission channels is reduced and the sampling rate is lowered. In response to the network quality score being higher than the score threshold, the number of transmission channels is increased and the sampling rate is improved; The changes in the number of transmission channels and the changes in the sampling rate are determined based on the feedback data.
4. The system as described in claim 3, characterized in that, The performance tuning module is further configured to: In response to the network quality score being lower than the score threshold, the packet parameters are downgraded and adjusted multiple times, including reducing the packet size and shortening the transmission interval; In response to the network quality score being higher than the score threshold, the packet parameters are upgraded and adjusted multiple times, including increasing the packet size and extending the transmission interval; The packet size and the sending interval are determined based on the feedback data, and the packet size is adjusted using a gradual, curve-like change.
5. A method for balancing network transmission bandwidth, characterized in that, The method includes: Monitor network transmission metrics and clock synchronization information; A network quality score is generated based on the network transmission metrics and the clock synchronization information. Based on the received audio data, feedback data is generated; the feedback data includes the buffer change trend, the number of packet losses caused by buffer overflow, and the clock offset. Based on the network quality score and the feedback data, a performance tuning strategy is determined; Based on the performance tuning strategy, adjust the audio transmission parameters.
6. The method as described in claim 5, characterized in that, The generation of feedback data based on the received audio data includes: Based on the received audio data, the feedback data is determined through a triple feedback mechanism, wherein the triple feedback mechanism includes: analyzing the trend of the buffer change; counting the number of packet losses caused by the buffer overflow; and monitoring the clock offset.
7. The method as described in claim 5, characterized in that, The step of determining a performance tuning strategy based on the network quality score and the feedback data includes: In response to the network quality score falling below the score threshold, the number of transmission channels is reduced and the sampling rate is lowered. In response to the network quality score being higher than the score threshold, the number of transmission channels is increased and the sampling rate is improved; The changes in the number of transmission channels and the changes in the sampling rate are determined based on the feedback data.
8. The method as described in claim 7, characterized in that, The process of determining the performance tuning strategy based on the network quality score and the feedback data further includes: In response to the network quality score being lower than the score threshold, the packet parameters are downgraded and adjusted multiple times, including reducing the packet size and shortening the transmission interval; In response to the network quality score being higher than the score threshold, the packet parameters are upgraded and adjusted multiple times, including increasing the packet size and extending the transmission interval; The packet size and the sending interval are determined based on the feedback data, and the packet size is adjusted using a gradual, curve-like change.
9. An apparatus for balancing network transmission bandwidth, the apparatus comprising at least one processor and at least one memory; The at least one memory is used to store computer instructions; The at least one processor is configured to execute at least a portion of the computer instructions to implement the method as described in any one of claims 5 to 8.
10. A computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, enable the implementation of the method according to any one of claims 5 to 8.