Broadcast low-latency switching method and electronic device
Patent Information
- Application Number
- CN202610880604.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-17
- Publication Date
- 2026-09-29
AI Technical Summary
[0003]本申请实施例提供了一种广播低时延切换方法以及电子设备,可以解决现有切换方式为人工切换,在切换后容易出现音频断点、不同步、延时不可控的问题
本申请提供的广播低时延切换方法包括:用于广播系统中的接收端,接收端与广播系统中的发送端连接,方法,包括:接收音频信号,音频信号是发送端通过IP网络主通道、模拟备份通道同步发送的,同步发送的音频信号设有相同的全局时间戳;基于预设周期检测网络状态,网络状态包括丢包率、网络抖动、心跳包响应时间中的至少一种;若确定所述网络状态满足切换条件,则基于全局时间戳、缓存的音频信号将播放的音频信号切换为模拟备份通道传输的音频信号。本申请实施例能够实现IP网络主通道与模拟备份通道的无缝切换,有效避免音频断点、不同步以及延时不可控的问题,有效满足考试这类对同步低断点的要求。
Smart Images

Figure CN122845489A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of broadcasting technology, and more specifically, to a low-latency broadcasting handover method and an electronic device. Background Technology
[0002] With the deepening of educational informatization, campus broadcasting systems have evolved from traditional fixed-voltage broadcasting to digital broadcasting based on IP networks. Currently, campus broadcasting mainly serves three functions: daily campus broadcasting (such as eye exercises and announcements), classroom sound reinforcement (such as classroom audio systems), and, most importantly, standardized English listening tests (such as the middle school entrance examination and the college entrance examination). English listening tests place extremely high demands on the stability and latency controllability of the broadcasting system. Any fluctuations or even interruptions in the IP network will directly affect the normal conduct of the examination, potentially leading to serious examination incidents. Existing IP broadcasting systems typically only have a backup analog broadcasting channel, but current switching methods are mostly based on manual switching, or there is no audio alignment after triggering the switch. After switching, problems such as audio breakpoints, asynchrony, and uncontrollable latency easily occur, failing to meet the requirements of examinations for low-breakpoint synchronization. Summary of the Invention
[0003] This application provides a low-latency broadcast handover method and electronic device, which can solve the problems of existing handover methods, which are manual and prone to audio interruptions, desynchronization, and uncontrollable delays after handover. To achieve this objective, this application provides the following solutions.
[0004] According to one aspect of the embodiments of this application, a broadcast low-latency handover method is provided for a receiving end in a broadcast system, the receiving end being connected to a transmitting end in the broadcast system, the method comprising: The audio signal is received, which is synchronously transmitted by the transmitting end through the main channel of the IP network and the analog backup channel, and the synchronously transmitted audio signals have the same global timestamp. The network status is detected based on a preset period, and the network status includes at least one of packet loss rate, network jitter, and heartbeat packet response time. If the network status is determined to meet the switching conditions, the played audio signal will be switched to the audio signal transmitted through the analog backup channel based on the global timestamp and the cached audio signal.
[0005] In one possible implementation, the setting of the global timestamp includes: The time synchronization between the transmitting and receiving ends is performed according to a preset time protocol and audio sampling frequency. Once synchronization is confirmed, a global timestamp is generated based on the current system time. This global timestamp is then embedded into the audio signals to be synchronized and transmitted to the main IP network channel and the analog backup channel, enabling the receiving end to use the global timestamp for phase calibration and dynamic compensation.
[0006] In one possible implementation, determining that the network state meets the switching conditions includes: The required level is determined based on the network status, and the level includes any one of the following: early warning level, buffer level, switching level, and circuit breaker level. If the level is determined to be either the switching level or the circuit breaker level, then the switching conditions are met. If the level is determined to be either warning level or buffer level, then the switching conditions are not met, and a preset strategy is executed according to the level that is met. The preset strategy includes preloading and packet loss resistance enhancement.
[0007] In one possible implementation, the step of switching the played audio signal to the audio signal transmitted through the analog backup channel based on the global timestamp and the cached audio signal includes: The switching operation is triggered, the audio signal in the buffer is played, the audio signal is exponentially attenuated based on the global timestamp, and the audio signal of the analog backup channel is exponentially enhanced. The audio signal buffered in the buffer is the audio signal transmitted in the most recent time period of the main channel of the IP network.
[0008] In one possible implementation, the preloading includes: Establish a connection with the simulated backup channel, and determine the size of the target buffer based on the network status; The buffer size is smoothly adjusted based on the target buffer size and the current buffer size, and the final size of the buffer is determined based on the boundary check structure of the adjustment result. The buffer is used to cache the audio signals transmitted through the main channel of the IP network.
[0009] In one possible implementation, the packet loss resistance enhancement includes: Obtain packet loss information based on network status, and determine the redundant packet ratio based on the packet loss information; Redundant packets are generated based on the stated redundancy ratio, and audio signals to be transmitted are generated based on the stated redundancy packets.
[0010] In one possible implementation, the method includes: The type of the currently transmitted audio signal is detected, including exam audio, classroom amplification audio, and daily broadcast audio. If it is determined that the audio signal includes the exam audio, then the audio signals other than the exam audio are muted, the bandwidth of the audio channel corresponding to the exam audio is locked, and a buffer corresponding to the exam audio is created.
[0011] In one possible implementation, the method includes: Collect the ambient noise of the playback environment corresponding to the audio signal, and calculate the gain coefficient based on the ambient noise. The gain of the receiver is determined based on the gain coefficient and the noise differences in different regions so that the receiver can play the audio signal using the gain coefficient.
[0012] According to one aspect of the embodiments of this application, a broadcast low-latency handover method is provided for a transmitter in a broadcast system, the transmitter being connected to a receiver in the broadcast system, the method comprising: Audio signals are synchronously transmitted to the receiving end through the main IP network channel and the simulated backup channel. The synchronously transmitted audio signals have the same global timestamp. The receiving end receives the audio signals and detects the network status based on a preset period. If the network status is determined to meet the switching conditions, the audio signal being played is switched to the audio signal transmitted by the simulated backup channel based on the global timestamp and the cached audio signal. The network status includes at least one of packet loss rate, network jitter, and heartbeat packet response time.
[0013] According to one aspect of the embodiments of this application, an electronic device is provided, including a processor and a memory, wherein the processor is communicatively connected to the memory, the memory stores program data, and the program data is used to execute the method described above.
[0014] The beneficial effects of the technical solutions provided in this application are: The low-latency switching method for broadcasting provided in this application includes: a receiving end in a broadcasting system, the receiving end being connected to a transmitting end in the broadcasting system; the method comprising: receiving an audio signal, the audio signal being synchronously transmitted by the transmitting end through an IP network main channel and an analog backup channel, the synchronously transmitted audio signals having the same global timestamp; detecting network status based on a preset period, the network status including at least one of packet loss rate, network jitter, and heartbeat packet response time; if it is determined that the network status meets the switching conditions, then switching the played audio signal to the audio signal transmitted through the analog backup channel based on the global timestamp and the buffered audio signal. This application embodiment can achieve seamless switching between the IP network main channel and the analog backup channel, effectively avoiding problems such as audio interruptions, asynchrony, and uncontrollable latency, effectively meeting the requirements for low-latency synchronization in examinations and similar events. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below.
[0016] Figure 1 A flowchart of a broadcast low-latency handover method provided in an embodiment of this application; Figure 2 Another flowchart of the broadcast low-latency handover method provided in the embodiments of this application; Figure 3 This is a structural diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0017] The embodiments of this application are described below with reference to the accompanying drawings. It should be understood that the embodiments described below with reference to the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions of the embodiments of this application.
[0018] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the terms “comprising” and “including” as used in embodiments of this application mean that the corresponding feature can be implemented as the presented feature, information, data, step, operation, element, and / or component, but do not exclude implementation as other features, information, data, step, operation, element, component, and / or combinations thereof supported by the art. It should be understood that when we say that an element is “connected” or “coupled” to another element, the one element can be directly connected or coupled to the other element, or it can mean that the one element and the other element establish a connection relationship through an intermediate element. Furthermore, “connected” or “coupled” as used herein can include wireless connection or wireless coupling. The term “and / or” as used herein indicates at least one of the items defined by the term; for example, “A and / or B” indicates implementation as “A,” or implementation as “A,” or implementation as “A and B.”
[0019] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0020] The technical solutions of this application and their effects are described below through several exemplary embodiments. It should be noted that the following embodiments can be referenced, borrowed from, or combined with each other. Identical terms, similar features, and similar implementation steps in different embodiments will not be repeated.
[0021] It should be noted that the collection, storage, use, processing, transmission, provision, and disclosure of data or user information involved in the technical solution of this application all comply with relevant laws and regulations and do not violate public order and good morals. The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, and displayed data) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use, and processing of related data must comply with relevant laws, regulations, and standards, and corresponding operation entry points are provided for users to choose to authorize or refuse. In addition, certain software, components, models, and other existing industry solutions may be mentioned in the embodiments of this application. These should be considered exemplary, and their purpose is merely to illustrate the feasibility of implementing the technical solution of this application, but does not mean that the applicant has already used or necessarily used such solutions.
[0022] The broadcast low-latency handover method and electronic device provided in this application are intended to solve at least one technical problem existing in the prior art.
[0023] Optionally, the low-latency handover method of this application can be used at the receiving end of a broadcast system, whereby the receiving end is connected to the transmitting end of the broadcast system. The transmitting end can be a broadcast host or core switch in the broadcast system, or other terminals capable of controlling the transmission of audio signals. The receiving end can be an audio system, decoder, network amplifier, or other device used to receive and process the audio signals transmitted by the transmitting end.
[0024] Options, such as Figure 1 As shown, the broadcast low-latency handover method of this application includes: S101: Receives audio signals.
[0025] Optionally, the audio signal is transmitted synchronously by the transmitting end through the main IP network channel and the analog backup channel, and the synchronously transmitted audio signals have the same global timestamp. The audio signal transmitted through the analog backup channel can be represented as an analog signal.
[0026] Optionally, when transmitting audio signals, an improved timestamp calibration algorithm based on PTP (Precise Time Protocol) can be used to synchronize the audio signals of the main IP network channel and the analog backup channel.
[0027] Optionally, the setting of the global timestamp includes: synchronizing the time of the sending end and the receiving end according to a preset time protocol and audio sampling frequency; determining that the synchronization is complete, generating a global timestamp based on the current system time, and embedding the global timestamp into the audio signal to be synchronized and transmitted to the main channel and analog backup channel of the IP network so that the receiving end can use the global timestamp for phase calibration and dynamic compensation.
[0028] Optionally, the preset time protocol can be PTP (Precise Time Protocol), Network Time Protocol (NTP), or other types of protocols. Before sending the audio signal, the transmitting end can synchronize its time with the receiving end according to the preset time protocol. During time synchronization, the receiving end's clock can be adjusted to match that of the transmitting end.
[0029] In one embodiment, time synchronization can be achieved by: Master election: designate the broadcast host (or core switch) as the PTP master clock source.
[0030] Slave-follower: All receivers (such as decoders and network amplifiers) act as slave clocks, receiving synchronization messages from the master clock and performing time synchronization based on these messages.
[0031] During time synchronization, the PTP protocol can be used to calculate the round-trip delay, eliminating the uncertainty caused by network transmission. Based on this round-trip delay, the time of the external clock can be adjusted to be the same as the master clock.
[0032] Optionally, during time synchronization, the clock frequency of the master clock can be locked to the audio sampling rate (e.g., 48kHz) to ensure that the local playback clock is from the same source as the master clock and to avoid "clock drift".
[0033] Optionally, before packaging the audio signal (RTP / UDP), the sending end can read the current PTP system time and generate a Global Stamp Time (GST). For the main IP network channel: the GST is embedded in the header or a specific field of the IP packet. The algorithm also controls the analog signal generator (or analog audio interface) to embed the same GST information in the frame header of the analog signal (transmitted through an analog backup channel). Regardless of whether it uses the main IP network channel or the analog backup channel, the "identity card" (timestamp) of each audio frame is completely consistent.
[0034] Optionally, when the receiver is a decoder, the buffer size and playback speed can be adjusted. Specifically, the decoder dynamically adjusts the buffer size based on network conditions. After receiving the data packet corresponding to the audio signal, the decoder reads the GST in the data packet and compares it with the local PTP clock. If network jitter is detected causing the data packet arrival time to fluctuate, the playback sampling rate of the DAC (digital-to-analog converter) will be fine-tuned. If the data packet arrives late (lagging), the player is instructed to speed up the playback speed (discarding a small number of interpolated samples). If the data packet arrives early (leading), the player is instructed to slow down the playback speed (inserting silence or repeating samples). Since analog signal transmission has almost no delay, while the audio signal transmission of the IP network main channel requires encoding, network transmission, and decoding, there is a delay (usually 20ms-100ms). The output delay of the IP network main channel (the delay between audio signal reception and playback) can be forced to lock at a fixed value (e.g., 50ms), and the analog channel can also be artificially delayed by the same 50ms through the buffer. In this way, the dual channels achieve "zero phase difference" at the output, meaning that the two audio paths completely overlap, providing a physical basis for the subsequent smooth switching between Fade-in and Fade-out.
[0035] S102: Detect network status based on a preset period.
[0036] Optionally, the network status includes at least one of packet loss rate, network jitter, and heartbeat response time. The preset period can be 10ms, 20ms, or other time lengths. The receiving end can receive audio signals and detect the network status based on the preset period and the audio signals.
[0037] Optionally, network status may include conditions that affect audio signal transmission, such as whether the network is down, whether the device is powered off, and the bit error rate.
[0038] S103: If the network status is determined to meet the switching conditions, the audio signal being played will be switched to the audio signal transmitted through the analog backup channel based on the global timestamp and the cached audio signal.
[0039] Optionally, determining whether the network status meets the switching conditions includes: determining the level of satisfaction based on the network status, where the level includes any one of warning level, buffer level, switching level, and circuit breaker level; if the level of satisfaction is determined to be the switching level or the circuit breaker level, then the switching conditions are satisfied; if the level of satisfaction is determined to be the warning level or the buffer level, then the switching conditions are not satisfied, and a preset strategy is executed according to the level of satisfaction, where the preset strategy includes preloading and packet loss enhancement.
[0040] In one embodiment, different levels of triggering conditions and system response actions can be as shown in Table 1, where jitter in Table 1 can be network jitter, and latency can be network round-trip latency or buffer underload time determined based on heartbeat response time.
[0041] Table 1
[0042] Optionally, the switching conditions may also include a time-aware threshold, a packet loss rate threshold, and a back-off threshold. If the status value in the network status is greater than at least one of these thresholds, the switching conditions are determined to be met, and the system switches to the simulated backup channel or the main IP network channel based on the met switching conditions.
[0043] In one embodiment, the time-aware threshold can be 150ms, with 200ms being the critical point for human perception, and 200ms set as the hard handover threshold. Specifically, when the network round-trip time (RTT) or buffer underload time approaches 150ms, handover must be triggered immediately, reserving 50ms for the algorithm to perform handover processing, ensuring that the user hears a sound interruption time of less than 50ms (or even imperceptible).
[0044] The packet loss rate threshold can be 5%, which is the watershed for speech / audio intelligibility. In listening tests, signal-to-noise ratio and clarity are crucial. When the packet loss rate is <5%, the human ear can barely distinguish the difference using FEC and interpolation algorithms. Once it exceeds 5%, the audio will exhibit noticeable "mechanical" or "hollow" sounds, requiring a decisive switch, as clear analog audio quality is superior to flawed digital audio quality. The switchback threshold can be designed based on hysteresis intervals; the switchback threshold (switching back from the analog backup channel to the IP network main channel) should be stricter than the switching threshold. Switchback conditions can be: a packet loss rate <1% and jitter <20ms for the IP network main channel for 10 consecutive seconds. This prevents the system from repeatedly switching between IP and analog when the network is in a "critical state" (i.e., the "ping-pong effect"), which would severely disrupt the testing environment. Automatic switchback should only be allowed after the network has completely stabilized.
[0045] Optionally, switching the played audio signal to the audio signal transmitted via the simulated backup channel based on a global timestamp and a buffered audio signal includes: determining whether to trigger a switching operation, playing the audio signal in the buffer, exponentially attenuating the audio signal based on the global timestamp, and exponentially amplifying the audio signal of the simulated backup channel. The audio signal buffered in the buffer is the audio signal transmitted via the main IP network channel in the most recent time period. Specifically, after triggering the switching operation, the audio signal stored in the buffer can be identified as the audio signal of the main IP network channel, and the audio signals are played sequentially according to their timestamps. During playback, the played audio signal is switched to the audio signal transmitted via the simulated backup channel based on the global timestamp.
[0046] Optionally, in digital signal processing, in order to conform to the logarithmic perception characteristics of human ear for volume, exponential curves or S-shaped curves are used instead of linear interpolation when performing exponential decay and exponential enhancement.
[0047] In one embodiment, an exponential curve is used to process the audio signal, and the algorithm includes two main functional forms: Fade-out function (fade out of the main IP network channel): Aout(t) = ek·t Fade-in function (alternate channel fade-in): Ain(t) = 1 - ek·(1-t).
[0048] In the formula, e can be a natural constant, t is the current time progress (i.e., the time progress of decay and enhancement), t∈[0,1] (0 is the start, 1 is the end), k is the decay / enhancement constant, which can be adjusted according to the actual listening experience, and the exponential slope can be around 6.9 (corresponding to a dynamic range of about -60dB), Aout(t) is the gain coefficient (0.0 to 1.0) corresponding to the audio signal of the main channel of the IP network at time t, and Ain(t) is the gain coefficient corresponding to the audio signal of the analog backup channel at time t.
[0049] Alternatively, the gain coefficients of the two channels can be determined using an equal-power crossfade model, which can be expressed as: Fade-out: Aout(t)=cos(π / 2·t) Fade-in:Ain(t)= sin(π / 2·t) The model satisfies Aout(t). 2 +Ain(t) 2 =1, ensuring that the total power of the mixed audio signal remains constant during playback.
[0050] In one embodiment, the comparison results of exponential curves and equal power crosses are shown in Table 2. The exponential curve or equal power cross can be selected for audio signal processing according to actual needs.
[0051] Table 2
[0052] Optionally, at the instant the switching is triggered (T0), a transition duration ΔT (200ms-500ms) can be determined. During this transition duration, the audio stream corresponding to the main IP network channel continues to play audio (either the audio signal in the buffer or the audio signal transmitted directly through the main IP network channel can be used), and Aout(t) is used as the corresponding gain coefficient. During this transition duration, the audio stream from the analog backup channel is read; this audio stream is initially muted, and its subsequent gain coefficient is Ain(t). During the switching process, the gain coefficient is calculated in real time based on Aout(t) and Ain(t). This exponential decay and gain method conforms to the logarithmic perception of volume by the human ear, avoiding the "middle collapse" phenomenon of linear transitions and achieving a smooth transition. Furthermore, during the switching process, background noise is controlled based on a noise threshold, and phase alignment is performed according to the global timestamp to avoid sound quality degradation caused by phase cancellation.
[0053] Optionally, the noise threshold value can be between 0.1 and 0.3, and the specific value can be adjusted according to the noise floor level.
[0054] Optionally, when the warning level is determined to be P1 based on the network status, a preloading operation can be performed, wherein the preloading includes: establishing a connection with the simulated backup channel, determining the size of the target buffer based on the network status; performing a smooth adjustment based on the target buffer size and the current buffer size, determining the final size of the buffer based on the boundary check structure of the adjustment result; and using the buffer to cache the audio signal transmitted by the main channel of the IP network.
[0055] In one embodiment, the following network quality metrics can be collected in real time: Network jitter: Current jitter value J(t); Packet Loss Rate: The current packet loss rate P(t); Network Delay: Current round-trip time D(t); Buffer status: Current buffer utilization U(t); The buffer size is calculated based on this network quality metric. A weighted dynamic calculation model can be used to calculate the target buffer size (Btarget(t)), and the formula is as follows: Btarget(t)=Bbase×[1+α×J(t) / Jmax+β×P(t) / Pmax+γ×D(t) / Dmax] Parameter description: Bbase: Base buffer size (can be set to 200ms by default); α, β, γ: Weighting coefficients (specific values can be α=0.6, β=0.3, γ=0.1); Jmax: Maximum allowable jitter (e.g., 50ms); Pmax: Maximum tolerable packet loss rate (e.g., 5%) Dmax: Maximum network latency (e.g., 300ms).
[0056] Optionally, after determining the target buffer size, to avoid audio problems caused by sudden changes in the buffer size, an exponential smoothing algorithm is used for smoothing adjustment. The formula for smoothing adjustment is: Bsmooth(t)=(1-λ)×Bsmooth(t-1)+λ×Btarget(t) In the formula, λ is the smoothing factor (which can take values from 0.1 to 0.3), and Bsmooth(t-1) is the size of the smoothing buffer at the previous time step.
[0057] In one embodiment, the base buffer size is 200ms, the current buffer size is equal to the base buffer size, and the weighting coefficients are: α=0.6; β=0.3; γ=0.1, the smoothing factor λ=0.2, the maximum allowable jitter is 50ms, the maximum tolerable packet loss rate is 0.05, and the maximum network latency is 300ms. During audio signal transmission, the network status is collected every 100ms, and the target buffer (Btarget(t)) size is calculated based on the collected network status. To avoid the target buffer being too large, the target buffer size can be calculated using the following formula: Btarget(t) = Bbase (1+α min(J(t) / Jmax,1.0)+β min(P(t) / Pmax,1.0)+γ min(D(t) / Dmax,1.0)) Based on Btarget(t), a smoothing adjustment is performed to obtain the smoothed buffer size Bsmooth(t). To prevent the smoothed buffer from increasing or decreasing excessively, boundary checks and buffer adjustments can be performed to obtain the final buffer size. A reasonable range can be set: a minimum buffer size of B_min = 100ms to prevent excessive shrinkage, and a maximum buffer size of B_max = 500ms to prevent excessive increase. The final buffer size B_final = max(B_min, min(B_max, round(Bsmooth(t)))). The difference between the smoothed buffer size and the current buffer size can be checked. If the difference is greater than 10ms (or other values), a boundary check is performed to obtain the final buffer size, and the buffer size is adjusted based on this final size.
[0058] Optionally, when dynamically adjusting the buffer size, the volume can also be adjusted synchronously to maintain a consistent listening experience. The volume adjustment formula can be:
[0059] In the formula, The adjusted volume. Based on the basic volume, This is the current buffer size. When adjusting the buffer size and volume, corresponding log entries are generated for later querying and optimization.
[0060] Optionally, when the warning level is P2, or when the packet loss rate is determined to be within a preset range, or when continuous packet loss occurs, packet loss mitigation enhancement processing is performed. This packet loss mitigation enhancement includes: obtaining packet loss information based on network status; determining a redundant packet ratio based on the packet loss information; generating redundant packets based on the redundant packet ratio; and generating an audio signal to be transmitted based on the redundant packets. During packet loss mitigation enhancement, the receiving end can also activate an interpolation algorithm to predict the lost audio signal using audio from preceding and following frames, and play the predicted audio signal to reduce the impact of packet loss.
[0061] Optionally, packet loss information can include packet loss rate, network jitter, whether packet loss is continuous, and the number of consecutive packet losses, etc. Redundant packets can be generated using forward error correction algorithms. The receiving end can maintain a sliding window to calculate the packet loss rate (PLR) and network jitter over a recent period (e.g., every 500ms or every 100 packets). In RTP / RTCP protocols, the receiving end sends packet loss statistics via RTCPXR messages; in WebRTC or QUIC protocols, the loss situation is calculated using the receiving end's ACK or NACK list. Besides packet loss rate, RTT (Round Trip Time) also needs to be monitored. If the RTT is very high, it indicates network congestion; excessively increasing FEC in this case may exacerbate congestion and requires careful handling.
[0062] Optionally, a mapping table can be established for packet loss resilience enhancement. For example: Packet loss rate <1%: FEC (Forward Error Correction) or 10% redundant packet setting is not enabled (i.e., 10% redundant packets are set in the audio signal to be transmitted).
[0063] Packet loss rate 1%-5%: Enable 25% redundant packets.
[0064] Packet loss rate > 5%: Enable 40%+ more redundant packets.
[0065] Advanced strategies (adaptive windows) can also be used to enhance packet loss resilience. For example, the window expansion technique mentioned on CSDN adjusts based on the characteristics of burst packet loss. If continuous packet loss (burst) is detected, not only should the number of redundant packets be increased, but the FEC encoding window should also be expanded (i.e., more original packets are packaged together for checksum calculation) to cover longer burst errors. To prevent the redundancy rate from drastically fluctuating between "high" and "low" (oscillating), a smoothing factor or hysteresis threshold is usually introduced. For example, the protection level is only increased when the packet loss rate is higher than the threshold for three consecutive cycles.
[0066] In one embodiment, when performing packet loss mitigation, the original audio signal data packets (SourcePackets) can be grouped. For example, each group can consist of k packets. Redundant packets for the audio signal can be generated using the Reed-Solomon (RS) or XOR algorithm. This algorithm can generate n redundant packets based on the calculated redundancy rate (redundancy packet generation logic: if the goal is to mitigate 20% packet loss and each group sends 10 packets, then at least 2-3 redundant packets need to be generated). The original data packets and redundant packets are combined with an FEC header (containing sequence number, block size, etc.) and sent to the receiving end via UDP. After receiving the audio signal data packets, the receiving end first places them in a buffer. If the original data packets are complete, they are sent directly to the decoder. If there is packet loss, matrix operations (Gaussian elimination, etc.) are performed using the received original packets + redundant packets to recover the lost data. As long as the total number of received packets (original + redundant) is greater than or equal to the number of original data packets, complete recovery is usually achievable. If the waiting time for redundant packets exceeds the playback time limit, recovery must be abandoned, and the packets must be discarded or packet loss hiding (PLC) must be implemented to avoid increasing latency.
[0067] Optionally, the method of this application includes: detecting the type of the currently transmitted audio signal, including exam audio, classroom amplification audio, and daily broadcast audio; if it is determined that the audio signal includes exam audio, then muting audio signals other than exam audio, locking the bandwidth of the audio channel corresponding to the exam audio and creating a buffer corresponding to the exam audio, and also preventing any audio source other than exam audio from preempting the audio channel and bandwidth.
[0068] Optionally, a preemptive priority queue can be established, in which P0 (highest) is the listening test signal, P1 is the classroom amplification signal, and P2 is the daily broadcast signal. When the P0 signal (exam audio) enters, the P1 / P2 signals are automatically muted, and the current channel bandwidth is locked to prevent any non-exam audio sources from being inserted.
[0069] In one embodiment, when a P0 signal is detected, digital signature verification and protocol parsing can be used to confirm whether the input signal is a legitimate examination signal. This examination signal is then synchronized with the examination management system to ensure the accuracy of the global timestamp. Furthermore, upon receiving the examination signal, the digital certificate and authorization information of the examination institution (which can be provided by the examination management system) can be verified. After successful verification, a dedicated audio processing channel and network bandwidth can be reserved for the P0 signal. The buffer is initialized, an independent circular buffer is created with a capacity of 300ms (for safety margin), and the system priority of the processing thread corresponding to the P0 signal is elevated to the highest level to achieve priority escalation.
[0070] Optionally, during mute processing, the P1 signal (such as background music / notification) can be muted immediately, and the P2 signal (such as audio signals for advertisements / other services) can be muted after a 50ms delay. For mixing logic switching, all non-exam-related mixing channels can be closed, the exam-specific mixer can be activated, and only the P0 signal can be allowed to pass.
[0071] Optionally, upon receiving an exam signal, the currently allocated network bandwidth (no less than 128kbps) can be locked, and any bandwidth adjustment operations can be prohibited. Furthermore, core audio processing resources can be locked, preventing other non-exam applications from accessing audio devices, and a "exam mode activated" status can be broadcast to all relevant subsystems so that the subsystems can switch to the operating state corresponding to the exam signal.
[0072] Optionally, when the P0 signal needs to exit, it can be verified whether the current exit conditions of the P0 signal are met (such as whether audio playback has ended), the current system resource usage status can be recorded, a resource status snapshot can be generated, and a security check can be performed to ensure that the exit operation will not affect the ongoing exam. Upon exit, a Fade-out process (lasting 100ms) can be performed on the P0 signal to ensure a smooth transition in audio output, gradually release locked network bandwidth, remove audio device access restrictions, and restore the mixer to its default mixing configuration. After the P0 signal has indeed exited, the P1 signal can be unmuted after a 150ms delay, the P2 signal can be unmuted after a 200ms delay, the system thread priority can be restored to normal levels, and a "Exam Mode Ended" status can be broadcast.
[0073] In one embodiment, the P0 signal exit condition may include whether the end time set by the examination management system has been reached or whether an "exam end" instruction has been received from the examination center. If the buffer time (e.g., 10 minutes) after the end time has been reached, the exit condition is deemed met. Alternatively, the exit condition may be determined to be met after receiving an "emergency end" instruction from the examination administrator and the examination administrator's permissions have been verified. The exit condition may also be determined to be met when the signal-abnormal terminal or broadcast system malfunctions. Signal abnormality interruption includes any one of the following: continuous P0 signal loss for more than 30 seconds, continuous signal quality degradation (packet loss rate > 8% for 10 seconds), or more than 3 failed digital signature verifications. System malfunctions include any one of the following: a serious error in the audio processing core, a complete network connection interruption for more than 60 seconds, or system resource utilization exceeding 95%.
[0074] After the exit conditions are met, verification can be performed. The exit conditions can be monitored in real time, and the detected conditions can be confirmed with a 5-second delay. The legality of the exit operation can be verified and the impact of the exit on the system can be assessed. Based on the assessment results, a decision can be made on whether to execute the exit.
[0075] Optionally, when exiting via the P0 signal fails, it can be checked whether the failure was due to resource release. If so, a forced cleanup process is executed, and detailed error logs are recorded for subsequent analysis and activation of the backup recovery mechanism. Furthermore, a state recovery timeout mechanism can be set (e.g., re-execute the exit operation after 30 seconds), a manual intervention interface is provided, and a complete state rollback mechanism is established.
[0076] Optionally, the P0 entry response time can be set to ≤200ms, the P0 exit response time to ≤300ms, and the state transition confirmation time to ≤50ms.
[0077] Alternatively, a multi-dimensional conflict detection algorithm can be used to monitor the following conflict types in real time: 1. Priority conflict: Signals of different priorities arrive simultaneously; 2. Bandwidth conflict: Total bandwidth demand exceeds available bandwidth; 3. Timing conflicts: Overlapping signal arrival times lead to competition for processing resources; If any of the above conflict types are detected, a priority arbitration algorithm or a dynamic bandwidth allocation algorithm can be used to resolve the conflict. Specifically, for the priority arbitration algorithm, a four-level priority arbitration mechanism can be used: P0 (highest priority): exam signal; P1 (high priority): emergency broadcast signal; P2 (medium priority): daily broadcast signal; P3 (low priority): background music signal. The arbitration rules can be: when a high-priority signal arrives, low-priority signals are automatically interrupted or muted; signals of the same priority are processed in the order of arrival time; in an emergency, upon receiving the P0 signal, all other signals can be forcibly interrupted. For bandwidth allocation, 30% dedicated bandwidth can be reserved for the P0 signal, or the bandwidth allocation of each signal can be dynamically adjusted according to real-time needs (such as increasing the bandwidth proportion of the P0 signal). When a high-priority signal ends, its occupied bandwidth is automatically reclaimed.
[0078] Optionally, the receiving end can also execute corresponding response actions according to the triggering conditions shown in Table 3.
[0079] Table 3
[0080] Specifically, for mild degradation, the audio encoding can be downgraded from AAC-LC to AMR-NB, and the audio signal sampling rate can be reduced, such as from 48kHz to 24kHz. Forward error correction (FEC) can be enabled, increasing redundant data by 10%, and the buffer size can be dynamically adjusted to 300ms. For moderate degradation, audio effects processing devices (such as reverb and equalizer) can be disabled, the priority of audio processing threads can be reduced, non-real-time data statistics functions can be paused, and low-power mode can be entered, turning off the receiver's LED indicators. For severe degradation, the system can automatically switch to the backup network card, reduce the audio quality to P2 level (such as setting the bitrate to 128kbps), disable all unnecessary background services, enter fail-safe mode, and retain only basic broadcast functions. For emergency degradation, the current state can be immediately saved to non-volatile memory, all network connections can be disabled, the system can switch to local audio playback mode, and enter standby mode, awaiting manual intervention.
[0081] Optionally, for the trigger conditions corresponding to mild and moderate degradation, machine learning algorithms can be used to dynamically adjust the degradation thresholds based on historical data. The new threshold is calculated as: new threshold = base threshold × (1 + historical failure rate × adjustment coefficient). The adjustment coefficient is set as follows: 0.1 in a stable environment, 0.3 in a general environment, and 0.5 in a severe environment. For the degradation recovery mechanism, the recovery condition is that system indicators return to normal for 5 consecutive minutes, with no new conflicts or failures, and manual confirmation or automatic detection passing. The recovery steps can be: progressively restoring disabled functional modules, dynamically increasing the audio quality level, re-establishing network connections, and restoring normal mode after completing system self-checks. All degradation operations are logged in detail, including: degradation timestamp, trigger condition details, specific actions performed, system status snapshot, and recovery time record.
[0082] Optionally, stress tests and regression tests can also be performed to check whether the downgrade operation and recovery can be performed normally. During stress testing, scenarios such as high packet loss rate network environment, high CPU load, and hardware failure can be simulated to test the execution of the downgrade operation. During regression testing, it can be verified that the downgrade does not affect core functions, ensure that the system state is correct after recovery, and test the switching between different downgrade levels.
[0083] Optionally, to achieve the localization of the broadcasting system, a domestically developed architecture can be used to achieve low-latency transmission of audio signals on the main channel of the IP network. This broadcasting system can employ an audio codec library optimized based on the RISC-V instruction set. RISC-V instruction set optimization can include compiler-level code generation optimization, deep tuning of RISC-V Vector Extensions (RVV), microarchitecture and hardware acceleration optimization, and assembly-level optimization of specific algorithms.
[0084] In one embodiment, compiler-level code generation optimization includes significantly improving code execution efficiency and reducing size by optimizing the backend strategy of the LLVM or GCC compiler during the software compilation phase. Redundant data movement instructions (such as `vmv`) often appear in the code generated by LLVM, meaning register values are repeatedly moved without modification. The compiler backend's `MachineCopyPropagationPass` can be fixed to correctly trace the register's Def-Use chain, thereby directly eliminating redundant `vmv` instructions, reducing instruction issuance, and alleviating register pressure. The C language standard requires low-precision operations such as `int8` / `int16` to be promoted to 32 bits before calculation, resulting in frequent `vsetvli` instructions to switch element widths (SEW) in RISC-V Vector Extensions (RVV), incurring significant overhead. By leveraging the explicit definition of overflow behavior in the RVV1.0 instruction set, native low-precision vector instructions (such as vrem.vv and vsra.vv) can be directly generated at the compiler's IR layer, bypassing unnecessary precision enhancement and truncation processes. This approach avoids frequent vset configuration instructions and significantly improves the processing efficiency of low-precision data (such as audio sampling points).
[0085] Control flow optimization (JumpThreading) can also be performed on control flow code like CoreMark's, which contains numerous branch judgments. The optimization scheme employs JumpThreading technology to analyze the control flow graph (CFG) at compile time. When the variable value before a branch is determined, an unconditional jump is used instead of a conditional jump, skipping unnecessary judgment logic. This reduces the probability of branch prediction failure and improves the execution efficiency of the instruction pipeline. Code density optimization (customized compressed instruction sets) can also be performed, customizing RV32C compressed instruction sets for specific scenarios such as network packet forwarding or audio control flow. Evolutionary algorithms can be used to analyze hot basic blocks and selectively replace standard 32-bit instructions with 16-bit compressed instructions to achieve code density optimization. Dynamic instruction compression ratios can be increased to 90%, significantly reducing instruction cache miss rates (by 30%~70%) and reducing instruction fetch power consumption.
[0086] Optionally, for in-depth tuning of RISC-V Vector Extensions (RVV), a dynamic LMUL configuration strategy can be used. LMUL (Register Set Size) determines the number of vector registers merged. GCC often defaults to LMUL=1, which limits performance. To improve performance, the -mrvv-max-lmul=dynamic option can be enabled, allowing the compiler to dynamically adjust LMUL based on the amount of data. For vrgather.vv (the gather operation), a smaller LMUL is preferred to reduce computational complexity (cost is related to LMUL²); for regular memory operations, an LMUL that covers the maximum amount of data is chosen to reduce the number of instructions. In hot functions such as x264 encoding, dynamically adjusting LMUL can bring performance on par with the high optimization levels of LLVM / Clang.
[0087] Vector length management (VLA vs VLS) is also possible, allowing for code writing specific to hardware VLENs, suitable for extreme performance scenarios. Portable code can also be generated using the vscale type. In the core audio processing loop, VLS mode or dynamic dispatch based on vlenb should be used as much as possible to avoid frequent execution of vsetvli within the loop. Tail handling optimization (TailHandling) can also be performed, distinguishing between "tail-unperturbed" and "tail-independent" processing. Specifically, if the tail data of vector computation does not need to be retained (e.g., in pure computation processes), the "tail-independent" mode should be forced to avoid hardware reading of old register values and reduce pipeline stalls.
[0088] Optionally, for microarchitecture and hardware acceleration optimization, advanced instruction fusion can be performed, which can merge two independent RISC-V instructions (such as LOAD+ALU or LOAD+BR) into a single micro-operation issue at the microarchitecture level. This improves the instruction-level parallelism (ILP) of a single-issue processor without changing the ISA, effectively achieving a "dual-issue" effect without increasing hardware area. This approach is particularly suitable for embedded audio processing cores, effectively masking memory access latency.
[0089] Multi-operand acceleration instructions can also be used, while traditional RISC-V instructions follow a "2-input, 1-output" model. Multi-operand interfaces (such as "3-input, 1-output") can be introduced specifically for accelerating FIR / IIR filters or SHA encryption algorithms. In FPGA verification, only 3% of the logic resource overhead is required to improve the performance of complex algorithms by 14%.
[0090] Floating-point unit (FPU) pipeline optimization is also possible. The 4Booth-Wallace algorithm can be used to replace traditional multipliers, reducing the number of partial products. A binary tree parallel structure is used to replace serial detection, reducing critical path latency. This significantly improves the FPU operating frequency (measured improvement of approximately 39%), making it suitable for high-precision audio floating-point operations.
[0091] Optionally, for assembly-level optimization of specific algorithms, especially audio or security-related algorithms, handwritten assembly or intrinsic optimization is essential. An example of SM4 / SM3 algorithm optimization is as follows: Circular shift: The RISC-V basic instruction set does not have a circular left shift instruction; it requires the combination of three instructions: SLLI, SRLI, and OR.
[0092] Bit-slicing: By optimizing data matrix transformations using the quadrant swapping method and combining it with Monte Carlo tree search to optimize S-boxes, SM4 throughput can be increased by 7 times. Bit-slicing techniques should also be used in bit operations for audio encoding and decoding (such as Huffman decoding) to utilize RISC-V bit manipulation instructions.
[0093] Optionally, during audio signal transmission, the buffer size can be adjusted based on network jitter to minimize end-to-end latency while ensuring the continuity and stability of the media stream. This can be achieved by intelligently predicting the optimal buffer size through real-time monitoring of network jitter characteristics, thus dynamically optimizing the buffer capacity.
[0094] Optionally, network jitter can be calculated based on the global timestamp of the audio signal, and the optimal buffer size can be calculated based on the current buffer size. The formula for calculating network jitter can be:
[0095] The optimal buffer size can be calculated as follows:
[0096] In the formula, This represents the network jitter value at the current moment. The average value of the jitter. The standard deviation of jitter, The current buffer size, For the optimal buffer size, t represents the predicted arrival time of the data packet, and t represents the actual arrival time of the data packet. This is the predicted global timestamp (the theoretical global timestamp of the data packet currently arriving at the receiving end). The optimal buffer size at time t+1 This is a smoothing factor (which can be between 0.1 and 0.3).
[0097] In one embodiment, network jitter can be monitored in real time, and the average jitter value for the current window can be calculated. and standard deviation It detects sudden jitter events and performs trend analysis, which can include: using linear regression analysis to analyze the trend of the most recent 10 sampling points and calculating the jitter change rate. Based on monitoring results, the current network status (normal network status, mild jitter, severe jitter) is determined, and the buffer size is adjusted accordingly. Specifically, the criteria for determining a normal network status are: The adjustment method is as follows: The adjustment objective is to maintain the minimum necessary buffer.
[0098] The criteria for determining mild shaking are: The adjustment method is as follows: The adjustment objective is to prevent potential packet loss. The criteria for determining severe jitter are: The adjustment method is as follows: Adjust the goal: Ensure uninterrupted streaming.
[0099] Alternatively, the smoothing factor can be dynamically calculated based on network stability. The specific calculation formula can be:
[0100] In the formula, k is the gain coefficient corresponding to the audio signal of the main channel of the IP network. The threshold is the standard deviation.
[0101] Alternatively, the CUSUM (CumulativeSum) algorithm can be used to detect sudden jitter, and the detection formula can be:
[0102] when A sudden jitter response is triggered at a certain time, where h is the sudden jitter threshold. Allowable deviation.
[0103] Optionally, a playback delay can be set for the buffer size. For joint optimization, the calculation formula can be:
[0104] in To adjust the coefficient, The target delay is set. During the buffer adjustment process, upper and lower limits (5ms-200ms) can be set to prevent the optimized buffer size from exceeding these limits. A hysteresis mechanism can be introduced (such as setting the buffer adjustment to be performed after 500ms or only performing the buffer size adjustment twice within 30s) to prevent frequent adjustments. A manual configuration interface (configuring the buffer size) is provided for special scenarios.
[0105] Optionally, the method of this application includes: acquiring the ambient noise of the playback environment corresponding to the audio signal, calculating the gain coefficient based on the ambient noise, and determining the gain of the receiver based on the gain coefficient and the noise difference in different areas so that the receiver can play the audio signal using the gain coefficient.
[0106] Alternatively, ambient noise can be collected in real time using IoT microphones deployed inside and outside the examination room. The gain coefficient can be calculated using the following formula: G=Kp×(Starget-Scurrent)+f(Nenv) Where G is the gain coefficient, Kp is the scaling factor, Starget is the target sound pressure level (e.g., 65dB in the examination room), current is the current sound pressure level, Nenv is the ambient noise, and f(Nenv) is the ambient noise compensation function.
[0107] Optionally, the weights of the gain coefficients corresponding to different network states and different levels of ambient noise can be preset. After the gain coefficients are calculated based on the network state and ambient noise, the final gain coefficient of the audio signal is calculated based on the product of the weights and the gain coefficients.
[0108] Alternatively, the formula for calculating the environmental noise compensation function can be: f(Nenv)=Kn×(Nenv-Nbase)A×eB×t In the formula, Kn is the noise sensitivity coefficient, which can range from 0.1 to 0.5; Nbase is the reference ambient noise sound pressure level, which can range from 35 to 45 dB; A is the noise response index, which can range from 1.0 to 1.5; B is the time decay coefficient, which can range from 0.01 to 0.1 s⁻¹; t is the time variable, which can be started from the last time the ambient noise changed abruptly; and e can be a natural constant.
[0109] Optionally, different environmental noise compensation functions can be selected according to actual application scenarios. When Nenv ≤ Nbase + 5dB (quiet environment), the value of the compensation function is 0; when Nbase + 5dB < Nenv ≤ Nbase + 20dB (moderate noise), the value of the compensation function can be: Kn×(Nenv-Nbase-5)A; when Nenv > Nbase + 20dB (high noise saturation), the value of the compensation function can be Kn×15A×e-B×(Nenv-Nbase-20).
[0110] Optionally, to adapt to different examination room environments, the noise sensitivity coefficient can be dynamically adjusted, and the calculation formula of the noise sensitivity coefficient can be: Kn=Kn0×[1+f×(Scurrent-Starget)2] Wherein: Kn0: initial noise sensitivity coefficient, f: gain adjustment coefficient (0.01~0.1).
[0111] In one embodiment, for a standard examination room environment, the parameters can be set as: Kp=0.3, Kn=0.2, Nbase=40dB, A=1.2, B=0.05s-1, Starget=65dB.
[0112] Optionally, a sudden noise detection mechanism can also be introduced. In this mechanism, the noise change rate ΔN / Δt is calculated, and when ΔN / Δt > a threshold value, the fast compensation mode is activated. In the fast compensation mode, Kn is increased by 30% (other proportions are also possible).
[0113] Optionally, to prevent gain oscillation, a gain change rate limit can be set, and moving average filtering and dead zone control (no adjustment within ±1dB) can be used to avoid frequent gain changes. Gain adjustment can also be performed according to the difference between off-campus noise and on-campus noise. Specifically, if the off-campus noise is high, the gain is appropriately increased; if the on-campus noise is high, the gain is decreased to prevent howling.
[0114] Optionally, distance attenuation compensation can also be performed. For example, according to a preset examination room distance model, a gain of 3-5dB is automatically added for remote speakers. The model can automatically calculate and compensate for sound pressure level attenuation caused by propagation distance according to the spatial layout of the examination room and the distribution characteristics of speakers. Through preset examination room distance parameters, the model implements 3-5dB gain compensation for remote speakers to ensure the consistency of sound pressure level at various positions in the examination room. Specifically, the distance attenuation compensation model adopts a layered architecture design and includes the following core components: The input layer includes: Geometric parameters of the examination room: length (L), width (W), height (H); Speaker layout parameters: speaker position coordinates (Xi,Yi,Zi); Reference point location: Standard listening position coordinates (Xr, Yr, Zr); The processing layer includes: Distance calculation module: Calculates the Euclidean distance from each speaker to the reference point; Attenuation prediction module: Predicts sound pressure level attenuation based on distance; Compensation calculation module: Determines the required gain compensation value; The output layer outputs the gain compensation values for each speaker: Gi (i=1,2,...,n).
[0115] Specifically, for each speaker i, the distance calculation module can calculate its distance to the reference point, and the calculation formula can be: di=
[0116] In the formula, di is the distance between speaker i and the reference point, (Xi, Yi, Zi) is the actual position of speaker i, and (Xr, Yr, Zr) is the position of the reference point.
[0117] For the reference distance d0, it can be the distance to the nearest speaker to the reference point or a preset reference distance. The theoretical attenuation can be calculated based on the inverse square law of sound wave propagation. ΔLi = 20 × log10(di / d0) In the formula, ΔLi is the theoretical attenuation.
[0118] The formula for calculating the attenuation prediction by the attenuation prediction module can be: Lp(d)=Lp(d0)-20×log10(d / d0)-E×(d-d0) Where: Lp(d): sound pressure level (dB) at a distance d; Lp(d0): Sound pressure level (dB) at the reference distance d0; d: Propagation distance (m); d0: Reference distance (m); E: Air absorption coefficient (Np / m); The formula for calculating the gain compensation value by the compensation calculation module can be: Gi = Kc × [20 × log10(di / dref)] in: Gi: The compensation gain (dB) of the i-th speaker; Kc: Compensation coefficient (0.6-0.8); di: Distance (m) from the i-th speaker to the reference point; dref: Reference distance (m).
[0119] The low-latency switching method for broadcasting provided in this application includes: a receiving end in a broadcasting system, the receiving end being connected to a transmitting end in the broadcasting system; the method comprising: receiving an audio signal, the audio signal being synchronously transmitted by the transmitting end through an IP network main channel and an analog backup channel, the synchronously transmitted audio signals having the same global timestamp; detecting network status based on a preset period, the network status including at least one of packet loss rate, network jitter, and heartbeat packet response time; if it is determined that the network status meets the switching conditions, then switching the played audio signal to the audio signal transmitted through the analog backup channel based on the global timestamp and the buffered audio signal. This application embodiment can achieve seamless switching between the IP network main channel and the analog backup channel, effectively avoiding problems such as audio interruptions, asynchrony, and uncontrollable latency, effectively meeting the requirements for low-latency synchronization in examinations and similar events.
[0120] According to one aspect of the embodiments of this application, a broadcast low-latency handover method is provided, such as... Figure 2 As shown, this method is used in a broadcast system's transmitting end, which is connected to the receiving end of the broadcast system.
[0121] Optionally, the switching method includes: S201: Audio signals are synchronously sent to the receiving end through the main channel and the analog backup channel of the IP network. The synchronously sent audio signals have the same global timestamp. The receiving end receives the audio signals and detects the network status based on a preset period. If it is determined that the network status meets the switching conditions, the audio signal being played is switched to the audio signal transmitted by the analog backup channel based on the global timestamp and the buffered audio signal.
[0122] Optionally, the network status includes at least one of packet loss rate, network jitter, and heartbeat response time.
[0123] Optionally, the setting of the global timestamp includes: synchronizing the time of the sending end and the receiving end according to a preset time protocol and audio sampling frequency; determining that the synchronization is complete, generating a global timestamp based on the current system time, and embedding the global timestamp into the audio signal to be synchronized and transmitted to the main channel and analog backup channel of the IP network so that the receiving end can use the global timestamp for phase calibration and dynamic compensation.
[0124] Optionally, determining whether the network status meets the switching conditions includes: determining the level of satisfaction based on the network status, where the level includes any one of warning level, buffer level, switching level, and circuit breaker level; if the level of satisfaction is determined to be the switching level or the circuit breaker level, then the switching conditions are satisfied; if the level of satisfaction is determined to be the warning level or the buffer level, then the switching conditions are not satisfied, and a preset strategy is executed according to the level of satisfaction, where the preset strategy includes preloading and packet loss enhancement.
[0125] Optionally, the audio signal being played is switched to the audio signal transmitted through the analog backup channel based on the global timestamp and the buffered audio signal, including: determining to trigger the switching operation, playing the audio signal in the buffer, exponentially attenuating the audio signal based on the global timestamp and exponentially enhancing the audio signal of the analog backup channel, wherein the audio signal buffered in the buffer is the audio signal transmitted in the most recent time period of the IP network main channel.
[0126] Optionally, preloading includes: establishing a connection with the simulated backup channel; determining the size of the target buffer based on network status; performing a smooth adjustment based on the target buffer size and the current buffer size; determining the final size of the buffer based on the boundary check structure of the adjustment result; and using the buffer to cache audio signals transmitted through the main IP network channel.
[0127] Optionally, packet loss mitigation enhancement includes: obtaining packet loss information based on network conditions, determining a redundant packet ratio based on the packet loss information, generating redundant packets based on the redundant packet ratio, and generating an audio signal to be transmitted based on the redundant packets.
[0128] Optionally, the method includes: detecting the type of the currently transmitted audio signal, including exam audio, classroom amplification audio, and daily broadcast audio; if it is determined that the audio signal includes exam audio, then muting audio signals other than exam audio, locking the bandwidth of the audio channel corresponding to the exam audio, and creating a buffer corresponding to the exam audio.
[0129] Optionally, the method includes: acquiring ambient noise of the playback environment corresponding to the audio signal; calculating a gain coefficient based on the ambient noise; and determining the gain of the receiver based on the gain coefficient and the noise differences in different regions so that the receiver can play the audio signal using the gain.
[0130] In one alternative embodiment, an electronic device is provided, such as Figure 3 As shown, Figure 3 The illustrated electronic device 4000 includes a processor 4001 and a memory 4003. The processor 4001 and the memory 4003 are connected, for example, via a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, which can be used for data interaction between the electronic device and other electronic devices, such as sending and / or receiving data. It should be noted that in practical applications, the transceiver 4004 is not limited to one type, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of this application.
[0131] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), a FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 4001 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.
[0132] Bus 4002 may include a pathway for transmitting information between the aforementioned components. Bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. Bus 4002 can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the figure, but this does not indicate that there is only one bus or one type of bus.
[0133] The memory 4003 may be ROM (Read-Only Memory) or other types of static storage devices capable of storing static information and instructions, RAM (Random Access Memory) or other types of dynamic storage devices capable of storing information and instructions, or EEPROM (Electrically Erasable Programmable Read-Only Memory), CD-ROM (Compact Disc Read-Only Memory) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium capable of carrying or storing computer programs and capable of being read by a computer, without limitation herein.
[0134] The memory 4003 stores computer programs that execute embodiments of this application, and its execution is controlled by the processor 4001. The processor 4001 executes the computer programs stored in the memory 4003 to implement the steps shown in the foregoing method embodiments.
[0135] Among them, electronic devices can be any electronic product that can interact with an object, such as personal computers, tablets, smartphones, personal digital assistants (PDAs), game consoles, interactive network television (IPTV), smart wearable devices, etc.
[0136] The electronic device may also include network devices and / or object devices. The network devices include, but are not limited to, a single network server, a server group consisting of multiple network servers, or a cloud consisting of a large number of hosts or network servers for cloud computing.
[0137] The networks in which the electronic devices are located include, but are not limited to, the Internet, wide area networks, metropolitan area networks, local area networks, and virtual private networks (VPNs).
[0138] The terms "first," "second," "third," "fourth," "1," "2," etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in a sequence other than that shown in the illustrations or text descriptions.
[0139] It should be understood that although arrows indicate various operation steps in the flowcharts of this application's embodiments, the order in which these steps are implemented is not limited to the order indicated by the arrows. Unless explicitly stated herein, in some implementation scenarios of this application's embodiments, the implementation steps in each flowchart can be executed in other orders as required. Furthermore, some or all steps in each flowchart, based on the actual implementation scenario, may include multiple sub-steps or multiple stages. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage can also be executed at different times. In scenarios where execution times differ, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and this application's embodiments do not limit this.
[0140] The above description is only an optional implementation method for some implementation scenarios of this application. It should be noted that for those skilled in the art, other similar implementation methods based on the technical concept of this application without departing from the technical concept of this application also fall within the protection scope of the embodiments of this application.
Claims
1. A broadcast low-latency handover method, characterized in that, For a receiving end in a broadcast system, the receiving end being connected to a transmitting end in the broadcast system, the method includes: The audio signal is received, which is synchronously transmitted by the transmitting end through the main channel of the IP network and the analog backup channel, and the synchronously transmitted audio signals have the same global timestamp. The network status is detected based on a preset period, and the network status includes at least one of packet loss rate, network jitter, and heartbeat packet response time. If the network status is determined to meet the switching conditions, the played audio signal will be switched to the audio signal transmitted through the analog backup channel based on the global timestamp and the cached audio signal.
2. The broadcast low-latency handover method according to claim 1, characterized in that, The setting of the global timestamp includes: The time synchronization between the transmitting and receiving ends is performed according to a preset time protocol and audio sampling frequency. Once synchronization is confirmed, a global timestamp is generated based on the current system time. This global timestamp is then embedded into the audio signals to be synchronized and transmitted to the main IP network channel and the analog backup channel, enabling the receiving end to use the global timestamp for phase calibration and dynamic compensation.
3. The broadcast low-latency handover method according to claim 1, characterized in that, Determining that the network state meets the handover conditions includes: The required level is determined based on the network status, and the level includes any one of the following: early warning level, buffer level, switching level, and circuit breaker level. If the level is determined to be either the switching level or the circuit breaker level, then the switching conditions are met. If the level is determined to be either warning level or buffer level, then the switching conditions are not met, and a preset strategy is executed according to the level that is met. The preset strategy includes preloading and packet loss resistance enhancement.
4. The broadcast low-latency handover method according to claim 3, characterized in that, The step of switching the played audio signal to the audio signal transmitted through the analog backup channel based on the global timestamp and the cached audio signal includes: The switching operation is triggered, the audio signal in the buffer is played, the audio signal is exponentially attenuated based on the global timestamp, and the audio signal of the analog backup channel is exponentially enhanced. The audio signal buffered in the buffer is the audio signal transmitted in the most recent time period of the main channel of the IP network.
5. The broadcast low-latency handover method according to claim 3, characterized in that, The preloading includes: Establish a connection with the simulated backup channel, and determine the size of the target buffer based on the network status; The buffer size is smoothly adjusted based on the target buffer size and the current buffer size, and the final size of the buffer is determined based on the boundary check structure of the adjustment result. The buffer is used to cache the audio signals transmitted through the main channel of the IP network.
6. The broadcast low-latency handover method according to claim 3, characterized in that, The packet loss resistance enhancement includes: Obtain packet loss information based on network status, and determine the redundant packet ratio based on the packet loss information; Redundant packets are generated based on the stated redundancy ratio, and audio signals to be transmitted are generated based on the stated redundancy packets.
7. The broadcast low-latency handover method according to claim 1, characterized in that, The method includes: The type of the currently transmitted audio signal is detected, including exam audio, classroom amplification audio, and daily broadcast audio. If it is determined that the audio signal includes the exam audio, then the audio signals other than the exam audio are muted, the bandwidth of the audio channel corresponding to the exam audio is locked, and a buffer corresponding to the exam audio is created.
8. The broadcast low-latency handover method according to claim 1, characterized in that, The method includes: Collect the ambient noise of the playback environment corresponding to the audio signal, and calculate the gain coefficient based on the ambient noise. The gain of the receiver is determined based on the gain coefficient and the noise differences in different regions so that the receiver can play the audio signal using the gain coefficient.
9. A broadcast low-latency handover method, characterized in that, For a transmitter in a broadcast system, the transmitter being connected to a receiver in the broadcast system, the method comprising: Audio signals are synchronously transmitted to the receiving end through the main IP network channel and the simulated backup channel. The synchronously transmitted audio signals have the same global timestamp. The receiving end receives the audio signals and detects the network status based on a preset period. If the network status is determined to meet the switching conditions, the audio signal being played is switched to the audio signal transmitted by the simulated backup channel based on the global timestamp and the cached audio signal. The network status includes at least one of packet loss rate, network jitter, and heartbeat packet response time.
10. An electronic device, characterized in that, The method includes a processor and a memory, the processor being communicatively connected to the memory, the memory storing program data, and the program data being used to execute the method as described in any one of claims 1-9.