A broadcast audio channel selection anti-howling processing method and system

By using hybrid coefficients to synchronously update audio processor parameters in the audio system, the transient noise and howling problems during audio channel switching are solved, achieving smooth audio switching and high-fidelity output, which is suitable for ship broadcasting systems.

CN122369489APending Publication Date: 2026-07-10HANGZHOU HUAYAN DIGITAL ELECTRON

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HANGZHOU HUAYAN DIGITAL ELECTRON
Filing Date
2026-06-09
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

Existing technologies suffer from transient noise and howling problems during audio channel switching in audio systems due to inconsistent parameter updates, which are particularly severe in distributed systems under the influence of bus jitter and clock drift.

Method used

By using hybrid coefficients to synchronously update audio processor parameters during the crossfade transition period, combined with the time synchronization mechanism of the digital bus, the parameter gradation and audio amplitude gradation are ensured to be synchronized. Linear interpolation and threshold jump strategies are used to handle different types of parameters, and the parameter update triggering time is optimized.

Benefits of technology

It achieves a smooth transition during audio switching, eliminates transient noise, ensures anti-feedback effect for voice and high-fidelity music output, and solves the auditory noise problem caused by independent parameter switching in existing technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122369489A_ABST
    Figure CN122369489A_ABST
Patent Text Reader

Abstract

This invention relates to the field of audio processing technology, and in particular to a method and system for anti-feedback processing of broadcast audio channel selection. The method includes receiving audio data frames via a digital bus and parsing the audio type flag and audio data stream from the audio data frames; initiating a crossfade transition period with a preset transition duration; obtaining mixing coefficients according to a predetermined gradient curve during the crossfade transition period; and weighted mixing the processed speech stream with a delayed music stream based on the mixing coefficients to obtain the output audio stream; and synchronously updating the current anti-feedback parameter set according to the mixing coefficients during the crossfade transition period, so that the current anti-feedback parameter set gradually changes with the change of the mixing coefficients. This invention solves the transient noise problem caused by independent parameter switching in the prior art, and while ensuring the anti-feedback effect of speech amplification, it achieves high-fidelity direct pass-through of the music signal, making the switching process smoother and more seamless.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio processing technology, and in particular to a method and system for anti-feedback processing of broadcast audio channel selection. Background Technology

[0002] In audio systems such as those used on ships and for public address systems, two types of signals typically need to be processed simultaneously: one is voice broadcasting, which requires noise reduction and anti-feedback processing to ensure clear instructions and avoid feedback; the other is background music playback, which requires the signal to be output directly without any processing to maintain high-fidelity sound quality. To balance both, existing technologies have proposed a channel selection architecture: a switching device is placed before the power amplifier, allowing the voice signal to pass through a dedicated noise reduction and anti-feedback processor, while the music signal bypasses this processor and is output directly.

[0003] Existing solutions suffer from the following problems in practical applications: The noise reduction-anti-feedback processor contains multiple sets of configurable parameters (such as notch filter center frequency, noise reduction threshold, adaptive filter step size, etc.). The acoustic environment varies at different call points (such as the cockpit, cabin, and passenger cabin), requiring different parameter sets to be loaded for each call point. When the system switches from one call point to another, the processor parameters need to be updated accordingly. In existing technologies, the audio crossfade controls the audio amplitude according to the mixing coefficients, while the processor parameters are written immediately or with a delay after detecting a switching event. This approach causes a jump in the coefficients of the internal filter of the processor at the moment the parameters are written. This jump generates transient noise at the output, which occurs during the crossfade transition period. Since the audio amplitude may not have fully switched at this time, the noise is exposed in the output signal, ruining the auditory effect of a smooth switch. Furthermore, the filter needs a certain convergence time to reach a steady state after the parameter switch, during which the feedback suppression capability decreases, potentially triggering feedback briefly. Furthermore, in distributed systems that use digital buses (such as FDCAN buses) to transmit audio data, there is jitter in the bus transmission between each call point node and the central processing node, and the clocks of each node may drift, resulting in a deviation between the actual arrival time of the audio data frame and the theoretical expected time. This deviation will further exacerbate the asynchrony problem between parameter switching and crossfading, leading to the generation of transient noise. Summary of the Invention

[0004] The main objective of this invention is to provide a broadcast audio channel selection anti-feedback processing method and system, which aims to solve the technical problems mentioned in the background art.

[0005] This invention proposes a broadcast audio channel selection anti-feedback processing method, comprising: Audio data frames are received via a digital bus, and audio type flags and audio data streams are parsed from the audio data frames; the audio type flags are used to indicate whether the current audio data stream is a speech signal or a music signal. In response to the audio type flag indicating a speech signal, the audio data stream is treated as a speech data stream, and the system enters speech processing mode. Obtain the current anti-feedback parameter set and nominal processing delay of the audio processor, and process the voice data stream according to the current anti-feedback parameter set to obtain the processed voice stream; Acquire a music data stream, apply the nominal processing delay to the music data stream, and obtain a delayed music stream; A crossfade transition period with a preset transition duration is initiated. During the crossfade transition period, a mixing coefficient is obtained according to a predetermined gradient curve. The processed speech stream and the delayed music stream are then weighted and mixed according to the mixing coefficient to obtain the output audio stream. During the crossfading transition period, the current anti-whistling parameter set is updated synchronously according to the mixing coefficient, so that the current anti-whistling parameter set changes gradually with the change of the mixing coefficient; In response to the audio type flag indicating a music signal, the audio data stream is treated as a music data stream and enters music processing mode: the nominal processing delay is applied to the music data stream and it is directly used as the output audio stream, and the audio processor is placed in a low-power or bypass state.

[0006] Preferably, the step of synchronously updating the current anti-whistling parameter set according to the mixing coefficient includes: Obtain the set of linear parameters and the set of nonlinear parameters in the current anti-whistling parameter set; For each linear parameter in the set of linear parameters, a gradual update is performed according to a linear interpolation formula during the cross-fading transition period; For each nonlinear parameter in the set of nonlinear parameters, the value of the mixing coefficient is monitored during the crossfading transition period. When the mixing coefficient reaches or exceeds a preset threshold for the first time, the nonlinear parameter is switched to a pre-stored target nonlinear parameter value.

[0007] Preferably, the step of monitoring the value of the mixing coefficient during the cross-fading transition period, and switching the nonlinear parameter to a pre-stored target nonlinear parameter value when the mixing coefficient first reaches or exceeds a preset threshold includes: During the cross-fading transition period, the current mixing coefficient is calculated for each sampling period, and it is determined whether the current mixing coefficient is greater than or equal to the preset threshold and whether the current switch has not been performed. When the conditions are met, execute the following sub-steps: Read the pre-stored target nonlinear parameter value associated with the source of the voice data stream; Write the target nonlinear parameter value into the corresponding register of the audio processor; Record the switch completion flag and prevent repeated switches. Preferably, the step of synchronously updating the current anti-whistling parameter set according to the mixing coefficient further includes: The howling suppression contribution of each parameter in the current anti-howling parameter set is obtained according to the pre-constructed contribution ranking model; The parameters in the current anti-whistling parameter set whose whistling suppression contribution is greater than or equal to a preset contribution threshold are classified as the first group of core parameters, and the remaining parameters are classified as the second group of non-core parameters. During the crossfading transition period, only the first set of core parameters are updated; at the same time, the difference between the current value and the target value of the second set of non-core parameters is recorded. After the crossfading transition period ends, a tiered delayed update strategy is adopted based on the absolute value of the difference: if the absolute value of the difference is greater than the second preset threshold, it is updated immediately; if the absolute value of the difference is less than or equal to the second preset threshold, it is updated in the next system idle cycle.

[0008] Preferably, the step of obtaining the mixing coefficient according to a predetermined gradient curve during the cross-fading transition period includes: A cosine curve is selected as the gradient curve for the mixing coefficient, wherein the mixing coefficient and derivative of the cosine curve are zero at the starting point and one at the ending point and the derivative is zero. Based on the change in the slope of the tangent of the cosine curve, the cross-fading transition period is divided into an acceleration gradient region, a linear gradient region, and a deceleration gradient region. The number of sample points during the transition period is calculated based on the sampling rate, and the current mixing coefficient value is calculated at each sample point according to the discretized form of the cosine curve. Within the acceleration and deceleration gradient regions, the mixing coefficients are calculated at a first sampling frequency; within the linear gradient region, the mixing coefficients are calculated at a second sampling frequency, wherein the first sampling frequency is higher than the second sampling frequency.

[0009] Preferably, the step of obtaining the current anti-feedback parameter set and nominal processing latency of the audio processor includes: During the system initialization phase, a unit pulse test signal is input to the audio processor, and the time difference from input to output is measured to obtain the static delay; Use the static delay as the nominal processing delay, and configure the read pointer offset of the delay compensation cache; In real-time operation, the music data stream is written to the delay compensation buffer, and the delayed data is read according to the read pointer offset; When the nominal processing delay fluctuates dynamically, the read pointer offset is adjusted in real time using an adaptive delay estimation algorithm, and the adjustment is limited to a preset range.

[0010] Preferably, the present invention further includes the following steps: The identifier field, control field, and data field of the audio data frame are obtained according to the FDCAN protocol standard; Extract the audio type flag bit from the data field or the identifier field; The audio type of the current audio data frame is determined according to the logical value of the audio type flag bit, and the audio data of consecutive audio data frames are reassembled into the audio data stream according to the receiving timing. When the audio type flag is detected to switch from music signal to voice signal, the timestamp of the received audio data frame is recorded; The nominal processing delay is dynamically adjusted based on the deviation between the received timestamp and the local clock, and the received timestamp is used as the start time of the crossfade transition period. The parameter update delay deviation caused by bus transmission jitter is calculated based on the mapping relationship between the received timestamp and the start time of the crossfade transition period; when the absolute value of the parameter update delay deviation exceeds the first preset threshold, it is determined that there is abnormal jitter in the current bus. In response to the determination of the abnormal jitter, an offset correction factor is obtained based on the mixing coefficient and the parameter update delay deviation, and the triggering time of the step of synchronously updating the current anti-whistling parameter set based on the mixing coefficient is adjusted according to the correction factor.

[0011] The present invention also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of a broadcast audio channel selection anti-feedback processing method.

[0012] The present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of a broadcast audio channel selection anti-feedback processing method.

[0013] The beneficial effects of this invention are as follows: By simultaneously using the mixing coefficient during the crossfade transition period for audio weighted mixing and audio processor parameter updates, this invention achieves complete temporal synchronization between parameter gradual changes and audio amplitude gradual changes, thereby solving the transient noise problem caused by independent parameter switching in the prior art. Specifically, during the crossfade transition period, this invention uses the same mixing coefficient to perform linear interpolation or threshold jump updates on the audio processor parameters, ensuring that the parameter change process is coordinated with the natural changes in audio amplitude. Since parameter updates and audio mixing share the same control curve, the transient noise generated by parameter jumps is masked by the audio signal during the same period and cannot be perceived by the user. Furthermore, because the parameters change gradually rather than abruptly, the filters inside the audio processor do not need to reconverge, thus avoiding howling during the switching process.

[0014] Furthermore, this invention utilizes the time synchronization mechanism of a digital bus to compensate for arrival time fluctuations caused by bus transmission jitter and node clock drift by recording the received timestamps of audio data frames and dynamically adjusting the delay compensation amount. The received timestamps are used as the synchronization reference for initiating the crossfade transition period. When abnormal bus jitter is detected, this invention obtains a correction factor based on the mixing coefficients and parameter update delay deviation, adjusting the trigger time of parameter updates. This ensures sample-level synchronization of the mixing coefficients, parameter updates, and audio mixing even under non-ideal bus conditions. Therefore, this invention achieves high-fidelity direct transmission of music signals while maintaining the anti-feedback effect of voice amplification, and the switching process is smoother and more seamless. Attached Figure Description

[0015] Figure 1 This is a schematic diagram of a method flow according to an embodiment of this application.

[0016] Figure 2 This is a schematic diagram of the system structure according to an embodiment of this application.

[0017] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0018] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0019] like Figure 1 As shown, this application provides a broadcast audio channel selection anti-feedback processing method, in which a digital signal processor performs the following steps: S1, receives audio data frames via a digital bus, and parses the audio type flag and audio data stream from the audio data frames; the audio type flag is used to indicate whether the current audio data stream is a speech signal or a music signal; S2, in response to the audio type flag indicating a speech signal, the audio data stream is treated as a speech data stream and the speech processing mode is entered; S21, obtain the current anti-feedback parameter set and nominal processing delay of the audio processor, and process the voice data stream according to the current anti-feedback parameter set to obtain the processed voice stream; S22, acquire the music data stream, apply the nominal processing delay to the music data stream, and obtain a delayed music stream; S23, initiate a crossfade transition period with a preset transition duration, obtain a mixing coefficient according to a predetermined gradient curve during the crossfade transition period, and perform a weighted mixing of the processed speech stream and the delayed music stream according to the mixing coefficient to obtain an output audio stream; ; In the formula, Indicates the output audio stream. The mixing coefficient is 0 at the beginning of the cross-diffuse transition period and 1 at the end of the cross-diffuse transition period, and increases monotonically in between. This represents the processed speech stream. Indicates a delayed music stream; S24, during the cross-fading transition period, the current anti-whistling parameter set is updated synchronously according to the mixing coefficient, so that the current anti-whistling parameter set gradually changes with the change of the mixing coefficient; S3, in response to the audio type flag indicating a music signal, the audio data stream is treated as a music data stream and enters music processing mode: the nominal processing delay is applied to the music data stream and it is directly used as the output audio stream, and the audio processor is placed in a low-power or bypass state.

[0020] As described in steps S1-S3 above, this invention is applied to a shipboard broadcasting system, which includes a central processing node (i.e., a digital signal processor) and multiple distributed call point nodes. Each call point node is equipped with an independent microphone and analog-to-digital converter for acquiring voice signals. Music signals come from independent digital music sources (e.g., media players or Bluetooth receiver modules). All audio data is transmitted to the central processing node via a digital bus (e.g., CAN bus or Universal Serial Bus), which integrates an audio processor, a delay compensation buffer, and a crossfade module. The audio processor uses a single-microphone input chip with noise reduction and anti-feedback functions, and its internal components include adaptive filters and notch filters, with parameter sets configurable via an external interface (e.g., I2C).

[0021] The central processing node performs the method described in steps S1-S3 according to the following steps.

[0022] In step S1, the central processing node receives audio data frames via a digital bus and parses the audio type flag and audio data stream from the audio data frames. Each audio data frame contains an audio data block of a predetermined length (e.g., a PCM sample point) and an additional flag field. This flag is set by the transmitting node according to the audio source: it is set to a voice flag when the audio comes from a microphone at a calling point, and a music flag when the audio comes from a music source. The central processing node distinguishes whether the current audio data stream is a voice signal or a music signal based on this flag. It should be noted that at any given time, the audio data frame transmitted on the bus contains only one type of audio data; that is, voice and music will not appear in the same frame simultaneously, but they can be transmitted alternately in a time-division multiplexing manner.

[0023] In step S2, in response to the audio type flag indicating a speech signal, the central processing node treats the audio data stream as a speech data stream and enters speech processing mode. This mode includes the following sub-steps: In step S21, the central processing node acquires the current anti-feedback parameter set and nominal processing delay of the audio processor. The current anti-feedback parameter set is a set of configurable parameters used within the audio processor for noise reduction and feedback suppression, including but not limited to notch filter center frequency, notch filter depth, adaptive filter step size, and noise gate threshold. The central processing node loads the default parameter set from non-volatile memory during system initialization and updates it as needed during operation. The nominal processing delay is the fixed time difference between receiving input and generating output by the audio processor, which can be obtained by consulting a datasheet or pre-measuring. The central processing node processes the speech data stream according to the current anti-feedback parameter set, i.e., writing speech data sample by sample into the audio processor's input buffer and reading its output buffer to obtain the processed speech stream. This processing process suppresses environmental noise in the speech and reduces the risk of acoustic feedback (feedback).

[0024] In step S22, the central processing node acquires the music data stream and applies the nominal processing delay to the music data stream to obtain a delayed music stream. The music data stream originates from the following sources: when the system is in music mode, the central processing node caches the received music audio data in a circular buffer; when the system switches to voice mode, this buffer already contains music data from the most recent period. The central processing node reads music samples from this buffer corresponding to the time preceding the nominal processing delay, thereby obtaining a delayed music stream that is time-aligned with the voice stream. This is achieved by setting a circular buffer with a depth equal to the nominal processing delay multiplied by the sampling rate; the write pointer increments with each received music sample, while the read pointer always lags the write pointer by a fixed offset.

[0025] In step S23, the central processing node initiates a crossfade transition period with a preset transition duration. The preset transition duration is a predetermined fixed value, such as 10 milliseconds, long enough to make the amplitude change imperceptible to the human ear, yet short enough to avoid a noticeable switching sensation. During the crossfade transition period, the central processing node generates mixing coefficients according to a predetermined gradient curve. In this embodiment, the gradient curve is chosen as a linear curve, meaning the mixing coefficients start at 0 and linearly increase to 1 over time. At each sampling moment, the central processing node weightedly mixes the processed speech stream with the delayed music stream according to the mixing coefficients to obtain the output audio stream. At the beginning of the transition period, the output is entirely the delayed music stream; at the end of the transition period, the output is entirely the processed speech stream; the output at intermediate moments is a weighted sum of the two.

[0026] In step S24, during the crossfade transition period, the central processing node synchronously updates the current anti-feedback parameter set according to the mixing coefficient, causing the current anti-feedback parameter set to gradually change with the mixing coefficient. Specifically, for each configurable parameter in the parameter set, the central processing node calculates and updates it according to a linear interpolation formula. At each sampling point, the central processing node writes the calculated updated parameter value to the corresponding register of the audio processor via the I2C or SPI bus, causing the internal algorithm parameters of the audio processor to gradually change synchronously with the mixing coefficient. Since the parameter update and audio weighted mixing use the same mixing coefficient, they are perfectly aligned in time: when the mixing coefficient is small, the output is still dominated by music, and the parameters are close to the values ​​before the switch; as the mixing coefficient gradually increases and speech gradually becomes dominant, the parameters also synchronously approach the target value. This synergy ensures that the audio processor does not generate transient noise caused by parameter abrupt changes during parameter switching, while avoiding feedback caused by mismatch between parameters and audio amplitude.

[0027] In step S3, in response to the audio type flag indicating a music signal, the central processing node treats the audio data stream as a music data stream and enters music processing mode. In this mode, the central processing node applies the nominal processing delay to the music data stream and directly outputs it as the audio stream. Simultaneously, the central processing node places the audio processor in a low-power or bypass state, i.e., stops inputting voice data to it, or sets its internal processing algorithm to pass-through mode, to reduce unnecessary power consumption and computational resource usage. When the music data stream arrives continuously, the output audio stream is the original music signal after delay compensation, without any noise reduction or anti-feedback processing, thus ensuring high fidelity in music playback.

[0028] Through the above steps, this invention achieves automatic switching between voice mode and music mode. Furthermore, when voice mode is activated, by simultaneously using mixing coefficients for audio mixing and parameter updates, the parameter changes of the audio processor and the audio amplitude changes are synchronized, thereby eliminating transient noise caused by independent parameter switching and solving the auditory noise problem caused by independent parameter switching and crossfading in existing technologies. Based on the above description, those skilled in the art can apply this method to various broadcast systems requiring mixed voice and music output without any inventive effort.

[0029] In one embodiment of the present invention, the following steps are further included: S11, Obtain the identifier field, control field, and data field of the audio data frame according to the FDCAN protocol standard; S12, extract the audio type flag bit from a predetermined byte position of the data field, or extract the audio type flag bit from the extension bit of the identifier field; wherein, each audio data frame carries at least one PCM audio sampling point, and each audio data frame contains an independent audio type flag bit; S13, determine the audio type of the current audio data frame according to the logical value of the audio type flag bit, and reassemble the audio data of consecutive audio data frames into the audio data stream according to the receiving timing. S14, when the audio type flag is detected to switch from music signal to voice signal, the receiving timestamp of the audio data frame is recorded. The receiving timestamp is obtained based on the time synchronization mechanism of the FDCAN bus, which synchronizes the clocks of each network node through the time reference frame in the bus message. S15, dynamically adjust the nominal processing delay based on the deviation between the received timestamp and the local clock to compensate for the audio data arrival time fluctuation caused by bus transmission jitter or node clock drift, and use the received timestamp as the start time of the crossfade transition period. S16, calculate the parameter update delay deviation caused by bus transmission jitter based on the mapping relationship between the received timestamp and the start time of the crossfade transition period; when the absolute value of the parameter update delay deviation exceeds the first preset threshold, it is determined that there is abnormal jitter in the current bus. S17, in response to the determination of the abnormal jitter, obtain an offset correction factor based on the mixing coefficient and the parameter update delay deviation, and adjust the trigger time of the step of synchronously updating the current anti-whistling parameter set based on the mixing coefficient according to the correction factor.

[0030] In a shipboard broadcasting system, multiple call points (e.g., bridge, engine room, passenger cabin) transmit audio data frames to the central processing node via the FDCAN bus. Due to jitter in bus transmission (i.e., uncertainty in frame arrival time) and potential clock drift at each node, a discrepancy arises between the actual arrival time of the audio frames received by the central processing node and the theoretically expected arrival time. This discrepancy affects the accuracy of the crossfade transition start time, thereby compromising the synchronization accuracy between the mixing coefficients and audio processor parameter updates, potentially leading to transient noise or feedback suppression failure. To address this issue, this embodiment utilizes the FDCAN bus's time synchronization mechanism. By recording the reception timestamp, dynamically adjusting the delay compensation amount, detecting abnormal jitter, and correcting the parameter update trigger time, sample-level synchronization accuracy is maintained even under non-ideal bus transmission conditions.

[0031] Specifically, as described in steps S11-S17 above, the central processing node obtains the identifier field, control field, and data field of the received audio data frame according to the FDCAN protocol standard. This parsing process is a conventional technique, namely, reading the frame structure conforming to the ISO 11898-1 standard from the receive buffer of the FDCAN controller.

[0032] Subsequently, the central processing node extracts the audio type flag bit from a predetermined byte position in the data field, or from the extended bits of the identifier field. In this embodiment, the audio type flag bit is configured to occupy the least significant bit of the first byte of the data field, where a logical value of "1" represents a speech signal and a logical value of "0" represents a music signal. Each audio data frame carries at least one Pulse Code Modulation (PCM) audio sample point, and each audio data frame contains an independent audio type flag bit. The central processing node determines the audio type of the current frame based on the logical value of the flag bit and reassembles the audio data of consecutive frames into an audio data stream according to the reception sequence. For example, when a frame with a flag bit of "1" is received, its PCM data is written to the speech buffer; when a frame with a flag bit of "0" is received, it is written to the music buffer. The above reassembly process itself belongs to conventional streaming media reception processing. The contribution of this embodiment is to bind the flag bit and audio data in the same frame, so that the type information arrives synchronously with the data.

[0033] When the central processing node detects a switch in the audio type flag from music to speech (i.e., the flag in the previous frame was "0" and the flag in the current frame is "1"), it records the received timestamp of that frame. This received timestamp is obtained based on the time synchronization mechanism of the FDCAN bus. Specifically, the FDCAN bus supports synchronizing the clocks of each network node by periodically sending time reference frames. The central processing node, acting as the time master node, broadcasts a time reference frame every second, carrying the current global clock value. Each calling node adjusts its local clock upon receiving the frame, ensuring that the clock deviation across the entire network is maintained within ±1 microsecond. The central processing node maintains a local hardware timer synchronized with the global clock. When it receives an audio data frame that triggers the switch, it reads the count value of this timer as the received timestamp. The resolution of this timestamp is 1 microsecond.

[0034] The central processing node dynamically adjusts the nominal processing delay based on the deviation between the received timestamp and the local clock. The nominal processing delay is defined as the inherent processing delay of the audio processor, obtained by inputting a unit pulse test signal to the audio processor and measuring the input-output time difference (e.g., 15 milliseconds). The current value of the local clock is recorded as the current time, and the difference between the received timestamp and the current time reflects bus transmission jitter and the drift of the local clock relative to the global clock. This deviation is added to the nominal processing delay to obtain the adjusted delay compensation. This adjusted delay compensation is used for delay compensation applied to the music data stream in step S22. Simultaneously, the central processing node uses the received timestamp as the start time of the crossfade transition period. Because the local clock is synchronized with the global clock, this method ensures that audio frames from different call points trigger the crossfade transition period at the central processing node with a unified time base, thereby avoiding accumulated synchronization errors caused by clock drift.

[0035] To further address abnormal jitter in bus transmission, the central processing node calculates the parameter update delay deviation caused by bus transmission jitter based on the mapping relationship between the received timestamp and the start time of the crossfade transition period. Specifically, the start time is the received timestamp itself, and the parameter update delay deviation is defined as the difference between the received timestamp and the theoretically expected arrival time. The theoretically expected arrival time is recursively derived from the bus period (e.g., 1 millisecond) and the received time of the previous voice frame. When the absolute value of this deviation exceeds a first preset threshold, the central processing node determines that there is abnormal jitter in the current bus. The first preset threshold is obtained by measuring the bus transmission jitter under normal operating conditions multiple times during the system development phase, and taking 1.5 times the maximum value as the threshold value. For example, if the normal jitter range is within ±50 microseconds, the first preset threshold can be set to 75 microseconds. This threshold value is pre-stored in non-volatile memory for runtime retrieval.

[0036] In response to the determination of abnormal jitter, the central processing node updates the delay deviation based on the mixing coefficient and the aforementioned parameters to obtain the offset correction factor. The calculation formula is as follows: ; in, Represents the mixing coefficient. This indicates the parameter update delay deviation, calculated based on the difference between the received timestamp of the current frame and the theoretically expected arrival time. This indicates the degree of deviation between the measured arrival time and the theoretically expected time. This represents the first preset threshold, which is a pre-stored constant. The physical meaning of this formula is: when... Compared to When the value is larger, the denominator is larger, and the correction factor is larger. The smaller the value; at the same time, the correction factor and the mixing coefficient... The correction factor is proportional to the transition period, resulting in a smaller correction magnitude in the early stages and a correspondingly larger correction magnitude in the later stages. The correction factor ranges from 0 to 1. Subsequently, the central processing node adjusts the trigger time of "synchronously updating the current anti-feedback parameter set according to the mixing coefficient" in step S24 based on the correction factor. The original trigger time was the time point corresponding to the mixing coefficient after the start of the crossfade transition period. The adjusted trigger time is equal to the original trigger time multiplied by the correction factor. That is, when abnormal jitter is detected, the central processing node accelerates or decelerates the parameter update process by compressing the time axis, thereby realigning it with the actual arrival time of the audio data stream. At the same time, the central processing node feeds back the adjusted trigger time information to the calling node that sent the audio data frame via the FDCAN bus. The calling node adjusts the transmission timing of subsequent frames accordingly, forming a closed-loop control. During the above adjustment process, the gradient curve (such as a cosine curve) of the mixing coefficient itself remains unchanged; only the mapping position of the parameter update action on the curve is changed.

[0037] Through the above steps, this embodiment achieves precise synchronization between mixing coefficient generation, parameter updating, and audio mixing under conditions of bus transmission jitter and clock drift. Specifically, this embodiment coordinates the FDCAN data transmission and reception and time synchronization mechanism with the parameter update triggering time of the audio processing domain, thereby solving the switching noise problem caused by non-ideal bus characteristics in existing technologies. Those skilled in the art can implement this solution without creative effort based on the above description.

[0038] In one embodiment of the present invention, the step of synchronously updating the current anti-whistling parameter set according to the mixing coefficient during the cross-fading transition period includes: S241, the parameters in the current anti-whistling parameter set are pre-divided into a linear parameter set and a nonlinear parameter set; S242, For each linear parameter in the set of linear parameters, a gradual update is performed according to the linear interpolation formula during the cross-fading transition period; S243, for each nonlinear parameter in the set of nonlinear parameters, monitor the value of the mixing coefficient during the crossfading transition period, and when the mixing coefficient reaches or exceeds a preset threshold for the first time, switch the nonlinear parameter to a pre-stored target nonlinear parameter value.

[0039] As described in steps S241-S243 above, this embodiment specifically discloses a method for synchronously updating the current anti-feedback parameter set of an audio processor based on the mixing coefficients during the crossfade transition period. Audio processors have diverse parameter types, and different parameters exhibit different response characteristics to linear interpolation. If the same linear interpolation strategy is applied to all parameters, some parameters (such as the notch filter center frequency) may cause nonlinear distortion of the filter transfer function during the interpolation process. To solve this problem, this embodiment classifies parameters into two categories: linear parameters and nonlinear parameters, and employs two different update strategies—gradual interpolation and threshold jump—to ensure a smooth parameter transition while avoiding filter instability.

[0040] Specifically, the digital signal processor pre-classifies the parameters in the current anti-feedback parameter set of the audio processor into linear parameter sets and nonlinear parameter sets. This distinction is based on the following: linear parameters are those whose values ​​have a linear mapping relationship with the audio processing effect and can be directly interpolated linearly in the numerical domain without causing changes in the filter structure; nonlinear parameters are those whose values ​​have a nonlinear mapping relationship with the filter transfer function (e.g., mapped to Z-domain coefficients through bilinear transformation), and directly interpolating these parameters linearly can lead to unexpected frequency response distortion or even pole drift in the intermediate state of the transfer function. Specifically, in the audio processor used in this embodiment (e.g., the XMOS XVF3800 voice processor), its linear parameters include the noise reduction intensity threshold (range 0.0 to 1.0, dimensionless), the adaptive filter step size (range 0.001 to 0.1, dimensionless), and the total gain (range -20dB to +20dB, expressed as linear amplitude). Nonlinear parameters include the notch filter center frequency (unit: Hertz) and the filter quality factor (dimensionless). The above parameter classifications can be obtained in advance by consulting the audio processor's datasheet or through experimental measurements, and stored in non-volatile memory for reading at runtime.

[0041] In step S242, for each linear parameter in the linear parameter set, the digital signal processor performs a gradual update according to a linear interpolation formula during the crossfade transition period. The linear interpolation formula is: ; In the formula, This represents the parameters in the updated anti-whistle parameter set. This represents the parameters in the anti-whistle parameter set before the update. This represents the mixing coefficient (which monotonically increases from 0 to 1 during the transition period). The target linear parameter value is obtained by reading the corresponding parameter value associated with the current call point from a pre-stored parameter table based on the source of the current voice data stream (i.e., the currently active call point identifier, such as the cockpit, cabin, or passenger cabin). This parameter table is configured by the host computer and written to flash memory during system initialization. At each sampling point during the transition period, the digital signal processor calculates the updated anti-feedback parameter set based on the mixing coefficients at the current moment and writes the calculation results to the corresponding register of the audio processor via the I2C bus. Since the value range of the linear parameter is continuous and linear, this interpolation method ensures that the parameter changes smoothly with the mixing coefficients, thereby avoiding transient noise caused by abrupt parameter changes.

[0042] In step S243, for each nonlinear parameter in the set of nonlinear parameters, the digital signal processor monitors the value of the mixing coefficient during the crossfade transition period. When the mixing coefficient first reaches or exceeds a preset threshold, the nonlinear parameter is switched from its current value to a pre-stored target nonlinear parameter value. The preset threshold is a predetermined value, obtained by: during the system development phase, determining the masking threshold of the human ear to transient noise generated by frequency abrupt changes through psychoacoustic experiments. Experimental results show that when the mixing coefficient reaches 0.5, the amplitude of the speech signal is attenuated by 6 dB relative to the peak value, at which point the human ear's sensitivity to transient noise superimposed on the speech is significantly reduced. Therefore, the preset threshold is set to 0.5. This threshold is stored in non-volatile memory. During the transition period, at each sampling point, the digital signal processor compares the current mixing coefficient with the threshold 0.5 and checks whether the current switch has not yet been performed (recorded via a Boolean flag). Once the condition is met, the processor immediately reads the target nonlinear parameter value associated with the current call point from the pre-stored parameter table, writes it to the audio processor via the I2C bus, and sets the switching flag to true to prevent subsequent repeated switching. For example, when switching from music mode to driver's voice mode, the notch filter's center frequency needs to switch from the default 2000 Hz to the corresponding 3200 Hz on the driver's console. At the instant the mixing coefficient reaches 0.5, the processor writes 3200 Hz in one go, without going through an intermediate frequency value. Since the voice signal amplitude is already large enough at this point, the transient noise generated by this frequency jump is effectively masked by the voice itself and is imperceptible to the user. Simultaneously, because the nonlinear parameter only switches once at the threshold point, frequency trajectory deviation or filter instability caused by linear interpolation is avoided.

[0043] Through the above steps, this embodiment achieves the classification and adaptation update of audio processor parameters. The gradual interpolation of linear parameters ensures a smooth transition for sensitive parameters such as noise reduction intensity and filter step size, while nonlinear parameters utilize psychoacoustic masking effects to jump abruptly at appropriate times, avoiding nonlinear distortion of the filter transfer function and ensuring the inaudibility of switching noise. Conventional parameter read / write operations are well-known techniques. The contribution of this embodiment lies in combining parameter classification with a hybrid coefficient synchronization mechanism, selecting different update strategies based on parameter type, thereby resolving the contradiction between parameter switching noise and filter stability without increasing hardware costs. Those skilled in the art, based on the above description, can apply this method to different types of audio processors without any inventive effort.

[0044] In one embodiment of the present invention, the step of monitoring the value of the mixing coefficient during the crossfading transition period, and switching the nonlinear parameter to a pre-stored target nonlinear parameter value when the mixing coefficient first reaches or exceeds a preset threshold includes: S2431, the preset threshold is set to a mixing coefficient of 0.5; S2432, during the cross-fading transition period, calculate the current mixing coefficient for each sampling period, and determine whether the current mixing coefficient is greater than or equal to the preset threshold and whether the current switch has not been performed. S2433, when the condition is met, execute the following sub-step: Read the pre-stored target nonlinear parameter value, which is associated with the source of the current voice data stream; Write the target nonlinear parameter value into the corresponding register of the audio processor; Record the switch completion flag to prevent repeated switches.

[0045] As described in steps S2431-S2433 above, the selection of the preset threshold is based on a subjective auditory test experiment: During the system development phase, test subjects with normal hearing are selected, and a standard speech sample superimposed with simulated transient noise is played in an anechoic chamber environment. The amplitude attenuation of the speech signal is changed, and the attenuation at which the test subject can just not perceive the transient noise is recorded. Based on the masking effect principle in psychoacoustics, when the amplitude of the speech signal is attenuated by 6 dB (i.e., the amplitude is attenuated to 0.5 times the original value), the superimposed transient noise is effectively masked. Therefore, the preset threshold is set to 0.5. This preset threshold is pre-stored in non-volatile memory for retrieval during runtime.

[0046] In this embodiment, the digital signal processor reads a preset threshold from the memory, which is equal to 0.5. It should be noted that this preset threshold is a dimensionless value of the mixing coefficients and is independent of parameters such as the audio processor's sampling rate and bit width; therefore, it is applicable to different hardware platforms.

[0047] In step S2432, during the crossfade transition period, the digital signal processor calculates the current mixing coefficient in each sampling cycle and determines whether two conditions are met: first, the current mixing coefficient is greater than or equal to the preset threshold of 0.5; second, the current switch has not yet been executed. The "not yet executed" state is recorded by a Boolean flag, which is initialized to false each time the voice processing mode is entered. In each sampling cycle, the processor first obtains the value of the current mixing coefficient from the crossfade module, then compares it with 0.5, and checks the flag. Only when both conditions are met simultaneously does the process proceed to step S2433.

[0048] In step S2433, when the condition is met, the digital signal processor executes the following sub-steps. First, it reads the pre-stored target nonlinear parameter value, which is associated with the source of the current voice data stream. "Association" refers to establishing and utilizing a preset mapping relationship. During the initialization or configuration phase, the system pre-builds a parameter configuration database (or lookup table). This database stores the correspondence between multiple "voice data stream sources" and "target nonlinear parameter sets." The "source" can be physical, such as a specific location on a ship (e.g., bridge, engine room, cabin, restaurant) or a specific input interface (e.g., a microphone with a specific number, an alarm signal input port); it can also be logical, such as the type of signal (e.g., artificial voice, automatically synthesized voice, background music, emergency alarm). When the system receives a voice data stream, it first parses the source identifier of the data stream. Subsequently, the system queries the parameter configuration database and locks the corresponding target nonlinear parameter value based on the source identifier. For example, if the detected voice data stream originates from an "aircraft cabin microphone," the system will automatically associate a set of target parameters with strong nonlinear processing capabilities (such as a high gain attenuation coefficient and a stronger noise suppression threshold) due to the typically high background noise in the cabin environment. If the voice data stream originates from a "conference room microphone," the system will associate a set of smoother nonlinear parameters to preserve the naturalness and detail of the voice, given the relatively quiet environment. The specific determination of the associated values ​​is based on existing technology and will not be elaborated here. The parameters in the target nonlinear parameter set include the notch filter center frequency and the filter quality factor. This parameter table is configured by the host computer and written to flash memory during system initialization. Next, the digital signal processor writes the target nonlinear parameter values ​​to the corresponding registers of the audio processor via the I2C bus. The write operation uses a batch write mode, sending the addresses and data of multiple registers at once to reduce bus occupancy time. Then, the processor sets the switching completion flag to true, preventing repeated switching operations during the current crossfade transition period. This flag is reset after the current voice processing mode ends (i.e., when the mixing coefficient returns to 0), so that it can be re-enabled for the next voice switch.

[0049] Through the above steps, this embodiment achieves a one-time jump of the nonlinear parameter at the exact moment when the mixing coefficient is exactly equal to 0.5. The timing of this switch ensures that transient noise occurs when the speech signal amplitude attenuates to half, utilizing the masking effect of the speech signal itself to eliminate the audibility of the switching noise. Simultaneously, since the switch is executed only once, it avoids audio processor state confusion caused by repeated writing. Conventional I2C write operations and register accesses are well-known technologies. The contribution of this embodiment lies in binding the switching timing to a specific value of the mixing coefficient and ensuring single execution through a status flag, thereby achieving synergy between noise masking and parameter updating without relying on complex algorithms. Those skilled in the art, based on the above description, can apply this method to audio processors with parameter switching capabilities without any inventive effort.

[0050] In one embodiment of the present invention, the step of synchronously updating the current anti-whistling parameter set according to the mixing coefficient further includes: S244, Obtain the howling suppression contribution of each parameter in the current anti-howling parameter set according to the pre-constructed contribution ranking model; S245, the parameters in the current anti-whistling parameter set whose whistling suppression contribution is greater than or equal to a preset contribution threshold are divided into the first group of core parameters, and the remaining parameters are divided into the second group of non-core parameters; S246, during the crossfading transition period, only the first set of core parameters are updated; at the same time, the difference between the current value and the target value of the second set of non-core parameters is recorded; S247, after the crossfading transition period ends, a hierarchical delayed update strategy is adopted according to the absolute value of the difference: if the absolute value of the difference is greater than the second preset threshold, it is updated immediately; if the absolute value of the difference is less than or equal to the second preset threshold, it is updated in the next system idle cycle.

[0051] As described in steps S244-S247 above, in order to improve the stability of the howling suppression effect, this embodiment classifies the parameters according to their contribution to howling suppression, prioritizes updating the core parameters during the transition period, and updates the non-core parameters according to the difference after the transition period ends, so as to prevent missing the update opportunity of key parameters due to bus congestion, thereby optimizing the utilization efficiency of bus resources to ensure the stability of the howling suppression effect.

[0052] In step S244, the digital signal processor obtains the feedback suppression contribution of each parameter in the current feedback parameter set according to a pre-constructed contribution ranking model. The contribution ranking model is constructed as follows: During system development, the audio processor is connected to a test platform, a standard speech signal is input, and feedback conditions are simulated. Using a controlled variable method, the value of each parameter is changed one by one, and the change in feedback suppression effect is measured (e.g., using stable gain margin or feedback trigger time as evaluation indicators). The change in feedback suppression effect caused by the change of each parameter is normalized to the interval between 0 and 1, and this is taken as the contribution of that parameter. This model can be represented as a mapping table from parameters to contributions, pre-stored in non-volatile memory. For different models of audio processors, the contribution mapping table may be different; this embodiment does not limit specific values, only the acquisition method.

[0053] In step S245, the digital signal processor classifies parameters whose howling suppression contribution to the current anti-howling parameter set is greater than or equal to a preset threshold into a first group of core parameters, and the remaining parameters into a second group of non-core parameters. The preset threshold is a predetermined value, obtained by estimating the maximum number of parameters that can be updated within the transition period based on the system's designed bus bandwidth and transition period length. The contribution ranking position corresponding to this number is used as the threshold. For example, if the bus bandwidth is 400 kilobits per second, the transition period is 5 milliseconds, and each parameter update requires the transmission of 3 bytes (address plus data), then a maximum of approximately 80 parameters can be updated.

[0054] In step S246, during the crossfade transition period, the digital signal processor only updates the first set of core parameters based on the mixing coefficients (i.e., updates according to linear interpolation or threshold hopping strategies). Simultaneously, the processor records the difference between the current value and the target value for each parameter in the second set of non-core parameters. The current value is read from the audio processor's register, and the target value is obtained from a pre-stored parameter table associated with the current call point. The difference is calculated as follows: for numerical parameters, the difference equals the target value minus the current value; for indexed parameters (such as notch filter frequency index), the difference is defined as 0 (if the indices are the same) or 1 (if the indices are different), indicating whether an update is needed.

[0055] In step S247, after the crossfade transition period ends, the digital signal processor employs a graded delayed update strategy based on the absolute value of the difference. Specifically, for each parameter in the second group of non-core parameters, it is determined whether the absolute value of its difference is greater than a second preset threshold. The second preset threshold is obtained by determining it based on the human ear's perception threshold of auditory differences caused by parameter changes. For example, for equalizer gain, the minimum perceptible gain change is approximately 0.5 dB, so the second preset threshold is set to a corresponding linear amplitude ratio of 0.94 or 1.06. For noise threshold, the minimum perceptible change is approximately 1 dB, corresponding to a linear amplitude ratio of 0.89 or 1.12. These threshold values ​​can be obtained in advance through subjective hearing tests. If the absolute value of the difference is greater than the second preset threshold, it indicates that the current value of the parameter differs significantly from the target value. Failure to update in time may affect sound quality or feedback suppression. Therefore, the digital signal processor immediately performs an update operation, writing the parameter to the audio processor via the bus. If the absolute value of the difference is less than or equal to the second preset threshold, it indicates that the difference is small and imperceptible to the human ear. Therefore, the processor postpones the update operation to the next system idle cycle. The system idle cycle is defined as the idle time period after the digital signal processor completes all real-time audio processing tasks (including audio data reception, crossfade calculation, parameter interpolation, etc.) and waits for the next audio frame to arrive. The processor determines whether it is in an idle cycle by checking the task queue and timer status, and performs delayed update parameter writing operations in batches within the idle cycle.

[0056] Through the above steps, this embodiment achieves hierarchical parameter updates under limited bus bandwidth conditions. The construction of the contribution ranking model, the setting of preset thresholds, the recording of differences, and the hierarchical delayed updates based on these differences constitute features that distinguish it from conventional batch update methods. The contribution of this embodiment lies in combining parameter contribution with update timing, prioritizing real-time updates of core parameters while flexibly scheduling updates for non-core parameters based on the magnitude of their differences. This reduces peak bus load without affecting howling suppression. Those skilled in the art, based on the above description, can apply this method to audio processor systems with different parameter scales without any inventive effort.

[0057] In one embodiment of the present invention, the step of obtaining the mixing coefficient according to a predetermined gradient curve during the cross-fading transition period includes: S231, a cosine curve is selected as the gradient curve of the mixing coefficient, wherein the mixing coefficient of the cosine curve is zero and the derivative is zero at the starting point, and the mixing coefficient is one and the derivative is zero at the ending point. S232, based on the change in the slope of the tangent of the cosine curve, the cross-fading transition period is divided into an acceleration gradient region, a linear gradient region, and a deceleration gradient region; S233, calculate the number of sample points during the transition period based on the sampling rate, and calculate the current mixing coefficient value at each sample point according to the discretized form of the cosine curve; S234, within the acceleration and deceleration gradient regions, the mixing coefficient is calculated at a first sampling frequency; within the linear gradient region, the mixing coefficient is calculated at a second sampling frequency, wherein the first sampling frequency is higher than the second sampling frequency.

[0058] As described in steps S231-S234 above, this embodiment divides the transition period into different regions based on the change in the slope of the tangent of the cosine curve, and adopts a differentiated sampling frequency for the interpolation update of the linear parameters, thereby optimizing the processor's computational load while ensuring smooth parameter tracking.

[0059] Specifically, a cosine curve is chosen as the gradient curve for the mixing coefficient. The expression for this cosine curve satisfies the following in the time domain: the mixing coefficient is zero and its derivative is zero at the starting point, and the mixing coefficient is one and its derivative is zero at the ending point. Specifically, the starting point of the transition period is taken as the zero point in time, and the duration of the transition period is... Then the mixing coefficient Represented as: ; The first derivative (slope) of the curve is: ; slope at and The value is zero at the point of zero. The second derivative reflects the rate of change of the slope, reaching its maximum value at [a certain point]. and The absolute value in the vicinity is relatively large.

[0060] Digital signal processors (DSPs) divide the crossfade transition period into three regions: an acceleration gradient region, a linear gradient region, and a deceleration gradient region, based on the change in the slope of the tangent line of the cosine curve. The division is based on the relative magnitude of the absolute values ​​of the slopes. Specifically, the time point at which the absolute value of the slope reaches its maximum is calculated. The time interval during which the absolute value of the slope rises from zero to a predetermined proportion (e.g., 80%) of its maximum value is defined as the acceleration gradient region; the time interval during which the absolute value of the slope decreases from 80% of its maximum value to 80% and then back to 80% is defined as the linear gradient region; and the time interval during which the absolute value of the slope decreases from 80% to zero is defined as the deceleration gradient region. Due to the symmetry of the cosine curve, the acceleration and deceleration gradient regions are related to... The transition regions are symmetrical, each accounting for approximately 25% of the total transition period, with the linear gradient region accounting for approximately 50%. The aforementioned proportion of 80% is a preset empirical value, obtained by comparing the mean square error of the parameter trajectory with the theoretical curve under different division proportions through simulation experiments, and selecting the proportion with the smallest error.

[0061] The digital signal processor calculates the total number of sample points during the transition period based on the sampling rate. Let the sampling rate be... (Unit: Hertz), transition period duration is (Unit: seconds), then the total number of sample points At each sample point, the processor calculates the current mixing coefficient value according to the discretized form of the cosine curve described above, i.e.: ; Within the acceleration and deceleration transition regions, due to the large absolute values ​​of the first derivatives of the mixing coefficients, parameter interpolation is more sensitive to sampling errors. Therefore, the digital signal processor employs an encrypted interpolation strategy for the linear parameters. This encrypted interpolation strategy involves further increasing the parameter update frequency beyond the original once-per-sample-point update, ensuring that the number of parameter updates exceeds the number of sample points within the acceleration and deceleration transition regions. Specifically, within these regions, the parameter update interruption period is shortened to half the original period, meaning the parameters are updated once every half-sample period. Correspondingly, the mixing coefficient values ​​are obtained from adjacent sample points through linear interpolation. The value is obtained. The purpose of the encrypted interpolation strategy is to match the parameter update frequency with the rate of change of the slope of the mixing coefficients, thereby reducing parameter trajectory jitter caused by insufficient sampling. Within the linear gradient region, since the rate of change of the slope is small, the digital signal processor adopts a conventional interpolation strategy, that is, updating the parameters once for each sample point. The conventional interpolation strategy is consistent with the linear interpolation update method described in step S242.

[0062] The digital signal processor uses the calculated mixing coefficient values ​​simultaneously for both the weighted mixing in step S23 and the parameter update in step S24, ensuring complete temporal synchronization between audio mixing and parameter updates. Within the encrypted interpolation region, since the parameter update frequency is higher than the sampling frequency of the mixing coefficients, the mixing coefficient values ​​used during parameter updates are obtained by linear interpolation of the mixing coefficients of adjacent sample points, ensuring that both processing paths share the same mixing coefficient trajectory.

[0063] Through the above steps, this embodiment achieves a differentiated parameter update frequency based on the slope characteristics of the cosine curve. Specifically, the encrypted interpolation strategy in the accelerating and decelerating transition regions allows for more frequent parameter updates in areas where the mixing coefficients change rapidly, preventing the parameter trajectory from deviating from the theoretical curve; the conventional interpolation strategy in the linear transition region saves computational resources. Conventional cosine curve calculation and discretization are well-known techniques. The contribution of this embodiment lies in coordinating the geometric characteristics of the transition period with the parameter update frequency, improving parameter tracking accuracy without increasing hardware costs. Those skilled in the art, based on the above description, can apply this method to audio processing systems with real-time parameter update requirements without any inventive effort.

[0064] In one embodiment of the present invention, the step of obtaining the current anti-feedback parameter set and nominal processing latency of the audio processor includes: S221, During the system initialization phase, a unit pulse test signal is input to the audio processor, and the time difference from input to output is measured to obtain the static delay; or, the algorithm pipeline delay is read from the audio processor's datasheet. S222, the static delay is used as the nominal processing delay, and the read pointer offset of the delay compensation cache is configured; S223, In real-time operation, the music data stream is written to the delay compensation buffer, and the delayed data is read according to the read pointer offset; S224, when the nominal processing delay of the audio processor fluctuates dynamically due to parameter switching or mode change, the read pointer offset is adjusted in real time through an adaptive delay estimation algorithm, and the adjustment is limited to a preset range.

[0065] As described in steps S221-S224 above, the processing latency of the audio processor may fluctuate dynamically due to parameter switching or mode changes (for example, when switching from music mode to voice mode, the internal algorithm switches from bypass state to active noise cancellation state, which may increase the processing latency by several milliseconds). If the latency compensation is inaccurate, the voice stream and the delayed music stream will be misaligned in time, resulting in phase interference or echo effects during crossfading. To solve this problem, this embodiment provides an adaptive adjustment mechanism for dynamic fluctuations in processing latency.

[0066] In this embodiment, the digital signal processor (DSP) pre-obtains the nominal processing delay of the audio processor. This nominal value can be obtained by inputting a unit pulse test signal to the audio processor and measuring the input-output time difference, or by reading it from the audio processor's datasheet. This acquisition method is conventional; the contribution of this embodiment lies in the subsequent dynamic adjustment. When the audio processor's processing delay fluctuates dynamically due to parameter switching or mode changes, the DSP adjusts the read pointer offset of the delay compensation buffer in real time using an adaptive delay estimation algorithm. The adaptive delay estimation algorithm employs a cross-correlation method: the DSP simultaneously acquires the input speech signal and the processed output speech stream from the audio processor, calculates the cross-correlation function of the two signals, and finds the delay time corresponding to the peak of the cross-correlation function; this delay time is the current actual processing delay. The actual processing delay is compared with the current nominal processing delay, and the difference is calculated.

[0067] To prevent the introduction of audio gaps during adjustment, the digital signal processor (DSP) limits the step size of each adjustment to a single sample point and the adjustment rate to no more than one sample point per millisecond. Simultaneously, the adjustment range of the read pointer offset is limited to a preset range, determined based on the depth margin of the delay compensation buffer (e.g., when the buffer depth is 1.2 times the nominal delay, the maximum allowed adjustment is ±0.1 times the buffer depth). These step size and range limitations ensure smooth changes in delay compensation without exceeding buffer boundaries, thereby guaranteeing time alignment between the speech and music streams under all operating conditions and preventing quality degradation in crossfade due to sudden delay changes.

[0068] Through the above steps, this embodiment achieves adaptive tracking of dynamic fluctuations in audio processor processing latency. Conventional circular buffer operations and cross-correlation calculations are well-known techniques. The contribution of this embodiment lies in associating latency measurement with parameter switching events and introducing step size and range limitations during the adjustment process, thereby solving the audio desynchronization problem caused by changes in processing latency while ensuring stability. Those skilled in the art can implement this solution without inventive effort based on the above description.

[0069] like Figure 2 As shown, the present invention also provides a broadcast audio channel selection anti-feedback processing system, comprising: The frame receiving and parsing module is used to receive audio data frames through a digital bus and parse the audio type flag and audio data stream from the audio data frames; the audio type flag is used to indicate whether the current audio data stream is a speech signal or a music signal. A voice mode response module is used to respond to the audio type flag indicating a voice signal, treat the audio data stream as a voice data stream, and enter the voice processing mode. The voice processing execution unit is used to obtain the current anti-feedback parameter set and nominal processing delay of the audio processor, and process the voice data stream according to the current anti-feedback parameter set to obtain the processed voice stream. A music delay compensation unit is used to acquire a music data stream, apply the nominal processing delay amount to the music data stream, and obtain a delayed music stream. A crossfade mixing unit is used to initiate a crossfade transition period with a preset transition duration, obtain mixing coefficients according to a predetermined gradient curve during the crossfade transition period, and perform weighted mixing of the processed speech stream and the delayed music stream according to the mixing coefficients to obtain an output audio stream. The parameter synchronization update unit is used to synchronously update the current anti-whistling parameter set according to the mixing coefficient during the cross-fading transition period, so that the current anti-whistling parameter set gradually changes with the change of the mixing coefficient; The music mode response module is used to respond to the audio type flag indicating a music signal, treat the audio data stream as a music data stream, and enter the music processing mode: apply the nominal processing delay to the music data stream and directly use it as the output audio stream, and put the audio processor into a low power or bypass state.

[0070] The present invention also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of a broadcast audio channel selection anti-feedback processing method.

[0071] The present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of a broadcast audio channel selection anti-feedback processing method.

[0072] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, apparatus, article, or method that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, apparatus, article, or method. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, apparatus, article, or method that includes that element.

[0073] The above description is merely a preferred embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.

Claims

1. A method for anti-feedback processing of broadcast audio channel selection, characterized in that, include: Audio data frames are received via a digital bus, and audio type flags and audio data streams are parsed from the audio data frames; the audio type flags are used to indicate whether the current audio data stream is a speech signal or a music signal. In response to the audio type flag indicating a speech signal, the audio data stream is treated as a speech data stream, and the system enters speech processing mode. Obtain the current anti-feedback parameter set and nominal processing delay of the audio processor, and process the voice data stream according to the current anti-feedback parameter set to obtain the processed voice stream; Acquire a music data stream, apply the nominal processing delay to the music data stream, and obtain a delayed music stream; A crossfade transition period with a preset transition duration is initiated. During the crossfade transition period, a mixing coefficient is obtained according to a predetermined gradient curve. The processed speech stream and the delayed music stream are then weighted and mixed according to the mixing coefficient to obtain the output audio stream. During the crossfading transition period, the current anti-whistling parameter set is updated synchronously according to the mixing coefficient, so that the current anti-whistling parameter set changes gradually with the change of the mixing coefficient; In response to the audio type flag indicating a music signal, the audio data stream is treated as a music data stream and enters music processing mode: the music data stream is directly used as the output audio stream after applying the nominal processing delay, and the audio processor is placed in a low-power or bypass state.

2. The broadcast audio channel selection anti-feedback processing method according to claim 1, characterized in that, The step of synchronously updating the current anti-whistling parameter set according to the mixing coefficient includes: Obtain the set of linear parameters and the set of nonlinear parameters in the current anti-whistling parameter set; For each linear parameter in the set of linear parameters, a gradual update is performed according to a linear interpolation formula during the cross-fading transition period; For each nonlinear parameter in the set of nonlinear parameters, the value of the mixing coefficient is monitored during the crossfading transition period. When the mixing coefficient reaches or exceeds a preset threshold for the first time, the nonlinear parameter is switched to a pre-stored target nonlinear parameter value.

3. The broadcast audio channel selection anti-feedback processing method according to claim 2, characterized in that, The step of monitoring the value of the mixing coefficient during the cross-fading transition period, and switching the nonlinear parameter to a pre-stored target nonlinear parameter value when the mixing coefficient first reaches or exceeds a preset threshold, includes: During the cross-fading transition period, the current mixing coefficient is calculated for each sampling period, and it is determined whether the current mixing coefficient is greater than or equal to the preset threshold and whether the current switch has not been performed. When the conditions are met, execute the following sub-steps: Read the pre-stored target nonlinear parameter value associated with the source of the voice data stream; Write the target nonlinear parameter value into the corresponding register of the audio processor; Record the switch completion status to prevent repeated switches.

4. The broadcast audio channel selection anti-feedback processing method according to claim 1, characterized in that, The step of synchronously updating the current anti-whistling parameter set according to the mixing coefficient further includes: The howling suppression contribution of each parameter in the current anti-howling parameter set is obtained according to the pre-constructed contribution ranking model; The parameters in the current anti-whistling parameter set whose whistling suppression contribution is greater than or equal to a preset contribution threshold are classified as the first group of core parameters, and the remaining parameters are classified as the second group of non-core parameters. During the crossfading transition period, only the first set of core parameters are updated; at the same time, the difference between the current value and the target value of the second set of non-core parameters is recorded. After the crossfading transition period ends, a tiered delayed update strategy is adopted based on the absolute value of the difference: if the absolute value of the difference is greater than the second preset threshold, it is updated immediately; if the absolute value of the difference is less than or equal to the second preset threshold, it is updated in the next system idle cycle.

5. The broadcast audio channel selection anti-feedback processing method according to claim 1, characterized in that, The step of obtaining the mixing coefficient according to a predetermined gradient curve during the cross-drying transition period includes: A cosine curve is selected as the gradient curve for the mixing coefficient, wherein the mixing coefficient and derivative of the cosine curve are zero at the starting point and one at the ending point and the derivative is zero. Based on the change in the slope of the tangent of the cosine curve, the cross-fading transition period is divided into an acceleration gradient region, a linear gradient region, and a deceleration gradient region. The number of sample points during the transition period is calculated based on the sampling rate, and the current mixing coefficient value is calculated at each sample point according to the discretized form of the cosine curve. Within the acceleration and deceleration gradient regions, the mixing coefficients are calculated at a first sampling frequency; within the linear gradient region, the mixing coefficients are calculated at a second sampling frequency, wherein the first sampling frequency is higher than the second sampling frequency.

6. The broadcast audio channel selection anti-feedback processing method according to claim 1, characterized in that, The steps of obtaining the current anti-feedback parameter set and nominal processing latency of the audio processor include: During the system initialization phase, a unit pulse test signal is input to the audio processor, and the time difference from input to output is measured to obtain the static delay; Use the static delay as the nominal processing delay, and configure the read pointer offset of the delay compensation cache; In real-time operation, the music data stream is written to the delay compensation buffer, and the delayed data is read according to the read pointer offset; When the nominal processing delay fluctuates dynamically, the read pointer offset is adjusted in real time using an adaptive delay estimation algorithm, and the adjustment is limited to a preset range.

7. The broadcast audio channel selection anti-feedback processing method according to claim 1, characterized in that, It also includes the following steps: The identifier field, control field, and data field of the audio data frame are obtained according to the FDCAN protocol standard; Extract the audio type flag bit from the data field or the identifier field; The audio type of the current audio data frame is determined according to the logical value of the audio type flag bit, and the audio data of consecutive audio data frames are reassembled into the audio data stream according to the receiving timing. When the audio type flag is detected to switch from music signal to voice signal, the timestamp of the received audio data frame is recorded; The nominal processing delay is dynamically adjusted based on the deviation between the received timestamp and the local clock, and the received timestamp is used as the start time of the crossfade transition period. The parameter update delay deviation caused by bus transmission jitter is calculated based on the mapping relationship between the received timestamp and the start time of the crossfade transition period; when the absolute value of the parameter update delay deviation exceeds the first preset threshold, it is determined that there is abnormal jitter in the current bus. In response to the determination of the abnormal jitter, an offset correction factor is obtained based on the mixing coefficient and the parameter update delay deviation, and the triggering time of the step of synchronously updating the current anti-whistling parameter set based on the mixing coefficient is adjusted according to the correction factor.

8. A method for anti-feedback processing of broadcast audio channel selection, characterized in that, include: The frame receiving and parsing module is used to receive audio data frames through a digital bus and parse the audio type flag and audio data stream from the audio data frames; The audio type flag is used to indicate whether the current audio data stream is a speech signal or a music signal; A voice mode response module is used to respond to the audio type flag indicating a voice signal, treat the audio data stream as a voice data stream, and enter the voice processing mode. The voice processing execution unit is used to obtain the current anti-feedback parameter set and nominal processing delay of the audio processor, and process the voice data stream according to the current anti-feedback parameter set to obtain the processed voice stream. A music delay compensation unit is used to acquire a music data stream, apply the nominal processing delay amount to the music data stream, and obtain a delayed music stream. A crossfade mixing unit is used to initiate a crossfade transition period with a preset transition duration, obtain mixing coefficients according to a predetermined gradient curve during the crossfade transition period, and perform weighted mixing of the processed speech stream and the delayed music stream according to the mixing coefficients to obtain an output audio stream. The parameter synchronization update unit is used to synchronously update the current anti-whistling parameter set according to the mixing coefficient during the cross-fading transition period, so that the current anti-whistling parameter set gradually changes with the change of the mixing coefficient; The music mode response module is used to respond to the audio type flag indicating a music signal, treat the audio data stream as a music data stream, and enter the music processing mode: apply the nominal processing delay to the music data stream and directly use it as the output audio stream, and put the audio processor into a low power or bypass state.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.