Audio output end anti-howling method, device, equipment and program product

By combining Bluetooth AoA positioning and audio input terminal operation data to predict the motion trajectory of the audio input terminal, and judging its distance from the audio output terminal in real time, the audio intervention strategy is dynamically adjusted. This solves the response lag and nonlinear distortion problems of traditional anti-feedback solutions, realizes active prevention of feedback, and improves audio quality and coverage of applicable scenarios.

CN121967968APending Publication Date: 2026-05-01IFLYTEK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
IFLYTEK CO LTD
Filing Date
2026-03-13
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing technologies cannot effectively prevent feedback at the audio output end, and traditional solutions suffer from problems such as slow response, high nonlinear distortion, poor generalization ability, and high false trigger rate.

Method used

By combining Bluetooth AoA positioning algorithm with relevant data from the audio input device, the system predicts the motion trajectory coordinates of the audio input device, determines its distance from the audio output device in real time, presets risk thresholds, and dynamically adjusts audio intervention strategies, including audio suppression and haptic feedback, to achieve active anti-feedback.

Benefits of technology

It achieves real-time perception of the relative position of the audio input and output terminals, supports anti-feedback processing for different layout areas, reduces the risk of feedback, improves audio quality and adaptability to various scenarios, and avoids feedback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121967968A_ABST
    Figure CN121967968A_ABST
Patent Text Reader

Abstract

The invention discloses an audio output end anti-howling method, device, equipment and program product, and belongs to the intelligent interaction equipment anti-howling technical field, the method comprises the following steps: predicting motion track coordinates of an audio input end according to multi-modal data of the audio input end, the multi-modal data at least comprising positioning data and operation related data; and calculating the distance between the predicted movement track coordinate and the audio output end, comparing the distance with a preset risk threshold value, and determining whether to trigger an anti-howling intervention action according to a comparison result. According to the audio output end anti-howling method, device, equipment and program product provided by the invention, whether the audio input end is close to the audio output end area is accurately judged through fusion of the arrival angle positioning algorithm and the audio input end operation related data in combination with the pre-configured coordinates of the audio output end; real-time sensing of the relative distance and orientation of the audio input end and the audio output end is achieved, and anti-howling processing of the audio output ends of different layout areas is supported.
Need to check novelty before this filing date? Find Prior Art

Description

Audio output anti-feedback methods, devices, equipment and software products Technical Field

[0001] This invention belongs to the field of anti-feedback technology for intelligent interactive devices, and specifically relates to an anti-feedback method, device, equipment and program product for audio output terminals. Background Technology

[0002] Feedback, also known as howling, is caused by a portion of the audio output signal being picked up again by an audio input device such as a microphone, amplified, and then output again by the audio output device. This results in signal superposition and amplification, ultimately producing a piercing howling noise. Feedback not only severely affects sound quality and listening experience but can also damage audio equipment. Therefore, howling suppression has always been a key technical challenge in various audio applications such as teaching, conference rooms, theaters, and live performances.

[0003] Current mainstream anti-feedback solutions have significant limitations: traditional solutions require passive handling after feedback occurs, unable to intervene proactively based on location trends, resulting in a delayed response; traditional solutions are susceptible to frequency fluctuations and nonlinear distortion, and struggle to balance feedback suppression with voice quality, leading to low suppression accuracy; traditional solutions cannot adapt to different audio output layouts and dynamic motion scenarios of audio inputs, resulting in high false / missed trigger rates and poor generalization ability; traditional solutions require the introduction of decorrelation techniques, leading to high voice distortion rates. In summary, none of the traditional solutions fundamentally prevent feedback. Summary of the Invention

[0004] To address the aforementioned issues, this invention provides an audio output anti-feedback method, comprising: predicting the motion trajectory coordinates of the audio input terminal based on multimodal data of the audio input terminal, wherein the multimodal data includes at least positioning data and operation-related data; calculating the distance between the predicted motion trajectory coordinates and the audio output terminal; comparing the distance with a preset risk threshold; and determining whether to trigger an anti-feedback intervention action based on the comparison result.

[0005] Furthermore, the positioning data includes at least the relative positions of the audio input and audio output terminals; the operation-related data includes at least the three-axis angular velocity and the three-axis attitude angle of the audio input terminal.

[0006] Furthermore, before predicting the motion trajectory coordinates, the process includes a preprocessing step for the collected multimodal data, including: extracting continuous time coordinates from the running-related data, and determining instantaneous motion features of the audio input terminal based on the continuous time coordinates, wherein the instantaneous motion features include at least velocity and acceleration.

[0007] Furthermore, predicting the motion trajectory coordinates of the audio input terminal based on the multimodal data of the audio input terminal includes: using the Kalman filter algorithm to filter the positioning data, velocity data, and acceleration data at multiple consecutive moments to obtain filtered data at multiple moments; and fitting the filtered data at the multiple moments to obtain the motion trajectory coordinates.

[0008] Furthermore, in the distance risk assessment, multiple preset risk level thresholds are included, such as a Level 1 intervention threshold, a Level 2 intervention threshold, and a Level 3 intervention threshold. When the distance is less than the Level 1 intervention threshold but greater than or equal to the Level 2 intervention threshold, a low-risk comparison result is output. When the distance is less than the Level 2 intervention threshold but greater than or equal to the Level 3 intervention threshold, a medium-risk comparison result is output. When the distance is less than the Level 3 intervention threshold, a high-risk comparison result is output.

[0009] Furthermore, the relative position is determined based on the Bluetooth AoA positioning algorithm combined with the pre-configured coordinates of the audio output end.

[0010] Furthermore, triggering the anti-feedback intervention action includes: generating an intervention trigger signal when the distance is less than a preset first-level intervention threshold; responding to the intervention trigger signal, reading real-time audio data from the audio input terminal and inputting the real-time audio data into a pre-trained audio intervention algorithm model; the audio intervention algorithm model separating the target speech feature signal and potential feedback signal from the real-time audio data; adjusting the suppression coefficient for the feedback signal based on the determined risk level; using the adjusted suppression coefficient to suppress the feedback signal, and reconstructing the suppressed signal with the target speech feature signal to output the processed audio signal.

[0011] Furthermore, triggering the anti-feedback intervention includes: based on the determined risk level, calling the corresponding tactile feedback parameters; based on the called tactile feedback parameters, generating a drive control command; sending the drive control command to the vibration motor at the audio input terminal, driving the vibration motor to perform the corresponding tactile feedback action, and monitoring the change in the risk level to determine whether to continue or terminate the feedback.

[0012] The present invention also provides an anti-feedback device for an audio output terminal, the device comprising: a trajectory prediction module configured to predict the motion trajectory coordinates of the audio input terminal based on multimodal data of the audio input terminal, wherein the multimodal data includes at least positioning data and operation-related data; and an intervention module configured to calculate the distance between the predicted motion trajectory coordinates and the audio output terminal, compare the distance with a preset risk threshold, and determine whether to trigger an anti-feedback intervention action based on the comparison result.

[0013] The present invention also provides an electronic device, comprising: a memory and a processor; the memory is connected to the processor and is used to store a program; the processor is used to implement the audio output anti-feedback method as described in the present invention by running the program in the memory.

[0014] The present invention also provides a computer program product, including computer program instructions; when the computer program instructions are executed by a processor, the processor causes the processor to execute the audio output anti-feedback method of the present invention.

[0015] Compared with existing technologies, the present invention has the following advantages: The audio output anti-feedback method, device, equipment, and program products of the present invention, through the fusion of arrival angle positioning algorithm and audio input end operation-related data, combined with the pre-configured coordinates of the audio output end, accurately determine whether the audio input end is close to the audio output end area, realize real-time perception of the relative distance and direction between the audio input end and the audio output end, support anti-feedback processing of audio output ends in different layout areas, increase the coverage of adaptable scenarios, and use a non-deep learning method to judge the trend of the audio input end approaching the audio output end 1-3 seconds in advance, and start the audio intervention algorithm when the distance is close to the threshold, thus weakening the feedback without waiting for it to occur. The replay signal achieves "active suppression guided by positional trends" to avoid feedback. The audio intervention algorithm can accurately separate the target speech from the replay signal, reduce the attenuation rate of the replay signal, and is highly robust to nonlinear distortion. At the same time, the audio output anti-feedback method of this invention classifies risk levels based on predicted distance and triggers a graded strategy of "vibration reminder + audio intervention" accordingly. It dynamically adjusts the intensity of audio intervention and vibration parameters to ensure that there is no feedback throughout the process and avoids unnecessary intervention when there is no risk. The vibration motor is deployed in the middle of the audio input end that fits the finger grip area. It enhances recognition through tactile reminders without affecting writing stability, which is superior to existing light and buzzer reminder solutions.

[0016] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures pointed out in the description, claims and drawings. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 shows a schematic flowchart of the audio output anti-feedback method according to an embodiment of the present invention; Figure 2 shows a schematic diagram of the audio output anti-feedback system according to an embodiment of the present invention; Figure 3 shows a schematic diagram of the electronic device according to an embodiment of the present invention. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] Application Overview Existing technical solutions for handling feedback noise mainly revolve around frequency signal processing and simple position sensing, and can be divided into the following four typical implementation methods: 1. Notch filter scheme based on fixed frequency suppression: This type of scheme integrates a notch filter at the audio output end, pre-setting common feedback frequencies in the range of 1kHz-8kHz. When a signal in this frequency band appears in the output signal of the audio output end, the signal is attenuated by hardware circuitry to suppress feedback. However, hardware-level notch filters can only suppress fixed frequency signals and cannot adapt to feedback frequency fluctuations caused by dynamic position changes at the audio input end. 2. Software scheme based on adaptive feedback cancellation: This type of scheme uses a digital signal processing chip to collect the input signal of the audio input end in real time, and uses algorithms such as Kalman filtering and normalized least mean square to estimate the acoustic path, generating a reverse signal to cancel the replay component. However, this scheme needs to be triggered after feedback occurs, resulting in response lag and nonlinear distortion at the audio output end. The robustness is poor, and it is also necessary to consider the technical problem of reducing speech distortion rate by introducing decorrelation technology to reduce the correlation between the target signal and the replay signal; passive reminder schemes based on a single distance sensor: a few schemes attempt to integrate an infrared distance sensor into the head of the audio input end. When the distance between the audio input end and the audio output end is less than a preset threshold, an LED light flashes or a buzzer sounds to remind the user to move away; however, this scheme can only sense vertical distance and cannot distinguish the position of the audio output end. The light / buzzer reminder has low recognition in bright light or noisy environments and cannot effectively prevent the user from approaching, and ultimately feedback will still occur; post-event suppression schemes based on traditional deep learning: the feedback is treated as noise and separated from the audio signal through deep learning; however, this scheme requires post-processing after the feedback occurs and does not combine position information for early intervention; and the model does not consider the characteristics of recurrent feedback, resulting in poor generalization ability in real-world scenarios and large fluctuations in suppression accuracy when the replay signal strength changes.

[0021] Exemplary methodsIn view of the above problems, this invention proposes an anti-feedback method for audio output terminals. In this method, the Angle of Arrival (AoA) positioning algorithm is fused with relevant data of the audio input terminal, and combined with the pre-configured coordinates of the audio output terminal, to accurately determine whether the audio input terminal is close to the audio output terminal area. This enables real-time perception of the relative distance and orientation between the audio input terminal and the audio output terminal, supports anti-feedback processing for audio output terminals in different layout areas, and increases the coverage of adaptable scenarios.

[0022] The following description, with reference to Figure 1, illustrates the audio output anti-feedback method provided by this embodiment of the invention. This embodiment can be applied to any scenario requiring anti-feedback. For example, in the field of smart education, with the rapid penetration of smart education and remote work, smart screens have become core interactive carriers integrating writing, voice input, and audio playback. Their accompanying smart pens, with built-in microphones supporting voice annotation and real-time explanation, have become key audio input devices in teaching scenarios. The corresponding large-screen speakers have become key output devices in teaching scenarios. The anti-feedback method in this embodiment can prevent feedback problems caused by close contact between the smart pen and the speaker during teaching. Similarly, in theaters and live performances, the anti-feedback method in this embodiment can prevent feedback problems caused by close contact between the microphone and the speaker during performances.

[0023] To address at least one problem existing in the prior art, this invention provides an audio output anti-feedback method. This method utilizes a software algorithm to sense the relative positions of the audio input and output terminals in real time, predict their proximity trend, and activate an audio intervention algorithm when the distance approaches a threshold. This achieves active suppression guided by position prediction, significantly improving the anti-feedback effect and audio quality. It should be noted that the software algorithm implementing this audio output anti-feedback method can run on any device with data processing capabilities, such as a local computer, a cloud server, or a smart mobile device.

[0024] Figure 1 shows one of the schematic flowcharts of the audio output anti-feedback method according to an embodiment of the present invention. As shown in Figure 1, the method includes: Step 101, determining the multimodal data of the audio input terminal. In this embodiment, the audio input terminal refers to any device with acoustic signal acquisition capabilities, such as a microphone or a terminal with a built-in microphone. The audio output terminal refers to any device with electroacoustic signal restoration and playback capabilities, such as a speaker, loudspeaker, or other loudspeaker terminal. It should be noted that the embodiments of the present invention do not specifically limit the types of audio input terminals and audio output terminals. The above examples are only non-exhaustive and non-limiting examples provided to help understand the present invention. The core innovation of the present invention lies in a dynamic anti-feedback control method based on forward prediction, and its protection scope is not limited to any specific type of audio input or output terminal listed in the embodiments.

[0025] In this embodiment, the audio input terminal is communicatively connected to the data processing terminal. The data processing terminal predicts the motion trajectory coordinates of the audio input terminal based on the multimodal data of the audio input terminal, and triggers the anti-feedback intervention action of the audio output terminal based on the comparison result of the motion trajectory coordinates with the distance of the audio output terminal and the preset risk threshold.

[0026] In this embodiment, the multimodal data of the audio input terminal includes at least positioning data and operation-related data. Specifically, the positioning data is determined by the Bluetooth 6.0 AoA positioning algorithm. The AoA positioning function supported by Bluetooth 6.0 technology of the audio input terminal can achieve coarse positioning with an accuracy of ±5cm within a range of 1-3 meters. The Bluetooth 6.0 gateway of the data processing terminal supports the AoD positioning function. The audio input terminal establishes a connection with the Bluetooth gateway of the data processing terminal, and the Bluetooth positioning sampling frequency is set to fbt=10Hz.

[0027] In this embodiment, the positioning data includes at least the relative positions of the audio input and audio output. The relative positions are determined based on the Bluetooth AoA positioning algorithm combined with the pre-configured coordinates of the audio output. The data processing terminal Bluetooth gateway deploys a dual-antenna array with a 10cm spacing. The relative positions are calculated by measuring the angle of arrival (Angle of Arrival) and the received signal strength (RSSI) of the Bluetooth signal at the audio input. Determining the relative positions of the audio input and audio output includes: AoA calculation: θ = arcsin((Δφ) / (Δφ)) In the formula λ) / (2πd)), θ represents the angle of arrival, Δφ represents the phase difference, λ represents the spacing between the dual antenna arrays deployed in the Bluetooth gateway of the data processing terminal, and d represents the diameter.

[0028] The relationship between RSSI and distance is estimated as follows: r = 10 (A - RSSI) / (10n)In the formula, r represents the straight-line distance between the audio input and audio output, A represents the received signal strength at the reference distance, and n represents the path loss exponent. It should be noted that in this example, the received signal strength at the reference distance and the path loss exponent are the core parameters characterizing the signal propagation characteristics in Bluetooth wireless communication.

[0029] To determine the relative positions of the audio input and output terminals, with the center of the audio output terminal as the origin, the positioning coordinates of the audio input terminal are represented as: (x, y) = (r cosθ, r Specifically, in this embodiment, the audio output terminal is equipped with a positioning calibration. The user establishes a two-dimensional coordinate system with a point on the data processing terminal as the origin through the configuration interface of the data processing terminal, inputs the physical coordinate range of the audio output terminal, and presets multiple calibration points in front of the configuration interface. The audio input terminal stays at each calibration point for 1 second, collecting raw Bluetooth positioning data of the audio input terminal. This raw Bluetooth positioning data includes the relative distance between the audio input terminal and the audio output terminal. The relative angle between the audio input and audio output terminals The calibration coefficients are solved using the least squares method. In this embodiment, the calibrated Bluetooth positioning data of the audio input terminal is represented as follows:

[0030]

[0031] In the formula, , Both represent calibration coefficients for relative distance. , All represent calibration coefficients for relative angles. By calibrating the position of the audio input terminal in this invention, the positioning error after calibration is ≤ ±5cm. The audio input terminal sends an AoA positioning signal every 100ms, and the data processing terminal calculates the real-time coordinates after calibration.

[0032] In this embodiment, the operation-related data includes at least the three-axis angular velocity and the three-axis attitude angle of the audio input terminal, and the operation-related data is sensed by the built-in sensor of the audio input terminal.

[0033] In one possible implementation, the operational data of the audio input terminal can be acquired using a gyroscope integrated within the audio input terminal. This gyroscope can sense the rotation angle, angular velocity, and acceleration of the audio input terminal in three-dimensional space in real time. Because the gyroscope is built into the audio input terminal, it is not affected by external factors when acquiring operational data, thus improving the reliability and stability of the data acquisition. The gyroscope acquires the three-axis angular velocity ω every 50ms.t =(ω x,t ,ω y,t ,ω z,t ) and three-axis attitude angle α t =(θ t ,φ t ,ψ t In this embodiment, the sampling frequency of the gyroscope is also illustrated with an example. The sampling frequency f of the gyroscope is... gyro It can be set to 20Hz. In this embodiment, there is no specific limitation on the sampling frequency of the gyroscope.

[0034] In this embodiment, before predicting the motion trajectory coordinates, a step of preprocessing the collected multimodal data is also included, including: removing abnormal data from the positioning data, wherein the abnormal data includes at least data that does not conform to the statistical distribution characteristics; extracting the continuous time coordinates from the running related data, and determining the instantaneous motion characteristics of the audio input end based on the continuous time coordinates, wherein the instantaneous motion characteristics include at least velocity and acceleration.

[0035] In this embodiment, the preprocessing of multimodal data includes: outlier removal from location data: using the 3σ criterion, if the current coordinate (x... t y t If the deviation from the mean of the previous time step exceeds three times the standard deviation, then the coordinates of the previous time step are used to replace the coordinates of the current time step. It should be noted that the method for removing outliers in the location data in this embodiment is only an example. Other relatively complex anomaly detection methods include principal component analysis (PCA), similarity method, and isolated forest, etc. This invention does not specifically limit the description of this particular outlier removal algorithm.

[0036] Motion parameter extraction: Based on the coordinates of three consecutive moments, calculate velocity and acceleration. Velocity is expressed as:

[0037] Acceleration is expressed as: .

[0038] Step 102: Predict the motion trajectory coordinates of the audio input terminal based on the multimodal data of the audio input terminal.

[0039] In this embodiment, predicting the motion trajectory coordinates of the audio input terminal based on the multimodal data of the audio input terminal includes: using the Kalman filter algorithm to filter the positioning data, velocity data, and acceleration data at multiple consecutive moments to obtain filtered data at multiple moments; and fitting the filtered data at the multiple moments to obtain the motion trajectory coordinates.

[0040] Specifically, in this embodiment, filtering and fitting the acquired multimodal data includes: constructing a state vector representing the motion state of the audio input terminal based on the acquired multimodal data. The state vector includes at least the position coordinates, velocity components, and acceleration components of the audio input terminal. The position coordinates are determined based on the positioning data of the audio input terminal, and the velocity and acceleration components are determined based on the velocity and acceleration in the relevant running data. Using the state vector as input, the extended Kalman filter algorithm is used to perform data fusion and filtering on the motion state of the audio input terminal at the current moment, and the optimal state estimate at the current moment is output. Based on the optimal state estimates of multiple consecutive historical moments, including the optimal state estimate at the current moment, a polynomial fitting algorithm is used to predict the motion trajectory coordinates of the audio input terminal.

[0041] Predicting the motion trajectory coordinates of the audio input terminal specifically includes the positioning position (x, y) and velocity (v) of the audio input terminal. x ,v y ), acceleration (a x ,a y Create a 6-dimensional state vector X t =[x,y,v x ,v y ,a x ,a y ] T To establish a prediction and update model for the Extended Kalman Filter (EKF) algorithm: First, define the relevant parameters in the EKF process. The noise covariance Q is represented by diagonal elements [0.01, 0.01, 0.001]; the observation noise covariance R is represented by diagonal elements [0.02, 0.02].

[0042] Define the prediction step: X t|t-1 = F X t-1 + G w t Where F represents the state transition matrix, X t-1 This represents a 6-dimensional column vector, fully characterizing the motion state of the audio input terminal in a two-dimensional plane at time t-1, including three key physical quantities: position, velocity, and acceleration. It corresponds to the 6 rows of the state transition matrix F, where G represents the identity matrix, and w... tLet X represent the process noise, which follows the order N(0,Q). The state transition matrix F is represented as: [1, Δt, (1 / 2)Δt², 0, 0, 00, 1, Δt, 0, 0, 00, 0, 1, 0, 0, 00, 0, 1, Δt, (1 / 2)Δt²0, 0, 0, 0, 1, Δt0, 0, 0, 0, 0, 1] Observation correction update step: State correction representation X t =X t|t-1 +K t (Z t -H X t|t-1 ), where the observation matrix H=[[1,0,0,0,0,0],[0,0,0,1,0,0]], and the Kalman gain K t =P t|t-1 H T (H P t|t-1 H T + R) -1 Observation vector Z t =[x t ,y t ] T P t|t-1 This represents the covariance matrix of the prediction of the current state from the previous moment.

[0043] The optimal state estimate at the current moment is expressed as: X_t^=[x_t^, y_t^, v_x,t^, v_y,t^, a_x,t^,a_y,t^] ]^T.

[0044] In this embodiment, cubic polynomial trajectory prediction is performed based on the optimal state estimation of the last 5 time steps. The motion trajectory in the x and y directions is fitted using a cubic polynomial to predict the coordinates within the next 2 seconds (20 time steps): P^ = {(x_t+1^, y_t+1^), ..., (x_t+20^, y_t+20^)}. Further, the fitting in the x and y directions is determined based on the predicted coordinates: x-direction fitting: x^(τ) = a_x τ³ + b_x τ² + c_x Fitting the y-direction to τ + d_x: y^(τ) = a_y τ³ + b_y τ² + c_y τ + d_y, where τ is the prediction time, and at the current time t, τ=0.

[0045] In this embodiment, the combination of EKF and cubic polynomial fitting can adapt to the nonlinear motion of the audio input end. EKF suppresses random noise in the positioning data by dynamically adjusting process noise and observation noise, thereby reducing the error between position estimation and original data. Compared with linear models (uniform speed, uniform acceleration), cubic polynomial can fit complex trajectories such as circular arcs and variable acceleration, reduce the prediction error rate, and ensure accurate judgment of the approach trend.

[0046] Step 103: Calculate the distance between the predicted motion trajectory coordinates and the audio output terminal, compare the distance with a preset risk threshold, and determine whether to trigger the anti-feedback intervention action based on the comparison result.

[0047] In this embodiment, the distance between the predicted motion trajectory coordinates and the audio output terminal is calculated and compared with a preset risk threshold, and the comparison result is output. Based on the comparison result, it is determined whether the intervention conditions are met, and the anti-feedback intervention action is triggered when the conditions are met.

[0048] In this embodiment, the distance risk assessment uses multiple preset risk level thresholds, including a Level 1 intervention threshold, a Level 2 intervention threshold, and a Level 3 intervention threshold. When the distance is less than the Level 1 intervention threshold but greater than or equal to the Level 2 intervention threshold, a low-risk comparison result is output; when the distance is less than the Level 2 intervention threshold but greater than or equal to the Level 3 intervention threshold, a medium-risk comparison result is output; and when the distance is less than the Level 3 intervention threshold, a high-risk comparison result is output. Specifically, in this embodiment, the Level 1 intervention threshold is the low-risk threshold, and the low-risk threshold D... low The threshold is set at 30cm, and the secondary intervention threshold is the medium-risk threshold, with the medium-risk threshold being D. mid The threshold is set at 20cm, and the level 3 intervention threshold is the high-risk threshold, D. high Set to 10cm.

[0049] In this embodiment, triggering the anti-feedback intervention action includes: generating an intervention trigger signal when the distance is less than a preset first-level intervention threshold; responding to the intervention trigger signal, reading real-time audio data from the audio input terminal and inputting the real-time audio data into a pre-trained audio intervention algorithm model; the audio intervention algorithm model separating the target speech feature signal and potential feedback signal from the real-time audio data; adjusting the suppression coefficient for feedback signal based on the determined risk level; using the adjusted suppression coefficient to suppress the feedback signal, and reconstructing the suppressed signal with the target speech feature signal to output the processed audio signal.

[0050] In this embodiment, the pre-training of the audio intervention algorithm model can be completed by installing it on the data processing terminal after training at the factory or by installing it on the data processing terminal. This embodiment does not specify the pre-training method of the audio intervention algorithm model. This invention focuses on the pre-training capability of the audio intervention algorithm model and the performance advantages it brings, rather than the specific time or physical location of its pre-training. As long as the model has undergone sufficient offline learning before intervention decision, it should fall within the protection scope of this embodiment.

[0051] In this embodiment, a non-deep learning method is used to determine the trend of the audio input end moving closer to the audio output end 1-3 seconds in advance. When the distance approaches the threshold, the audio intervention algorithm is activated. The replay signal can be weakened without waiting for feedback to occur, thus achieving "active suppression guided by position trend" and avoiding feedback. The audio intervention algorithm can accurately separate the target speech from the replay signal, reduce the attenuation rate of the replay signal, and has strong robustness to nonlinear distortion.

[0052] In this embodiment, the trend of the audio input end moving closer to the audio output end is determined by a non-deep learning method. Specifically, the trend determination is completed in milliseconds or even sub-milliseconds by using a non-deep learning method of Kalman filtering and polynomial fitting, ensuring the true realization of "early intervention" and extremely low overall system response latency.

[0053] The specific steps to trigger anti-feedback intervention include: collecting real-time audio data. mic (t) The audio input terminal collects one frame of audio signal every 16ms and transmits it to the data processing terminal. The audio data is preprocessed, including pre-emphasis, framing, and adding Hanning windows, in preparation for audio intervention input.

[0054] In this embodiment of the invention, the continuous raw audio signal is converted into a short-time stable frame signal suitable for subsequent howling detection / suppression by preprocessing the audio data. At the same time, the recognition of high-frequency features is improved and the inter-frame distortion is reduced. For example, the pre-emphasis process of audio data is also illustrated: for audio signals with low high-frequency component energy that will attenuate during propagation / acquisition, this embodiment of the invention uses pre-emphasis to increase the amplitude of high-frequency components, enhance the recognition of howling features, and balance the energy distribution of high and low frequencies.

[0055] A first-order high-pass filter is used to process each sampling point, with the formula: s′(q)=s(q) α s(q 1), where s′(q) represents the q-th sampling point after pre-emphasis, s(q) represents the q-th sampling point of the original audio signal, and α represents the pre-emphasis coefficient. For example, in this embodiment of the invention, the value range of the pre-emphasis coefficient α is 0.9~0.97. Initial condition: s( 1)=0 (The first sampling point is directly taken as s′(0)=s(0)).

[0056] Regarding the voice / feedback signal of this invention, since the audio signal collected by the audio input terminal is usually 8kHz / 16kHz, the pre-emphasis coefficient α is set to 0.9375. This value can effectively enhance the high frequency of the high-frequency feedback band above 2kHz without introducing too much noise.

[0057] Calculate the minimum distance d between each predicted coordinate and the audio output region. t+i d t+i ∈[20cm,30cm) is considered low risk, d t+i ∈[10cm,20cm) is considered medium risk, d t+i A depth of less than 10cm is considered high-risk.

[0058] In this embodiment, triggering the anti-feedback intervention action includes: calling the corresponding tactile feedback parameters based on the determined risk level; generating a drive control command based on the called tactile feedback parameters; sending the drive control command to the vibration motor at the audio input terminal to drive the vibration motor to perform the corresponding tactile feedback action, and monitoring changes in the risk level to determine whether to continue or terminate the feedback.

[0059] The audio output anti-feedback method of this invention classifies risk levels based on predicted distance and triggers a graded strategy of "vibration reminder + audio intervention" accordingly. It dynamically adjusts the intensity of audio intervention and vibration parameters to ensure that no feedback occurs throughout the process and avoids unnecessary intervention when there is no risk. The vibration motor is deployed in the middle section of the audio input end, which fits the finger grip area. The tactile reminder enhances the recognition without affecting writing stability, which is superior to existing light and buzzer reminder solutions.

[0060] The tiered vibration alert process specifically includes: the data processing terminal sending vibration control commands to the audio input via Bluetooth, adjusting parameters according to the risk level. The vibration motor parameters are initialized to three levels: low (50Hz / 0.5mm), medium (100Hz / 1.0mm), and high (200Hz / 1.5mm). No vibration is triggered for low risk; a slight vibration (low level) lasts for 200ms for medium risk; and a strong vibration (high level) lasts for 400ms for high risk. This process repeats every 100ms until the risk is eliminated.

[0061] In this embodiment, the vibration motor is deployed in the middle of the audio input end (fitting the finger grip area) to ensure accurate user perception without affecting stability.

[0062] In this embodiment, a "distance-risk-inhibition" mapping relationship is established through the synergy of audio intervention and risk level. The closer the audio input end and the audio output end are, the stronger the inhibition. The inhibition coefficient λ is inversely proportional to the real-time distance. In this embodiment, the inhibition coefficient λ is adjusted in the range of λ=0.3~0.8. It should be noted that this embodiment does not make specific limitations on the selection of the inhibition coefficient.

[0063] In this embodiment, when the distance is less than the preset first-level intervention threshold, the data processing terminal starts the audio intervention algorithm and adjusts the suppression intensity according to the risk level. Specifically, in this embodiment, the potential howling signal is controlled by the suppression coefficient λ. The specific process includes: reading the pre-processed audio frame data, inputting it into the pre-trained audio intervention algorithm model, the model extracting the target speech features through signal separation technology, weakening the replay signal, i.e., the potential howling signal, and adjusting the suppression coefficient λ according to the risk level. When the risk is low, weak suppression and sound quality preservation measures are adopted, and the suppression coefficient λ is set to 0.3. When the risk is medium, medium suppression measures are adopted, and the suppression coefficient λ is set to 0.5. When the risk is high, strong suppression and priority howling elimination measures are adopted, and the suppression coefficient λ is set to 0.8. The processed target speech signal is output through the audio output terminal. Before output, the volume is normalized to avoid sudden volume changes.

[0064] This embodiment also includes a step of providing feedback on the intervention effect. The specific process of providing feedback on the intervention effect includes: collecting feedback data every 100ms and dynamically adjusting the strategy, if the real-time distance d t+i When the value increases above the low-risk threshold, it indicates that the risk has been eliminated. At this point, the suppression coefficient λ is reduced to 0 and the vibration stops. If a residual howling signal remains after processing, the suppression coefficient λ is increased. If the real-time distance remains within the current risk range and there is no residual signal, the current intervention parameters are maintained. In this embodiment, the presence or absence of the residual howling signal is determined by spectrum analysis. It should be noted that this invention does not specifically limit the determination method.

[0065] In this embodiment, the suppression coefficient λ is increased by an increment of 0.05. In practical applications, the suppression coefficient λ can be adaptively adjusted according to the residual howling signal.

[0066] The audio output anti-feedback method in this embodiment also includes the step of updating trajectory prediction results and risk level. Specifically, the periodic trajectory prediction update is repeated every 100ms to predict the motion trajectory coordinates of the audio input end, update the trajectory prediction results and risk level, and dynamically adjust the audio intervention intensity and vibration parameters to ensure that no feedback occurs throughout the process and the voice quality is stable.

[0067] Exemplary device Accordingly, this embodiment of the invention also provides an audio output anti-feedback system. Figure 2 shows a schematic diagram of the audio output anti-feedback system in this embodiment of the invention. As shown in Figure 2, the audio output anti-feedback system is installed in the data processing terminal of this invention and includes: a trajectory prediction module 201, configured to predict the motion trajectory coordinates of the audio input terminal based on the multimodal data of the audio input terminal, wherein the multimodal data includes at least positioning data and operation-related data; and an intervention module 202, configured to calculate the distance between the predicted motion trajectory coordinates and the audio output terminal, compare the distance with a preset risk threshold, and determine whether to trigger an anti-feedback intervention action based on the comparison result.

[0068] In one embodiment, the intervention module 202 is configured to calculate the distance between the predicted motion trajectory coordinates and the audio output terminal, compare the distance with a preset risk threshold, and output the comparison result; based on the comparison result, determine whether the intervention conditions are met, and trigger the anti-feedback intervention action when the conditions are met.

[0069] In one embodiment, the audio output anti-feedback system further includes a data acquisition module 203 configured to acquire and preprocess multimodal data.

[0070] The positioning data collected by the acquisition module 203 includes at least the relative position of the audio input end and the audio output end. The relative position is determined based on the Bluetooth AoA positioning algorithm combined with the pre-configured coordinates of the audio output end. The relevant data includes at least the three-axis angular velocity and the three-axis attitude angle of the audio input end.

[0071] In one embodiment, the acquisition module 203 is further configured to remove abnormal data from the positioning data, wherein the abnormal data includes at least data that does not conform to statistical distribution characteristics; extract continuous time coordinates from the running related data, and determine the instantaneous motion characteristics of the audio input terminal based on the continuous time coordinates, wherein the instantaneous motion characteristics include at least velocity and acceleration.

[0072] In one embodiment, the trajectory prediction module 201 is configured to predict the motion trajectory coordinates of the audio input terminal based on the multimodal data of the audio input terminal, including: using a Kalman filter algorithm to filter the positioning data, velocity data, and acceleration data at multiple consecutive time points to obtain filtered data at multiple time points; and fitting the filtered data at the multiple time points to obtain the motion trajectory coordinates.

[0073] In one embodiment, the audio output anti-feedback system further includes a pre-configuration module. This module is configured to preset multiple risk level thresholds in distance risk assessment, including a level 1 intervention threshold, a level 2 intervention threshold, and a level 3 intervention threshold. It is also configured to pre-configure the audio output location, positioning parameters, and risk thresholds to achieve scenario adaptation.

[0074] In one embodiment, the output comparison results include: when the distance is less than the first-level intervention threshold and greater than or equal to the second-level intervention threshold, a low-risk comparison result is output; when the distance is less than the second-level intervention threshold and greater than or equal to the third-level intervention threshold, a medium-risk comparison result is output; and when the distance is less than the third-level intervention threshold, a high-risk comparison result is output.

[0075] In one embodiment, the intervention module 202 is configured to trigger an anti-feedback intervention action, including: generating an intervention trigger signal when the distance is less than a preset first-level intervention threshold; in response to the intervention trigger signal, reading real-time audio data from the audio input terminal and inputting the real-time audio data into a pre-trained audio intervention algorithm model; the audio intervention algorithm model separating the target speech feature signal and potential feedback signal from the real-time audio data; adjusting the suppression coefficient for feedback signal based on the determined risk level; using the adjusted suppression coefficient to suppress the feedback signal, reconstructing the suppressed signal with the target speech feature signal, and outputting the processed audio signal.

[0076] In one embodiment, the intervention module 202 is configured to trigger an anti-feedback intervention action, including: calling the corresponding tactile feedback parameters based on the determined risk level; generating a drive control command based on the called tactile feedback parameters; sending the drive control command to the vibration motor at the audio input terminal to drive the vibration motor to perform the corresponding tactile feedback action, and monitoring changes in the risk level to determine whether to continue or terminate the feedback.

[0077] The audio output anti-feedback system in this embodiment belongs to the same concept as the audio output anti-feedback method provided in the above embodiments of this invention. It can execute the audio output anti-feedback method provided in any of the above embodiments of this invention and has the corresponding functional modules and beneficial effects of the method. Technical details not described in detail in this embodiment can be found in the specific processing content of the audio output anti-feedback method provided in the above embodiments of this invention, and will not be repeated here.

[0078] Exemplary electronic devices This invention also provides an electronic device. Figure 3 shows a schematic diagram of the electronic device structure in this invention. As shown in Figure 3, the electronic device includes a memory 300 and a processor 301.

[0079] The memory 300 is connected to the processor 301 and is used to store programs.

[0080] The processor 301 is used to implement the audio output anti-feedback method in the above embodiments by running the program stored in the memory 300.

[0081] Specifically, the aforementioned electronic device may also include: a communication interface 302, an input device 303, an output device 304, and a bus 305.

[0082] The processor 301, memory 300, communication interface 302, input device 303, and output device 304 are interconnected via a bus. The bus 305 may include a pathway for transmitting information between the various components of the computer system.

[0083] The processor 301 can be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or an application-specific integrated circuit (ASIC), or one or more integrated circuits used to control the execution of the program of the present invention.

[0084] Processor 301 may include a main processor, as well as a baseband chip, modem, etc.

[0085] The memory 300 stores a program that executes the technical solution of this invention, and may also store an operating system and other key business functions. Specifically, the program may include program code, which includes computer operation instructions. More specifically, the memory 300 may include read-only memory (ROM), other types of static storage devices capable of storing static information and instructions, random access memory (RAM), other types of dynamic storage devices capable of storing information and instructions, disk storage, flash memory, etc.

[0086] The communication interface 302 may include a device that uses any transceiver to communicate with other devices or communication networks, such as Ethernet, Radio Access Network (RAN), Wireless Local Area Network (WLAN), etc.

[0087] The processor 301 executes the program stored in the memory 300 and calls other devices, which can be used to implement the various steps of the audio output anti-feedback method provided in the above embodiments of the present invention.

[0088] Exemplary computer program products and storage media In addition to the methods and devices described above, embodiments of the present invention may also be computer program products, which include computer program instructions that, when executed by a processor, cause the processor to perform the steps in the audio output anti-feedback method described in the embodiments of the present invention.

[0089] Computer program products can be written in any combination of one or more programming languages ​​to perform the operations of the embodiments of the present invention. The programming languages ​​include object-oriented programming languages ​​such as Java and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0090] Furthermore, embodiments of the present invention may also be storage media storing a computer program, which is executed by a processor in the steps of the audio output anti-feedback method described in the present invention embodiments.

[0091] For the foregoing method embodiments, in order to simplify the description, they are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, because according to the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.

[0092] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For apparatus embodiments, since they are basically similar to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.

[0093] The steps in the methods of the various embodiments of the present invention can be adjusted, merged, or deleted in order according to actual needs, and the technical features described in the various embodiments can be replaced or combined.

[0094] The modules and sub-modules in the systems and terminals provided in the various embodiments of the present invention can be merged, divided, and deleted according to actual needs.

[0095] In the embodiments provided by this invention, it should be understood that the disclosed terminals, devices, and methods can be implemented in other ways. For example, the terminal embodiments described above are merely illustrative. For instance, the division of modules or sub-modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple sub-modules or modules may be combined or integrated into another module, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection through some interfaces, devices, or modules, and may be electrical, mechanical, or other forms.

[0096] The modules or submodules described as separate components may or may not be physically separate. The components that constitute a module or submodule may or may not be physical modules or submodules; that is, they may be located in one place or distributed across multiple network modules or submodules. Some or all of the modules or submodules can be selected to achieve the purpose of this embodiment's solution, depending on actual needs.

[0097] Furthermore, the functional modules or sub-modules in the various embodiments of the present invention can be integrated into one processing module, or each module or sub-module can exist physically separately, or two or more modules or sub-modules can be integrated into one module. The integrated modules or sub-modules described above can be implemented in hardware or in the form of software functional modules or sub-modules.

[0098] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed in this invention can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0099] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software unit executed by a processor, or a combination of both. The software unit can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0100] Finally, it should be noted that in this invention, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the term "comprising" or any other variations thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0101] Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for preventing feedback at an audio output terminal, characterized in that, The method includes: predicting the motion trajectory coordinates of the audio input terminal based on multimodal data of the audio input terminal, wherein the multimodal data includes at least positioning data and operation-related data; calculating the distance between the predicted motion trajectory coordinates and the audio output terminal; comparing the distance with a preset risk threshold; and determining whether to trigger an anti-feedback intervention action based on the comparison result.

2. The audio output anti-feedback method according to claim 1, characterized in that, The positioning data includes at least the relative positions of the audio input and audio output terminals; the operation-related data includes at least the three-axis angular velocity and the three-axis attitude angle of the audio input terminal.

3. The audio output anti-feedback method according to claim 1, characterized in that, Before predicting the motion trajectory coordinates, the method further includes a step of preprocessing the collected multimodal data, including: extracting continuous time coordinates from the running-related data, and determining the instantaneous motion feature quantity of the audio input terminal based on the continuous time coordinates, wherein the instantaneous motion feature quantity includes at least velocity and acceleration.

4. The audio output anti-feedback method according to claim 3, characterized in that, Predicting the motion trajectory coordinates of the audio input terminal based on multimodal data from the audio input terminal includes: using a Kalman filter algorithm to filter positioning data, velocity data, and acceleration data at multiple consecutive time points to obtain filtered data at multiple time points; and fitting the filtered data at the multiple time points to obtain the motion trajectory coordinates.

5. The audio output anti-feedback method according to claim 1, characterized in that, In the distance risk assessment, multiple preset risk level thresholds are included, such as a Level 1 intervention threshold, a Level 2 intervention threshold, and a Level 3 intervention threshold. When the distance is less than the Level 1 intervention threshold but greater than or equal to the Level 2 intervention threshold, a low-risk comparison result is output. When the distance is less than the Level 2 intervention threshold but greater than or equal to the Level 3 intervention threshold, a medium-risk comparison result is output. When the distance is less than the level 3 intervention threshold, a high-risk comparison result is output.

6. The audio output anti-feedback method according to claim 2, characterized in that, The relative position is determined based on the Bluetooth AoA positioning algorithm combined with the pre-configured coordinates of the audio output end.

7. The audio output anti-feedback method according to claim 5, characterized in that, Triggering anti-feedback intervention includes: generating an intervention trigger signal when the distance is less than a preset first-level intervention threshold; responding to the intervention trigger signal, reading real-time audio data from the audio input terminal and inputting the real-time audio data into a pre-trained audio intervention algorithm model; the audio intervention algorithm model separating the target speech feature signal and potential feedback signal from the real-time audio data; adjusting the suppression coefficient for the feedback signal based on the determined risk level; using the adjusted suppression coefficient to suppress the feedback signal, and reconstructing the suppressed signal with the target speech feature signal to output the processed audio signal.

8. The audio output anti-feedback method according to claim 5 or 7, characterized in that, Triggering anti-feedback intervention includes: based on the determined risk level, calling the corresponding tactile feedback parameters; based on the called tactile feedback parameters, generating a drive control command; sending the drive control command to the vibration motor at the audio input terminal, driving the vibration motor to perform the corresponding tactile feedback action, and monitoring the change in the risk level to determine whether to continue or terminate the feedback.

9. An anti-feedback device for audio output, characterized in that, The device includes: a trajectory prediction module configured to predict the motion trajectory coordinates of the audio input terminal based on multimodal data from the audio input terminal, wherein the multimodal data includes at least positioning data and operation-related data; and an intervention module configured to calculate the distance between the predicted motion trajectory coordinates and the audio output terminal, compare the distance with a preset risk threshold, and determine whether to trigger an anti-feedback intervention action based on the comparison result.

10. An electronic device, characterized in that, include: Memory and processor; The memory is connected to the processor and is used to store programs; The processor is used to implement the audio output anti-feedback method as described in any one of claims 1-8 by running the program in the memory.

11. A computer program product, characterized in that, Includes computer program instructions; when the computer program instructions are executed by the processor, the processor causes the processor to perform the audio output anti-feedback method as described in any one of claims 1-8.