AI original sound elimination method and device

By applying AI models in an audio processor, real-time identification and separation of vocals and accompaniment is solved, the problem that traditional technology cannot obtain music accompaniment in real time is solved, and efficient and real-time acoustic cancellation effect is achieved.

CN120148546AInactive Publication Date: 2025-06-13SHENZHEN XINGYIN CENTURY TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510326418.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-06-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional AI elimination technology cannot obtain plug-and-play music accompaniment in real time, and it is complex in operation and ineffective in effect.

Method used

Using the AI-based real-time acoustic cancellation method, the audio signal is processed in real time through the audio processor, the AI ​​model is used to identify and separate the vocals and accompaniment parts, and the accompaniment signal after the vocals is eliminated.

Benefits of technology

Real-time removal of vocals, obtain high-quality accompaniment signals, support multiple audio input sources, and connect through external devices, with a delay time of less than 50 milliseconds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148546A_ABST
    Figure CN120148546A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of AI original sound elimination, in particular to an AI original sound elimination method and device.The AI original sound elimination device comprises an audio input module, an AI processing module, an audio processor and an audio output module, audio signals from Bluetooth audio, external input, mobile phone OTG audio or PC audio are received, the audio signals are analyzed through an AI model, and the AI original sound is eliminated. The device identifies and separates human voice and accompaniment parts, then processes audio signals in real time through an audio processor so as to remove the human voice part, and finally outputs accompaniment signals after human voice elimination, and the device further provides a user interface which allows a user to adjust the intensity of human voice elimination, the volume of accompaniment and other audio parameters. Therefore, an accompaniment signal without human voice is obtained in real time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of AI original sound cancellation, and particularly to an AI original sound cancellation method and device. Background Art

[0002] Nowadays, mobile devices are basically popular, but the uneven quality of recording devices on the market results in problems such as background noise, noise or sound loss in voice recordings, thus affecting the communication effect. To solve this problem, AI cancellation technology has been designed.

[0003] Traditional karaoke requires professional equipment for vocal cancellation. This operation is not only complex but also has poor effects. Most early audio vocal removal technologies relied on traditional methods such as spectral analysis and equalizer adjustment, and it was impossible to obtain a plug-and-play music accompaniment in real time.

[0004] The purpose of the present invention is to provide a real-time original sound cancellation method and system based on AI technology, which can process audio signals in real time through an audio processor (DSP), remove the vocal part, and output an accompaniment signal after removing the vocal. This method supports multiple audio input sources, including Bluetooth devices, external input (AUX), mobile phone OTG audio, and PC audio, and can be connected through external devices, Bluetooth devices, smartphones and other devices to obtain an accompaniment signal after removing the vocal in real time. Summary of the Invention

[0005] In order to overcome the problem that traditional AI cancellation technology cannot obtain a plug-and-play music accompaniment in real time.

[0006] The technical solution of the present invention is: an AI original sound cancellation method and device, including the following steps:

[0007] S1, receiving an audio signal from a Bluetooth audio, an external input, a mobile phone OTG audio or a PC audio;

[0008] S2, using an AI model to analyze the audio signal, identify and separate the vocal and accompaniment parts;

[0009] S3, processing the audio signal in real time through an audio processor to remove the vocal part;

[0010] S4, outputting an accompaniment signal after removing the vocal.

[0011] Preferably, the AI model is a voice separation model based on deep learning, which can learn the characteristics of the vocal and accompaniment through training data and apply them in real-time audio processing.

[0012] Preferably, the audio processor converts the audio signal from the time domain to the frequency domain through real-time Fourier transform, and uses the AI model to separate the vocal and accompaniment in the frequency domain.

[0013] Preferably, the method supports multiple audio input sources, including Bluetooth devices, external inputs, mobile phone OTG audio, and PC audio, and can be connected through devices such as external devices, Bluetooth devices, and smartphones.

[0014] Preferably, the method further includes preprocessing the audio signal, including noise reduction, equalization, and compression, to improve the separation effect of the AI model.

[0015] Preferably, the method can process the audio signal in real time with a delay time of less than 50 milliseconds, ensuring that the user can hear the accompanied sound after eliminating the human voice in real time when singing or playing an instrument.

[0016] Preferably, the method further includes a user interface that allows the user to adjust the intensity of eliminating the human voice, the volume of the accompaniment, and other audio parameters.

[0017] Preferably, the method supports multi-channel audio processing, can process multiple audio input sources simultaneously, and outputs multiple accompanied sound signals after eliminating the human voice.

[0018] Preferably, the method further includes post-processing the accompanied sound signal after eliminating the human voice, including reverberation, echo, and sound effect enhancement, to improve the sound quality of the accompaniment.

[0019] Preferably, the AI original sound elimination device includes:

[0020] Audio input module: used to receive audio signals from Bluetooth audio, external input, mobile phone OTG audio, or PC audio;

[0021] AI processing module: used to analyze the audio signal using an AI model to identify and separate the human voice and the accompanied sound part;

[0022] Audio processor: includes a DSP processing module, an audio reconstruction module, and a post-processing module, used to process the audio signal in real time, remove the human voice part. The DSP processing module uses the parallel computing ability of the digital signal processor to process the audio signal efficiently and in real time. The audio reconstruction module restores the processed frequency-domain signal to the time-domain signal and ensures the naturalness and integrity of the signal. The post-processing module further optimizes the reconstructed audio signal to improve the sound quality and listening experience;

[0023] Audio output module: used to output the accompanied sound signal after eliminating the human voice.

[0024] Advantages of the present invention:

[0025] By receiving audio signals from Bluetooth audio, external input, mobile phone OTG audio, or PC audio, and then using an AI model to analyze the audio signals, identify and separate the vocal and accompaniment parts, and then using an audio processor to process the audio signals in real time to remove the vocal part, and finally output the accompaniment signal after removing the vocal part, so as to obtain an accompaniment signal without vocals in real time, thus solving the problem of not being able to obtain a plug-and-play music accompaniment in real time. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 FIG. shows a schematic diagram of the system architecture of an AI original sound cancellation method and device of the present invention;

[0027] Figure 2 FIG. shows a schematic diagram of the training process of the AI model of an AI original sound cancellation method and device of the present invention;

[0028] Figure 3 FIG. shows a schematic diagram of the real-time audio processing process of an AI original sound cancellation method and device of the present invention;

[0029] Figure 4 FIG. shows a flowchart of the operation of an AI original sound cancellation method and device of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0030] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0031] The technical solution of the present invention is: an AI original sound cancellation method and device, including the following steps:

[0032] S1, receiving audio signals from Bluetooth audio, external input, mobile phone OTG audio, or PC audio;

[0033] S2, using an AI model to analyze the audio signals, identify and separate the vocal and accompaniment parts;

[0034] S3, using an audio processor to process the audio signals in real time to remove the vocal part;

[0035] S4, outputting the accompaniment signal after removing the vocal part.

[0036] The AI model is a voice separation model based on deep learning, which can learn the characteristics of vocals and accompaniments through training data and be applied in real-time audio processing.

[0037] The audio processor converts the audio signal from the time domain to the frequency domain through real-time Fourier transform, and uses an AI model to separate the vocals and accompaniment in the frequency domain.

[0038] The method supports multiple audio input sources, including Bluetooth devices, external inputs, mobile phone OTG audio, and PC audio, and can be connected through devices such as external devices, Bluetooth devices, and smartphones.

[0039] The method also includes preprocessing the audio signal, including noise reduction, equalization, and compression, to improve the separation effect of the AI model.

[0040] The method can process audio signals in real time, with a delay time of less than 50 milliseconds, ensuring that users can hear the accompaniment without vocals in real time when singing or playing an instrument.

[0041] The method also includes a user interface that allows users to adjust the intensity of vocal elimination, the volume of the accompaniment, and other audio parameters.

[0042] The method supports multi-channel audio processing, can process multiple audio input sources simultaneously, and outputs multiple accompaniment signals without vocals.

[0043] The method also includes post-processing the accompaniment signal without vocals, including reverberation, echo, and sound effect enhancement, to improve the sound quality of the accompaniment.

[0044] The AI original sound elimination device includes:

[0045] An audio input module: used to receive audio signals from Bluetooth audio, external input, mobile phone OTG audio, or PC audio;

[0046] An AI processing module: used to analyze the audio signal using an AI model, identify and separate the vocal and accompaniment parts;

[0047] An audio processor: includes a DSP processing module, an audio reconstruction module, and a post-processing module, used to process audio signals in real time, remove the vocal part. The DSP processing module uses the parallel computing power of the digital signal processor to process the audio signal efficiently and in real time. The audio reconstruction module restores the processed frequency domain signal to the time domain signal and ensures the naturalness and integrity of the signal. The post-processing module further optimizes the reconstructed audio signal to improve the sound quality and listening experience;

[0048] An audio output module: used to output the accompaniment signal without vocals.

[0049] The technical solution of the present invention mainly includes the following steps:

[0050] The device receives audio signals from Bluetooth audio, external input (AUX), mobile phone OTG audio, or PC audio through the audio input module. The audio signals can be mono or stereo, and the system supports multi-channel audio input.

[0051] The device uses an AI model to analyze the audio signals, identify and separate the vocal and accompaniment parts. The AI model is a deep learning-based voice separation model that can learn the characteristics of vocals and accompaniments from training data and apply them in the process of real-time audio processing. The training data of the AI model includes a large number of audio samples with vocals and accompaniments. By learning these samples, the AI model can achieve the effect of accurately identifying and separating vocals and accompaniments.

[0052] The device processes the audio signals in real time through an audio processor (DSP) to remove the vocal part. The audio processor converts the audio signals from the time domain to the frequency domain through real-time Fourier transform (FFT), uses the AI model to separate the vocals and accompaniments in the frequency domain, and the processed audio signals are converted back to the time domain through inverse Fourier transform (IFFT) to output the accompaniment signal with vocals removed.

[0053] The device outputs the accompaniment signal with vocals removed through the audio output module. The output signal can be connected to external devices, Bluetooth devices, smartphones, etc. Users can hear the accompaniment with vocals removed in real time through devices such as headphones and speakers.

[0054] The device also provides a user interface that allows users to adjust the intensity of vocal removal, the volume of the accompaniment, and other audio parameters. Users can adjust the parameters according to their needs to obtain the best accompaniment effect.

[0055] Please refer to Figure 1 , the system architecture of the present invention includes an audio input module, an AI processing module, an audio processor (DSP), and an audio output module. The audio input module receives audio signals from Bluetooth audio, external input (AUX), mobile phone OTG audio, or PC audio. The AI processing module uses an AI model to analyze the audio signals and simultaneously identify and separate the vocal and accompaniment parts. At the same time, the audio processor (DSP) processes the audio signals in real time to remove the vocal part. Finally, the audio output module outputs the accompaniment signal with vocals removed.

[0056] Please refer to Figure 2, the training process of the AI model includes data collection, data preprocessing, model training, model deployment, model optimization, and model evaluation. In the data collection stage, a large number of audio samples with human voices and accompaniments are collected; in the data preprocessing stage, first, an adaptive filter is used to suppress environmental noise, then the audio stream is cut into short-time frames with a frame length of 20 - 40 ms, a Hamming window is applied to each frame to reduce spectral leakage, and then the fast Fourier transform (FFT) is performed on each frame of audio to generate the amplitude spectrum and phase spectrum. Finally, the amplitude spectrum is normalized to compress the numerical range to [0, 1]; in the model training stage, the U-Net architecture is adopted. Its encoder-decoder structure is suitable for the separation task of frequency-domain signals. The encoder part is used to extract features, and the decoder part is used to generate the vocal mask and accompaniment mask. In the design of the loss function, the mean square error (MSE) is used to measure the difference between the predicted mask and the true mask, and the perceptual loss is combined to ensure that the separated audio has a natural sense in terms of hearing. When training the AI model, the dataset is divided into a training set (80%), a validation set (10%), and a test set (10%). The Adam optimizer is used, and the initial learning rate is set to 0.001 and gradually decayed. The loss of the validation set is monitored during the training process to prevent overfitting; in the model deployment stage, the trained model is deployed to a DSP or an embedded device; in the model optimization stage, the network depth or width is increased to improve the feature extraction ability, and the attention mechanism is introduced to enhance the model's attention to important frequency bands. At the same time, more diverse training data is added, such as music in different languages and styles, and the noise and reverberation in real scenarios are simulated to improve the robustness of the model. At the same time, multi-task learning is combined to improve the performance of the model; in the model evaluation stage, the separation effect of the model is evaluated by using the test set, and the objective metrics SI-SNR and PESQ are calculated. Among them, SI-SNR is used to measure the signal-to-noise ratio of the separated signal, and PESQ is used to evaluate the auditory quality of the audio. The spectrograms before and after separation are compared to check for spectral holes or distortions.

[0057] Please refer to Figure 3, The real-time audio processing process includes audio signal reception, Fourier transform (FFT), AI model separation, inverse Fourier transform (IFFT), and audio signal output. The audio input stage device receives the audio signal through the audio input module; in the audio preprocessing stage, an adaptive filter is used to suppress environmental noise, and then the audio stream is cut into short-time frames with a frame length of 20 - 40 ms. A Hamming window is applied to each frame to reduce spectral leakage, and dynamic range control (DRC) is used to balance volume fluctuations and prevent overload or distortion; in the AI model inference stage, a fast Fourier transform (FFT) is performed on each frame of audio to generate an amplitude spectrum and a phase spectrum. Then, the amplitude spectrum is input into a pre-trained U-Net model to output a vocal mask and an accompaniment mask. Finally, the accompaniment mask is multiplied point by point with the original amplitude spectrum to suppress the vocal component and obtain the processed amplitude spectrum; in the audio reconstruction stage, an inverse FFT (IFFT) is performed on the processed amplitude spectrum combined with the original phase spectrum to restore it to a time-domain signal. Then, 50% frame overlap is used to reduce boundary distortion, and weighted superposition is performed on the overlapping part, such as Hamming window superposition, to ensure smooth signal transition. Finally, the processed short-time frames are spliced in chronological order to generate a complete time-domain audio stream; in the post-processing stage, a harmonic synthesis algorithm is used to fill in spectral holes, such as filling in missing mid-frequency instrument tones through a generative adversarial network (GAN). Then, dynamic range compression (DRC) is performed on the audio signal to prevent excessive volume fluctuations, and a limiter is used to prevent signal overload and ensure that the output signal is within a safe range; in the audio output stage, the processed accompaniment signal is output to the target device.

[0058] Please refer to Figure 4 , The specific working process of the present invention is as follows:

[0059] S1, Receive audio signals from multiple input sources, and then resample the input signals with different sampling rates to a unified fixed sampling rate. Among them, Bluetooth audio is transmitted using the A2DP protocol, AUX / OTG is input through analog or digital signals, and the PC side is transmitted through USB or HDMI interfaces;

[0060] S2, Use an adaptive filter to suppress environmental noise, and then cut the audio stream into short-time frames of 20 - 40 ms / frame. A Hamming window is applied to each frame to reduce spectral leakage, and at the same time, dynamic range control is used to balance volume fluctuations to prevent overload;

[0061] S3, Perform a fast Fourier transform (FFT) on each frame of audio to generate an amplitude spectrum and a phase spectrum. Input the amplitude spectrum into a pre-trained U-Net model to output a vocal mask and an accompaniment mask. Multiply the accompaniment mask point by point with the original amplitude spectrum to suppress the vocal component. Finally, combine the processed amplitude spectrum with the original phase spectrum and perform an inverse FFT (IFFT) to restore it to a time-domain signal;

[0062] S4. Use the DSP multi-core architecture to execute steps such as preprocessing, FFT, AI inference, and IFFT in parallel. At the same time, adopt 50% frame overlap to reduce boundary distortion, and control the end-to-end delay within 20 ms to meet the real-time singing requirements.

[0063] S5. Fill the spectral holes through harmonic synthesis and simulate spatial effects such as concert halls and recording studios to enhance the stereo field.

[0064] S6. Support wired (3.5 mm, Line Out) and wireless (Bluetooth, Wi-Fi) outputs. For the high latency of Bluetooth, adopt dynamic buffer adjustment to achieve audio-visual synchronization, process multi-channel audio, separate the vocals in the center channel, and output the processed sound.

[0065] For user interaction, information such as real-time spectrograms, vocal cancellation intensity indicator bars, and volume levels can be displayed on the mobile phone App or the device screen, providing an intuitive operation interface and real-time feedback, thus enhancing the user experience.

[0066] For Bluetooth audio input:

[0067] Users transmit the music in the mobile phone to the device through a Bluetooth device. The device receives the Bluetooth audio signal through the audio input module, analyzes the audio signal using an AI model, then identifies and separates the vocal and accompaniment parts, and then processes the audio signal in real time through an audio processor (DSP). Finally, the vocal part is removed and the accompaniment signal after vocal cancellation is output. Users can hear the accompaniment after vocal cancellation in real time through headphones.

[0068] For external input (AUX):

[0069] Users transmit the music in the music player to the system through external input (AUX). The system receives the external input audio signal through the audio input module, analyzes the audio signal using an AI model, then identifies and separates the vocal and accompaniment parts, and then processes the audio signal in real time through an audio processor (DSP). Finally, the vocal part is removed and the accompaniment signal after vocal cancellation is output. Users can hear the accompaniment after vocal cancellation in real time through speakers.

[0070] For mobile phone OTG audio:

[0071] Users transmit the music in the mobile phone to the system through the mobile phone OTG interface. The system receives the mobile phone OTG audio signal through the audio input module, analyzes the audio signal using an AI model, then identifies and separates the vocal and accompaniment parts, and then processes the audio signal in real time through an audio processor (DSP). Finally, the vocal part is removed and the accompaniment signal after vocal cancellation is output. Users can hear the accompaniment after vocal cancellation in real time through Bluetooth headphones.

[0072] For PC audio:

[0073] The user transfers the music in the computer to the system through the PC. The system receives the PC audio signal through the audio input module, analyzes the audio signal using the AI model, then identifies and separates the vocal and accompaniment parts, and then processes the audio signal in real time through the audio processor (DSP). Finally, the vocal part is removed and the accompaniment signal after removing the vocal is output. The user can hear the accompaniment after removing the vocal in real time through the external speaker.

[0074] In a real-time audio processing system, the sampling rates of the input audio signals may vary. To ensure the compatibility and consistency of subsequent processing, it is necessary to unify the audio signals with different sampling rates to a fixed sampling rate. The following are the specific steps of resampling:

[0075] S1. Determine the input and output sampling rates. The input sampling rate fin is read from the audio signal header file or metadata, and the output sampling rate fout is fixed at 48 kHz;

[0076] S2. The resampling ratio R = fout / fin. If R > 1 (such as 44.1 kHz → 48 kHz), upsampling is required. If R < 1 (such as 96 kHz → 48 kHz), downsampling is required;

[0077] S3. When upsampling, before interpolation, it is necessary to design a low-pass filter to remove the high-frequency mirror components introduced by interpolation. The cut-off frequency of the filter is fin / 2. When downsampling, before decimation, it is necessary to design a low-pass filter to remove the frequency components higher than fout / 2 to prevent aliasing distortion. The cut-off frequency of the filter is fout / 2;

[0078] S4. Insert L - 1 zero values between every two original sampling points, where L is the upsampling factor, L = fout / fin. Perform low-pass filtering on the signal after inserting zeros to remove the high-frequency mirror components. The filtered signal is the upsampled audio signal. For example: input sampling rate: 44.1 kHz, output sampling rate: 48 kHz, upsampling factor L = 48000 / 44100 ≈ 1.088. Insert L - 1 zero values between every two sampling points, and then perform low-pass filtering;

[0079] S5. Perform low-pass filtering on the original signal to remove the frequency components higher than fout / 2. Decimate one point from the filtered signal every M points, where M is the downsampling factor, M = fin / fout. The decimated signal is the downsampled audio signal. For example: input sampling rate: 96 kHz, output sampling rate: 48 kHz, downsampling factor M = 96000 / 48000 = 2. Perform low-pass filtering on the original signal, and then decimate one point every 2 points.

[0080] The steps of the Fast Fourier Transform (FFT) are as follows:

[0081] S1. Cut the audio signal into short-time frames, with each frame having a length of N sampling points. Usually, the frame length is 20 - 40 ms. For example, at a sampling rate of 48 kHz, 1024 points correspond to approximately 21.3 ms;

[0082] S1. Apply a window function to each frame of the signal to reduce spectral leakage. The window function formula is:

[0083] n = 0, 1,......, N - 1

[0084] The windowed signal: x ω (n) = x(n)·ω(n).

[0085] S2. Perform FFT on the windowed signal x ω (n) to obtain the frequency-domain representation X(k):

[0086] k = 0, 1,......, N - 1

[0087] X(k) is a complex number, containing amplitude and phase information.

[0088] Amplitude spectrum:

[0089] Phase spectrum:

[0090] Normalize the amplitude spectrum to compress the numerical range to [0, 1]:

[0091]

[0092] The phase spectrum φ(k) remains unchanged for subsequent IFFT reconstruction.

[0093] The steps of the Inverse FFT (IFFT) are as follows:

[0094] S1. Input the processed amplitude spectrum |Y(k)| and the original phase spectrum φ(k);

[0095] S2. Convert the amplitude spectrum and the phase spectrum into complex form:

[0096] Y(k) = |Y(k)|·e jφ(k)

[0097] where, e jφ(k) = cos(φ(k)) + j·sin(φ(k)).

[0098] The output of the IFFT is complex, but the actual audio signal is real. Therefore, the real part is taken:

[0099] y real (n) = Re(y(n))

[0100] S3. When framing, 50% overlap (such as 512-point overlap) is adopted. Therefore, the frames output by the IFFT need to be overlapped and added. The window function is applied to each frame signal again, and the overlapping parts are weighted and superimposed to ensure smooth signal transition:

[0101] y final (n) = y prev (n) · ω(n) + y current (n) · ω(n)

[0102] where y final (n) is the overlapping part of the previous frame, and y current (n) is the signal of the current frame.

[0103] The present invention provides a real-time original sound cancellation method and system based on AI technology, which can process audio signals in real time through an audio processor (DSP) to remove the human voice part, and finally output the accompanied sound signal after removing the human voice. This method supports multiple audio input sources, including Bluetooth devices, external inputs (AUX), mobile phone OTG audio, and PC audio, and can be connected through external devices, Bluetooth devices, smart phones, etc. to obtain the accompanied sound signal after removing the human voice in real time.

Claims

1. An AI original sound elimination method, characterized in that: The following steps are included: S1, receives audio signals from Bluetooth audio, external input, mobile phone OTG audio or PC audio; S2 uses AI models to analyze audio signals, identify and separate vocals and accompaniment. S3, processes the audio signal in real time through the audio processor to remove the vocal part; S4, outputs the accompaniment signal after eliminating the human voice.

2. The AI ​​sound removal method according to claim 1, characterized in that: The AI ​​model is a speech separation model based on deep learning, which can learn the characteristics of human voice and accompaniment through training data and apply it in real-time audio processing.

3. The AI ​​sound removal method according to claim 1, characterized in that: The audio processor converts the audio signal from the time domain to the frequency domain through real-time Fourier transform, and uses the AI ​​model to separate the human voice and accompaniment in the frequency domain.

4. The AI ​​sound removal method according to claim 1, characterized in that: The method supports multiple audio input sources, including Bluetooth devices, external input, mobile phone OTG audio and PC audio, and can be connected through external devices, Bluetooth devices, smart phones and other devices.

5. The AI ​​sound removal method according to claim 1, characterized in that: The method also includes preprocessing the audio signal, including noise reduction, equalization and compression, to improve the separation effect of the AI ​​model.

6. The AI ​​sound removal method according to claim 1, characterized in that: The method can process audio signals in real time with a delay time of less than 50 milliseconds, ensuring that the user can hear the accompaniment after the human voice is eliminated in real time when singing or playing.

7. The AI ​​sound removal method according to claim 1, characterized in that: The method also includes a user interface that allows the user to adjust the intensity of vocal removal, the volume of the accompaniment, and other audio parameters.

8. The AI ​​sound removal method according to claim 1, characterized in that: The method supports multi-channel audio processing, can process multiple audio input sources simultaneously, and output multiple accompaniment signals after eliminating human voices.

9. The AI ​​sound removal method according to claim 1, characterized in that: The method also includes post-processing the accompaniment signal after eliminating the human voice, including reverberation, echo and sound effect enhancement, so as to improve the sound quality of the accompaniment.

10. The AI ​​sound canceller device according to claim 1, characterized in that: include: Audio input module: used to receive audio signals from Bluetooth audio, external input, mobile phone OTG audio or PC audio; AI processing module: used to analyze audio signals using AI models, identify and separate vocals and accompaniment; Audio processor: It includes DSP processing module, audio reconstruction module and post-processing module, which are used to process audio signals in real time and remove the vocal part. The DSP processing module uses the parallel computing capability of the digital signal processor to process the audio signals efficiently and in real time. The audio reconstruction module restores the processed frequency domain signal to the time domain signal and ensures the naturalness and integrity of the signal. The post-processing module further optimizes the reconstructed audio signal to improve the sound quality and listening experience. Audio output module: used to output the accompaniment signal after eliminating the human voice.