Voiceprint-driven voice noise reduction method and terminal equipment
By collecting user voiceprint features to generate voiceprint packets, using AI algorithms to separate and compare noise signals, and combining feature completion algorithms to optimize speech integrity, the problem of speech noise reduction in complex environments for terminal devices has been solved, achieving efficient and low-power multi-scenario adaptation.
Patent Information
- Application Number
- CN202512058255.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-31
- Publication Date
- 2026-03-17
AI Technical Summary
Existing terminal devices have poor voice noise reduction technology, which cannot effectively filter noise in complex environments and is prone to damaging valid voice, resulting in decreased communication clarity, weak anti-interference ability, and difficulty in adapting to diverse usage scenarios.
By collecting user voice samples, extracting unique voiceprint features to generate voiceprint packages, using AI algorithms to analyze input audio in real time, separating and comparing voiceprint data, filtering noise signals, combining feature completion algorithms to optimize voice integrity, and adapting to the terminal's built-in components to achieve noise reduction.
It achieves precise filtering of all noise, preserves the essence of voice, adapts to diverse scenarios, reduces power consumption, requires no additional hardware, and improves communication accuracy.
Smart Images

Figure CN121687091A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech processing technology, specifically to a voiceprint-driven speech noise reduction method and terminal device. Background Technology
[0002] Existing voice noise reduction technologies for terminal devices (mobile phones, walkie-talkies) mostly employ methods such as spectrum filtering, volume threshold screening, and fixed frequency band suppression. Their core logic is to passively identify and weaken noise signals. These technologies have significant drawbacks: First, they lack specificity, only addressing fixed frequencies and single types of noise, failing to effectively filter environmental interference at the same frequency as the voice (such as conversations or background noise in complex environments). Second, they easily damage valid voice data; while suppressing noise frequencies, they simultaneously weaken the corresponding frequency band details of the user's voice, leading to voice distortion and affecting communication clarity. Third, they have weak anti-interference capabilities; when the intensity of environmental noise approaches or exceeds the intensity of the user's voice, the noise reduction effect significantly diminishes, sometimes even failing to distinguish between voice and noise, thus failing to meet the voice communication needs in complex and noisy environments.
[0003] Meanwhile, mobile phones and walkie-talkies are used in a variety of scenarios. Mobile phones often face outdoor crowds and traffic noise, while walkie-talkies are mostly used in strong interference scenarios such as construction sites and workshops. Existing noise reduction technologies are difficult to adapt to the noise reduction needs of all scenarios, and some noise reduction solutions rely on cloud computing power or additional hardware, which increases terminal power consumption and cost, thus limiting adaptability. Summary of the Invention
[0004] In view of this, the present invention provides a voiceprint-driven speech noise reduction method and terminal device to solve the above problems.
[0005] To address the above technical problems, this invention provides a voiceprint-driven speech denoising method, comprising: Collect voice samples from terminal users, extract unique voiceprint features determined by the user's physiological vocal organs, generate a unique voiceprint package for the user, and store it locally on the terminal. The terminal audio acquisition module collects the user's input audio during voice communication in real time. The input audio includes the user's valid voice and environmental noise. The input audio is analyzed in real time using AI algorithms to separate the voiceprint feature information in the audio and obtain the input voiceprint data. The input voiceprint data is compared with the user-specific voiceprint package pre-stored on the terminal. A similarity threshold is set, and audio signals with similarity reaching the threshold are retained, while noise signals with similarity not reaching the threshold are filtered out. Extract and preserve syllable parameters and rhythm information from the audio signal, optimize speech integrity through feature completion algorithm, and output the result.
[0006] As an optional approach, voiceprint features include the fundamental frequency, formants, and intensity variation patterns of the user's speech. Furthermore, the voiceprint package supports dynamic updates, expanding the voiceprint variant library by collecting speech samples from users in different scenarios.
[0007] As an optional approach, the AI algorithm includes a voiceprint separation module and a feature recognition module. The voiceprint separation module separates the voiceprint signal and noise signal superimposed in the input audio, while the feature recognition module extracts the core parameters of the voiceprint and resists environmental interference.
[0008] As an optional approach, syllable parameters include the pause duration, stress intensity, and tone of phonemes and / or syllables in the audio. The feature completion algorithm completes the fragmented speech signal covered by noise based on the user's voiceprint features and the rules of syllable parameters.
[0009] As an optional approach, the similarity threshold is adaptively adjusted, and the terminal dynamically optimizes the matching accuracy based on the intensity of ambient noise, balancing noise reduction effect and voice integrity.
[0010] On the other hand, the present invention also provides a voiceprint-driven speech noise reduction terminal device for implementing the above-mentioned voiceprint-driven speech noise reduction method, comprising: The storage module is used to collect voice samples from terminal users, extract unique voiceprint features determined by the user's physiological vocal organs, generate a unique voiceprint package for the user, and store it locally on the terminal. The acquisition module is used to acquire the user's input audio during voice communication in real time through the terminal audio acquisition module. The input audio includes the user's valid voice and environmental noise. The processing module is used to analyze the input audio in real time using AI algorithms, separate the voiceprint feature information in the audio, and obtain the input voiceprint data. The output module is used to compare the input voiceprint data with the user-specific voiceprint package pre-stored in the terminal, set a similarity threshold, retain audio signals with similarity reaching the threshold, and filter noise signals with similarity not reaching the threshold. The update module is used to extract syllable parameters and combine rhythm information from the preserved audio signal, optimize speech integrity through feature completion algorithm, and output high-fidelity effective speech.
[0011] As an alternative approach, the terminal device includes a mobile phone or a walkie-talkie, and the processing module is a SoC module.
[0012] As an optional approach, the acquisition module is a terminal microphone.
[0013] The beneficial effects of this invention are as follows: This invention uses the user's unique voiceprint as the core filtering criterion, and is not limited by noise type, frequency, or intensity. It can filter all interference signals without a target voiceprint, adapt to the complex usage scenarios of mobile phones and walkie-talkies, and has no noise reduction blind spots.
[0014] It filters out only irrelevant noise without altering the user's voiceprint features and syllable parameters. Combined with feature completion algorithms, it fully preserves the essence of speech, avoiding speech distortion caused by traditional noise reduction and improving communication accuracy. No additional hardware is required; it can be implemented using the terminal's built-in storage, computing power, and audio components. Mobile phones utilize the powerful computing capabilities of their SoCs for efficient operation, while walkie-talkies are adapted to lightweight algorithms, balancing low power consumption and low latency, resulting in broad compatibility. Attached Figure Description
[0015] Figure 1 This is a schematic diagram of the voiceprint-driven speech noise reduction method of the present invention. Detailed Implementation
[0016] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the various embodiments of the present invention will be described in detail below with reference to the accompanying drawings. However, those skilled in the art will understand that many technical details have been presented in the various embodiments of the present invention to enable the reader to better understand the present invention. However, the technical solutions claimed in the present invention can be implemented even without these technical details and various changes and modifications based on the following embodiments.
[0017] Please see Figure 1 This embodiment provides a voiceprint-driven speech denoising method, which uses a user's unique voiceprint as the recognition basis and completes voiceprint matching and denoising through local computing power. Its implementation process includes: The method involves collecting voice samples from terminal users, extracting unique voiceprint features determined by the user's physiological vocal organs, generating a user-specific voiceprint package, and storing it locally on the terminal. Next, the terminal's audio acquisition module collects the user's input audio during voice communication in real time. This input audio includes the user's valid speech and environmental noise. Then, an AI algorithm analyzes the input audio in real time, separating the voiceprint feature information to obtain input voiceprint data. The input voiceprint data is compared with the user-specific voiceprint package pre-stored on the terminal. A similarity threshold is set, retaining audio signals with similarity reaching the threshold and filtering out noise signals with similarity below the threshold. Finally, syllable parameters and rhythmic information are extracted from the retained audio signals, and a feature completion algorithm is used to optimize speech integrity and output high-fidelity valid speech. This method is applicable to terminal devices such as mobile phones and walkie-talkies, relying on the terminal's built-in components to achieve noise reduction. The AI algorithm includes, but is not limited to, extracting time-domain or frequency-domain features from the input audio and performing matching analysis based on voiceprint-related feature vectors or embedding representations to distinguish target speaker speech components from non-target noise components. Furthermore, those skilled in the art will understand that the voiceprint separation process can be achieved through correlation matching, feature similarity calculation, or voiceprint embedding conditionalization, and the present invention does not limit this.
[0018] As an optional approach, voiceprint features include the fundamental frequency, formants, and intensity variation patterns of the user's speech. The voiceprint package supports dynamic updates, expanding the voiceprint variant library by collecting speech samples from users in different scenarios. This dynamic update mechanism allows users to record samples at normal speaking speeds, while shouting, or during rapid conversations, thereby constructing a voiceprint variant library and improving adaptability to different speech states. The AI algorithm includes a voiceprint separation module and a feature recognition module. The voiceprint separation module separates the superimposed voiceprint signal from the noise signal in the input audio, while the feature recognition module extracts the core voiceprint parameters and possesses anti-environment interference capabilities. Through the above processing, this embodiment can resist the interference of environmental noise on voiceprint edge features, ensuring accurate extraction of input voiceprint data in complex and noisy environments.
[0019] Furthermore, syllable parameters include the pause duration, stress intensity, and tone of phonemes and syllables in the audio. The feature completion algorithm, based on the user's voiceprint features and syllable parameter patterns, completes the fragmented speech signal covered by noise. This completion process optimizes speech coherence and clarity, filtering only unrelated noise without altering the user's voiceprint features and syllable parameters, thus avoiding speech distortion caused by traditional noise reduction. Simultaneously, the similarity threshold is adaptively adjusted, with the terminal dynamically optimizing matching accuracy based on environmental noise intensity to balance noise reduction effect and speech integrity. The threshold is increased when noise is strong to ensure noise reduction accuracy, while it is appropriately reduced when noise is weak to ensure speech integrity, thereby adapting to diverse scenarios such as outdoor crowds, traffic noise, or construction equipment noise.
[0020] On the other hand, this embodiment also provides a voiceprint-driven voice noise reduction terminal device for implementing the above method, including a storage module, a collection module, a processing module, an output module, and an update module. The storage module is used to collect voice samples from terminal users, extract unique voiceprint features determined by the user's physiological vocal organs, generate a user-specific voiceprint package, and store it locally on the terminal, while supporting high-speed local access to reduce latency and power consumption. The collection module is used to collect the user's input audio during voice communication in real time through the terminal's audio collection module. This input audio includes the user's valid voice and environmental noise. The processing module is used to analyze the input audio in real time using AI algorithms, separate the voiceprint feature information in the audio, and obtain the input voiceprint data. This module integrates a voiceprint separation module, a feature recognition module, a matching module, and a completion module based on the terminal's computing power to complete the entire process. The output module is used to compare the similarity of the input voiceprint data with the user-specific voiceprint package pre-stored on the terminal, set a similarity threshold, retain audio signals with similarity reaching the threshold, filter noise signals with similarity below the threshold, and output a high-fidelity voice signal optimized for noise reduction through audio output components such as speakers or earpieces. The update module is used to extract syllable parameters and combined rhythm information from the preserved audio signal, optimize speech integrity through feature completion algorithm, output high-fidelity effective speech, and support users to manually enter new speech samples, dynamically update voiceprint packets and algorithm parameters, and improve noise reduction adaptability.
[0021] In one optional implementation, the terminal device in this embodiment can be a mobile phone or a walkie-talkie. The processing module is a SoC module, utilizing the terminal's built-in SoC computing power, eliminating the need for additional hardware components and achieving lightweight deployment. The mobile phone operates efficiently using its powerful SoC computing power, while the walkie-talkie is equipped with a lightweight algorithm, balancing low power consumption and low latency. The acquisition module is the terminal microphone, and the algorithm optimizes the audio signal acquired by the microphone, enhancing the accuracy of voiceprint feature capture and improving the recognition of voiceprint signals in the original audio. This device does not rely on cloud computing power or additional hardware, and is suitable for the voice communication needs of mobile phones in noisy outdoor environments and walkie-talkies in highly interference-prone environments such as construction sites and workshops.
[0022] For example, in mobile applications, users record clean voice samples in different states using the phone's voice capture function. The phone's processing module extracts voiceprint features, generates a unique voiceprint package and variant library, and stores it locally. When making voice calls in noisy outdoor traffic environments, the phone's microphone captures input audio containing both voice and traffic noise. The processing module runs AI algorithms to separate and extract the input voiceprint, compares it with the pre-stored voiceprint package, and filters out traffic noise and interference from other people's voices. The algorithm completes the voice fragments slightly covered by noise, optimizes voice clarity, and outputs high-fidelity call voice through the earpiece to ensure clear communication.
[0023] In walkie-talkie applications, users record clean voice samples from their work environment, generating a unique voiceprint package stored locally to adapt to noisy construction site environments. When sending work instructions amidst the noise of construction equipment, the walkie-talkie captures audio containing both the instruction voice and equipment noise. The processing module uses a lightweight AI algorithm to quickly match the voiceprint, filter out equipment noise, and retain the instruction voice. The algorithm completes the fragmented features in the instruction voice, outputting clear instructions to ensure efficient team reception and avoid misreading of instructions due to noise.
[0024] The embodiments of the present invention have been described in detail above. The various embodiments in the specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section. It should be noted that those skilled in the art can make several improvements and modifications to the present invention without departing from the principles of the invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.
Claims
1. A voiceprint-driven speech noise reduction method, characterized in that, The method comprises the following steps: Collecting terminal user voice samples, extracting user physiological pronunciation organ determined exclusive voiceprint features, generating user exclusive voiceprint package, and storing in terminal local; Through the terminal audio acquisition module, real-time acquisition of user voice communication process input audio, the input audio contains user effective voice and environmental noise; Through AI algorithm real-time analysis of input audio, separate voiceprint feature information in audio, get input voiceprint data; Compare the input voiceprint data with the pre-stored user exclusive voiceprint package of the terminal, set the similarity threshold, retain the audio signal with similarity reaching the threshold, filter the noise signal with similarity not reaching the threshold; Extract syllable parameters and / or combined rhythm information in the retained audio signal, and output after optimizing the voice integrity through feature completion algorithm; or, directly output the voice signal filtered by voiceprint.
2. The voiceprint-driven speech noise reduction method of claim 1, wherein, The voiceprint features include the fundamental frequency, formant, and intensity variation of user voice, and the voiceprint package supports dynamic updating, which expands the voiceprint variant library by collecting user voice samples in different scenarios.
3. The voiceprint-driven speech noise reduction method of claim 1, wherein, The AI algorithm includes a voiceprint separation module and a feature recognition module, the voiceprint separation module separates the superimposed voiceprint signal and noise signal in the input audio, and the feature recognition module extracts voiceprint core parameters and resists environmental interference.
4. The voiceprint-driven speech noise reduction method of claim 1, wherein, The syllable parameters include the pause duration, stress intensity, and tone intonation of phonemes and / or syllables in the audio, and the feature completion algorithm is based on user voiceprint features and syllable parameter rules to complete the fragmented voice signal covered by noise.
5. The voiceprint-driven speech noise reduction method of claim 1, wherein, The similarity threshold is self-adaptive, the terminal dynamically optimizes the matching accuracy according to the environmental noise intensity, and balances the noise reduction effect and voice integrity.
6. A voiceprint-driven speech noise reduction terminal device for implementing the voiceprint-driven speech noise reduction method according to any one of claims 1-5, characterized in that, The method comprises the following steps: A storage module is used to collect terminal user voice samples, extract user physiological pronunciation organ determined exclusive voiceprint features, generate user exclusive voiceprint package, and store in terminal local; An acquisition module is used to collect input audio in user voice communication process through terminal audio acquisition module, the input audio contains user effective voice and environmental noise; A processing module is used to analyze input audio in real time through AI algorithm, separate voiceprint feature information in audio, and get input voiceprint data; An output module is used to compare input voiceprint data with pre-stored user exclusive voiceprint package of the terminal, set the similarity threshold, retain the audio signal with similarity reaching the threshold, and filter the noise signal with similarity not reaching the threshold; An update module is used to extract syllable parameters and combined rhythm information in the retained audio signal, optimize voice integrity through feature completion algorithm, and output high-fidelity effective voice.
7. The voiceprint-driven speech noise reduction endpointing device of claim 6, wherein, The terminal device includes a mobile phone and a talkback machine, and the processing module is a SoC module.
8. The voiceprint-driven speech noise reduction endpointing device of claim 6, wherein, The acquisition module is a terminal microphone.