Voiceprint recognition method and system for high-noise environment
Through array microphone and noise suppression technology, combined with audio mark matching, traditional voiceprint recognition has solved the problems of low recognition rate and poor adaptability in high noise environments, and achieved high accuracy and real-time adaptation of voiceprint recognition effects.
Patent Information
- Application Number
- CN202510226375.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional voiceprint recognition systems have low recognition rate in noisy environments and cannot effectively filter out environmental noise, resulting in a decrease in recognition accuracy and poor environmental adaptability under different noise conditions.
The sound information is obtained through the array microphone, the phase difference and delay are calculated to determine the target sound direction, the target sound is enhanced and the noise is suppressed, the environmental noise characteristic model is constructed, the background noise is filtered out, and the audio marks are embedded to match the voiceprint signal.
Significantly improve recognition rate and accuracy in high noise environments, enhance system security and robustness, prevent identity forgery, and achieve real-time adaptation to different noise conditions.
Smart Images

Figure CN120108400A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of voiceprint recognition, and in particular to a voiceprint recognition method and system for use in a high-noise environment. Background Art
[0002] Voiceprint recognition is a type of biometric technology that authenticates or identifies a person based on the unique characteristics of their voice. Each person's voice has specific voiceprint characteristics, such as pitch, timbre, speaking speed, and pronunciation, due to differences in the structure of the vocal tract, larynx, and mouth. Voiceprint recognition technology is widely used in security authentication, mobile payment, smart assistants, and telephone banking.
[0003] Its basic process includes sound collection, feature extraction, voiceprint library comparison and recognition judgment. With the development of deep learning and signal processing technology, the accuracy and robustness of voiceprint recognition have been significantly improved, especially in complex environments. However, this field still faces some challenges, such as environmental noise impact, user identity forgery, and adaptability to voiceprint changes. In the future, voiceprint recognition is expected to be combined with other biometric technologies to improve security and user experience.
[0004] Traditional voiceprint recognition systems often face the dilemma of low recognition rate in noisy environments and cannot accurately capture the target sound, resulting in poor recognition results. In addition, the existing technology has insufficient adaptability to background noise and cannot effectively filter out environmental noise, further reducing recognition accuracy. It also has poor environmental adaptability under different noise conditions and cannot maintain stable recognition performance, limiting its effectiveness and universal applicability in complex application scenarios. Summary of the invention
[0005] The purpose of the present invention is to provide a voiceprint recognition method for a high noise environment, aiming to solve the technical problems existing in the prior art identified in the background technology.
[0006] The present invention is implemented as follows: a voiceprint recognition method for a high noise environment, the method comprising:
[0007] Pre-capture the user's voice signal, analyze the voice features, extract the voiceprint features, combine the extracted voiceprint features with the user's identity information to form a user-specific voiceprint fingerprint, and establish a voiceprint library for storage;
[0008] In a high-noise environment, the positions of all array microphones and the sound information collected by each microphone in different directions are obtained, and the direction of the target sound is determined by calculating the phase difference and delay between the signals received by each microphone;
[0009] Enhance the identified target sound and suppress the noise from other directions. Perform spectrum analysis on the enhanced sound information, identify the characteristic frequency of the noise, and generate an environmental noise characteristic model.
[0010] According to the environmental noise feature model, the noise frequency in the current environment is identified, and the background noise is filtered out to extract the voiceprint signal of the target sound;
[0011] An audio tag is embedded in the voiceprint signal, and the voiceprint signal embedded with the audio tag is matched with the voiceprint fingerprint in the voiceprint library for similarity to determine whether the target voice is the user's voice.
[0012] As a further solution of the present invention, the method of analyzing the sound features, extracting the voiceprint features, combining the extracted voiceprint features with the user's identity information to form a user-specific voiceprint fingerprint, and establishing a voiceprint library for storage specifically includes:
[0013] Pre-recording the user's voice information, and pre-processing the recorded voice information, dividing the pre-processed voice information into a number of short time frame audio segments;
[0014] Simulate the auditory characteristics of the human ear, obtain voiceprint features, and combine the extracted voiceprint features with the user's identity information to form a complete set of voiceprint fingerprint data;
[0015] Create a voiceprint library structure to store the voiceprint fingerprints of all users, including user IDs and corresponding voiceprint feature sets.
[0016] As a further solution of the present invention, the steps of obtaining the positions of all array microphones and the sound information collected by each microphone in different directions, and determining the direction of the target sound by calculating the phase difference and delay between the signals received by each microphone, specifically include:
[0017] Read the audio signals from different directions collected by each microphone and convert them into digital signals;
[0018] Perform short-time Fourier transform on the signals collected by different microphones to obtain the time-frequency characteristics of each signal;
[0019] The direction angle of the target sound is calculated based on the time delay from the target sound to each microphone and the microphone position relationship, and the signals collected by different microphones in the same time window are compared to calculate the phase difference:
[0020]
[0021] Among them, Δt irepresents the signal delay received by the i-th microphone, d represents the distance between adjacent microphones in the array, θ represents the angle between the sound source and the microphone array, and c represents the propagation speed of sound waves in the air. represents the phase difference between the i-th microphone and the reference microphone, and f represents the frequency of the sound source signal.
[0022] As a further solution of the present invention, the identified target sound is enhanced, noise from other directions is suppressed, spectrum analysis is performed on the enhanced sound information, characteristic frequencies of the noise are identified, and an environmental noise characteristic model is generated, which specifically includes:
[0023] Based on the calculated directional angle of the target sound, a directional beam is formed for the microphone array, the sensitivity of the array is concentrated in the direction of the target sound, the target sound is enhanced and interference signals from other directions are suppressed;
[0024] Perform fast Fourier transform on the enhanced target sound, convert the target sound from the time domain to the frequency domain, extract the frequency domain features, and identify the existing noise frequency components;
[0025] Based on the recognition results, the frequency components and power spectral density distribution of the noise are analyzed, and an environmental noise model is constructed according to the analysis results.
[0026] As a further solution of the present invention, the noise frequency existing in the current environment is identified according to the environmental noise feature model, and the background noise is filtered out to extract the voiceprint signal of the target sound;
[0027] The audio signal collected by the microphone in real time is divided into several short-time frame audio segments for analysis, and the spectrum characteristics of each frame are extracted;
[0028] Based on the environmental noise model, the spectrum extracted in real time is analyzed to determine the noise components in the current signal and identify the characteristic frequency of the noise;
[0029] Set a frequency threshold, filter and mark the noise characteristic frequencies that exceed the frequency threshold, subtract the marked noise spectrum from the spectrum of the current audio signal to obtain the denoised signal spectrum, and use a filter to separate the noise;
[0030] The filtered frequency domain signal is converted back to the time domain, reorganized into the target sound signal, and the voiceprint features are extracted.
[0031] As a further solution of the present invention, embedding an audio tag in a voiceprint signal, performing similarity matching between the voiceprint signal embedded with the audio tag and the voiceprint fingerprint in the voiceprint library, and determining whether the target sound is the user's voice specifically includes:
[0032] Based on the extracted voiceprint features, a preset audio tag is embedded in the corresponding voiceprint signal;
[0033] Confirm the identity information of the current user to be matched, and extract the voiceprint fingerprint related to the target user from the voiceprint database;
[0034] Compare the voiceprint signal embedded in the audio tag with the voiceprint fingerprint in the voiceprint library, calculate the similarity, and determine whether the sound is the user's voice:
[0035]
[0036] Among them, D represents the comparison similarity, x i is the feature of the embedded voiceprint signal, y i Represents the characteristics of the voiceprint fingerprint, n is the total number of characteristics;
[0037] A decision threshold is set, and the calculated similarity is compared with the preset decision threshold. If the similarity is higher than the decision threshold, the target voice is determined to be the user's voice.
[0038] Another object of the present invention is to provide a voiceprint recognition system for a high noise environment, the system comprising:
[0039] The user voiceprint feature extraction module is used to pre-capture the user's voice signal, analyze the voice features, extract the voiceprint features, combine the extracted voiceprint features with the user's identity information to form a user-specific voiceprint fingerprint, and establish a voiceprint library for storage;
[0040] The array microphone sound information acquisition module is used to obtain the positions of all array microphones and the sound information collected by each microphone in different directions in a high-noise environment, and determine the direction of the target sound by calculating the phase difference and time delay between the signals received by each microphone;
[0041] The target sound enhancement module is used to enhance the identified target sound and suppress the noise from other directions, perform spectrum analysis on the enhanced sound information, identify the characteristic frequency of the noise, and generate an environmental noise characteristic model;
[0042] The target sound extraction module is used to identify the noise frequency existing in the current environment according to the environmental noise feature model, filter out the background noise, and extract the voiceprint signal of the target sound;
[0043] The voiceprint matching judgment module is used to embed an audio tag in the voiceprint signal, perform similarity matching between the voiceprint signal embedded with the audio tag and the voiceprint fingerprint in the voiceprint library, and judge whether the target voice is the user's voice.
[0044] The beneficial effects of the present invention are:
[0045] The application of this voiceprint recognition method in a high-noise environment has shown significant comprehensive beneficial effects. By building a personalized voiceprint database, the system can combine voiceprint features with user identity information to ensure the uniqueness of each voiceprint fingerprint. This feature not only enhances the security and reliability of the system, but also effectively prevents identity forgery and deception, and improves the accuracy of recognition.
[0046] By using array microphones for multi-channel signal processing, the system can accurately separate target sounds and locate them in real time in complex acoustic environments, significantly improving the recognition rate. By enhancing and analyzing the spectrum of the target sound, the system can timely identify and model the characteristics of environmental noise, giving it real-time adaptability and ensuring that the target sound can still be accurately recognized under changing noise conditions.
[0047] When analyzing audio signals in real time, the system can accurately separate noise by setting frequency thresholds and using filters to ensure the clarity and integrity of the target sound. This process significantly reduces the interference of environmental noise on the recognition effect. In addition, by embedding audio tags, the recognizability of voiceprint signals is enhanced, and similarity calculation is used to further improve accuracy and robustness. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 A flow chart of a voiceprint recognition method for a high noise environment provided by an embodiment of the present invention;
[0049] Figure 2 A flowchart of forming a user-specific voiceprint fingerprint and establishing a voiceprint library for storage provided by an embodiment of the present invention;
[0050] Figure 3 A flowchart for calculating the phase difference and delay between the signals received by each microphone and determining the direction of the target sound provided by an embodiment of the present invention;
[0051] Figure 4 A flowchart of performing spectrum analysis on enhanced sound information and identifying characteristic frequencies of noise provided by an embodiment of the present invention;
[0052] Figure 5 A flowchart of an embodiment of the present invention for identifying the noise frequency existing in the current environment, filtering out the background noise, and extracting the voiceprint signal of the target sound;
[0053] Figure 6 A flowchart for determining whether a target sound is a user's sound provided by an embodiment of the present invention;
[0054] Figure 7 This is a structural block diagram of a voiceprint recognition system for a high noise environment provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0055] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0056] It is understood that the terms "first", "second", etc. used in this application may be used herein to describe various elements, but unless otherwise specified, these elements are not limited by these terms. These terms are only used to distinguish a first element from another element. For example, without departing from the scope of this application, a first xx script may be referred to as a second xx script, and similarly, a second xx script may be referred to as a first xx script.
[0057] Figure 1 A flow chart of a voiceprint recognition method for a high noise environment provided by an embodiment of the present invention, such as Figure 1 As shown, the method includes:
[0058] S100, pre-capturing the user's voice signal, analyzing the voice features, extracting the voiceprint features, combining the extracted voiceprint features with the user's identity information to form a user-specific voiceprint fingerprint, and establishing a voiceprint library for storage;
[0059] This step records the user's voice information through a high-quality microphone system to ensure that the recorded audio signal reflects the user's voice characteristics as realistically as possible. The recorded sound information is then preprocessed, including denoising, echo cancellation, and normalization to improve signal quality, thereby providing a clear basis for subsequent feature extraction. Next, the preprocessed sound information is divided into several short-time frame audio segments. This processing method makes the analysis of the sound more detailed, so that richer voiceprint features can be captured. These short-time audio segments are used to extract features by simulating the auditory characteristics of the human ear. The extracted voiceprint features include multi-dimensional information such as pitch, timbre, accent, and speaking speed, and then generate a voiceprint fingerprint combined with the user's identity information to form a complete set of voiceprint fingerprint data. These fingerprint data are unique and difficult to copy, and can effectively represent the voice characteristics of each user.
[0060] This step builds a personalized and accurate voiceprint database, which provides a reliable foundation for subsequent voiceprint recognition. By combining voiceprint features with user identity information, the uniqueness of each voiceprint fingerprint is ensured, which not only enhances the security and reliability of the system, but also improves the accuracy of recognition. In addition, pre-sound recording and analysis help the system adapt to the changes in the characteristics of different users and enhance the flexibility of the system. By building a high-quality voiceprint library, the system can identify user voices in a complex real environment, significantly improving the success rate of voiceprint recognition in a high-noise environment, thus laying a solid foundation for the effective execution of subsequent steps.
[0061] like Figure 2 As shown, the method of analyzing the sound features, extracting the voiceprint features, combining the extracted voiceprint features with the user's identity information to form a user-specific voiceprint fingerprint, and establishing a voiceprint library for storage specifically includes:
[0062] S110, pre-recording the user's voice information, and pre-processing the recorded voice information, dividing the pre-processed voice information into a number of short time frame audio segments;
[0063] S120, simulating the auditory characteristics of human ears to obtain voiceprint features, and combining the extracted voiceprint features with the user's identity information to form a set of complete voiceprint fingerprint data;
[0064] S130, creating a voiceprint library structure to store the voiceprint fingerprints of all users, including user IDs and corresponding voiceprint feature sets.
[0065] S200, in a high noise environment, obtaining the positions of all array microphones and the sound information collected by each microphone in different directions, and determining the direction of the target sound by calculating the phase difference and delay between the signals received by each microphone;
[0066] This step reads the audio signals collected by each microphone in different directions and converts them into digital signals. This conversion process ensures that the signal can be effectively processed and analyzed, allowing the subsequent short-time Fourier transform (STFT) to be performed. Through STFT, the time-frequency characteristics of each signal can be obtained, thus providing the necessary data basis for the subsequent sound direction calculation.
[0067] After obtaining the time-frequency features, the system calculates the direction angle of the target sound based on the time delay from the target sound to each microphone and the relationship between the microphone positions. Specifically, by using the basic physical principles of sound propagation and combining the time difference between the signals received by each microphone, the system can locate the direction of the target sound. This process is accomplished by comparing the signals from different microphones in the same time window and calculating the phase difference. The mathematical expression shown in the formula provides theoretical support for this calculation.
[0068] By using an array microphone, this method can separate the target sound in a complex sound environment, greatly improving the recognition rate compared to the traditional single microphone system. This multi-channel signal processing technology can not only effectively reduce the interference of background noise, but also locate the direction of the target sound in real time, allowing the system to focus on the user's voice more accurately. In addition, by calculating the phase difference and delay, it can achieve rapid response and adaptation to the sound source, thereby maintaining high performance stability in a dynamic environment.
[0069] like Figure 3 As shown, the acquisition of the positions of all array microphones and the sound information collected by each microphone in different directions, and the determination of the direction of the target sound by calculating the phase difference and delay between the signals received by each microphone, specifically includes:
[0070] S210, reading the audio signals from different directions collected by each microphone, and converting the audio signals into digital signals;
[0071] S220, performing short-time Fourier transform on the signals collected by different microphones to obtain time-frequency characteristics of each signal;
[0072] S230, calculating the direction angle of the target sound based on the time delay from the target sound to each microphone and the microphone position relationship, and comparing the signals collected by different microphones in the same time window to calculate the phase difference:
[0073]
[0074] Among them, Δt i represents the signal delay received by the i-th microphone, d represents the distance between adjacent microphones in the array, θ represents the angle between the sound source and the microphone array, and c represents the propagation speed of sound waves in the air. represents the phase difference between the i-th microphone and the reference microphone, and f represents the frequency of the sound source signal.
[0075] S300, enhancing the identified target sound, suppressing noise from other directions, performing spectrum analysis on the enhanced sound information, identifying characteristic frequencies of the noise, and generating an environmental noise characteristic model;
[0076] In this step, the system uses a microphone array to form a directional beam based on the pre-calculated direction angle of the target sound, which can effectively focus the sensitivity of the array in the direction of the target sound. This process not only helps to increase the ability to capture the target sound, but also effectively suppresses interference signals from other directions. Through this beamforming technology, the system can improve the signal-to-noise ratio of the target sound in a complex acoustic environment, thereby providing a higher quality signal for subsequent sound analysis.
[0077] After the target sound is enhanced, the system performs a fast Fourier transform (FFT) on it to convert the time domain signal into a frequency domain signal. This conversion process is the basis for extracting frequency domain features, allowing the system to identify the existing noise frequency components. By analyzing these frequency domain features, the system can effectively identify the frequency characteristics and power spectrum density distribution of the noise, thereby generating an environmental noise model. This model not only reflects the noise characteristics in the current environment, but also provides an important basis for subsequent noise filtering.
[0078] This step has efficient signal processing capabilities and improved environmental adaptability. Through directional beamforming technology, the system can effectively capture target sounds in noisy environments. The application of this technology significantly improves the accuracy of voiceprint recognition. In addition, the fast Fourier transform enables the system to analyze and identify noise frequencies in real time, thereby quickly building an environmental noise model. Such real-time response capabilities not only improve the robustness of the system, but also enhance its adaptability in different environments, ensuring that the target sound can still be accurately identified under changing noise conditions.
[0079] like Figure 4 As shown, the method of enhancing the identified target sound, suppressing the noise from other directions, performing spectrum analysis on the enhanced sound information, identifying the characteristic frequency of the noise, and generating an environmental noise characteristic model specifically includes:
[0080] S310, based on the calculated direction angle of the target sound, forming a directional beam for the microphone array, focusing the sensitivity of the array in the direction of the target sound, enhancing the target sound and suppressing interference signals from other directions;
[0081] S320, performing a fast Fourier transform on the enhanced target sound, converting the target sound from the time domain to the frequency domain, extracting frequency domain features, and identifying existing noise frequency components;
[0082] S330, based on the recognition result, analyzing the frequency components and power spectrum density distribution of the noise, and constructing an environmental noise model according to the analysis result.
[0083] S400, identifying the noise frequency existing in the current environment according to the environmental noise feature model, filtering out the background noise, and extracting the voiceprint signal of the target sound;
[0084] This step will divide the real-time collected audio signal into several short-time frame audio segments. This frame processing enables the system to analyze the signal changes more finely and extract the spectrum characteristics of each frame. By performing spectrum analysis on each short-time frame, the system can capture the frequency composition of the sound in a short period of time, thereby providing basic data for noise recognition.
[0085] After extracting the spectral features, the system conducts an in-depth analysis of the real-time extracted spectrum based on the environmental noise model. The goal of this analysis process is to determine the noise components in the current signal and identify the characteristic frequencies of the noise. In this way, the system can accurately identify the noise characteristics that exist in a specific environment, thus laying the foundation for subsequent noise filtering. Next, a frequency threshold is set so that the characteristic frequencies of noise that exceed the threshold are screened and marked. In this way, the system can effectively identify the noise components that interfere with the target sound, and subtract them from the spectrum of the current audio signal to generate a denoised signal spectrum.
[0086] The key to this process is to use filters to separate the noise frequency. By converting the filtered frequency domain signal back to the time domain, the system can reconstruct the target sound signal and extract the voiceprint features from it. This conversion from frequency domain to time domain not only ensures the integrity of the target sound, but also improves the accuracy of subsequent voiceprint recognition.
[0087] This step analyzes the audio signal in real time and extracts the spectrum features. The system can quickly identify and adapt to different noise environments, thereby effectively improving the success rate of voiceprint recognition. In addition, by setting frequency thresholds and using filters to accurately separate noise, the system ensures the clarity and integrity of the target sound and significantly reduces interference caused by environmental noise. This process not only improves the system's adaptability to background noise, but also enhances its robustness in complex acoustic environments, making voiceprint recognition technology more reliable and effective in real application scenarios.
[0088] like Figure 5 As shown, according to the environmental noise feature model, the noise frequency existing in the current environment is identified, and the background noise is filtered out to extract the voiceprint signal of the target sound;
[0089] S410, dividing the audio signal collected by the microphone in real time into a plurality of short time frame audio segments for analysis, and extracting the frequency spectrum features of each frame;
[0090] S420, analyzing the spectrum extracted in real time based on the environmental noise model, determining the noise component in the current signal, and identifying the characteristic frequency of the noise;
[0091] S430, setting a frequency threshold, screening and marking noise characteristic frequencies exceeding the frequency threshold, subtracting the marked noise spectrum from the spectrum of the current audio signal to obtain a denoised signal spectrum, and using a filter to separate the noise;
[0092] S440, converting the filtered frequency domain signal back to the time domain, reorganizing it into a target sound signal, and extracting voiceprint features.
[0093] S500, embedding an audio tag in a voiceprint signal, performing similarity matching between the voiceprint signal embedded with the audio tag and a voiceprint fingerprint in a voiceprint library, and determining whether the target voice is the user's voice.
[0094] This step embeds a preset audio tag in the corresponding voiceprint signal based on the extracted voiceprint features. This audio tag not only serves as a unique identifier, but also enhances the recognition efficiency of the voiceprint signal. By embedding audio tags in voiceprint signals, the system can ensure that each signal is closely associated with its corresponding user identity information, thereby providing higher accuracy in the subsequent matching process.
[0095] Next, the system confirms the identity information of the current user to be matched and extracts the voiceprint fingerprint related to the target user from the voiceprint library. This step ensures that the system can perform similarity calculations for specific users, avoids unnecessary calculations and processing, and improves overall efficiency.
[0096] After the extraction is completed, the system compares the voiceprint signal embedded with the audio tag with the voiceprint fingerprint in the voiceprint library, and determines whether the target voice is the user's voice by calculating the similarity. At this time, the system uses the Euclidean distance calculation method. The D shown in the formula represents the similarity, xi is the feature of the embedded voiceprint signal, and yi is the feature of the voiceprint fingerprint. Through this calculation, the system can evaluate the similarity between two voiceprint signals in a multi-dimensional feature space, and thus draw a conclusion on whether they match.
[0097] Finally, the system sets a decision threshold and compares the calculated similarity with the preset decision threshold. If the similarity is higher than the decision threshold, the target voice is determined to be the user's voice. This mechanism ensures the high reliability and accuracy of voiceprint recognition and avoids misjudgment or missed judgment.
[0098] This step not only improves the recognizability of the voiceprint signal by embedding audio tags, but also ensures the uniqueness of the user's identity, thereby preventing forgery and deception. In addition, by using the similarity calculation method, the system can effectively distinguish similar sounds and user voices in the multi-dimensional feature space, significantly improving the accuracy and robustness of recognition. This method maintains a high degree of adaptability in complex high-noise environments, making voiceprint recognition technology more reliable and effective. By setting appropriate decision thresholds, the system can also flexibly adapt to changes in sound features under different environmental conditions, further enhancing the practicality and stability of voiceprint recognition.
[0099] like Figure 6 As shown, the method of embedding an audio tag in a voiceprint signal, performing similarity matching between the voiceprint signal embedded with the audio tag and the voiceprint fingerprint in the voiceprint library, and determining whether the target sound is the user's voice specifically includes:
[0100] S510, embedding a preset audio tag in a corresponding voiceprint signal based on the extracted voiceprint feature;
[0101] S520, confirming the identity information of the current user to be matched, and extracting the voiceprint fingerprint related to the target user from the voiceprint database;
[0102] S530, compare the voiceprint signal embedded with the audio tag with the voiceprint fingerprint in the voiceprint library, calculate the similarity, and determine whether the voice is the user's voice:
[0103]
[0104] Among them, D represents the comparison similarity, x i is the feature of the embedded voiceprint signal, y i Represents the characteristics of the voiceprint fingerprint, n is the total number of characteristics;
[0105] S540, setting a decision threshold, comparing the calculated similarity with the preset decision threshold, and if the similarity is higher than the decision threshold, determining that the target voice is the user's voice.
[0106] Figure 7 The structure block diagram of the voiceprint recognition system for high noise environment provided by the embodiment of the present invention is as follows: Figure 7 As shown, the system comprises:
[0107] The user voiceprint feature extraction module 100 is used to pre-capture the user's voice signal, analyze the voice features, extract the voiceprint features, combine the extracted voiceprint features with the user's identity information to form a user-specific voiceprint fingerprint, and establish a voiceprint library for storage;
[0108] The array microphone sound information acquisition module 200 is used to obtain the positions of all array microphones and the sound information collected by each microphone in different directions in a high noise environment, and determine the direction of the target sound by calculating the phase difference and delay between the signals received by each microphone;
[0109] The target sound enhancement module 300 is used to enhance the identified target sound, suppress the noise from other directions, perform spectrum analysis on the enhanced sound information, identify the characteristic frequency of the noise, and generate an environmental noise characteristic model;
[0110] The target sound extraction module 400 is used to identify the noise frequency existing in the current environment according to the environmental noise feature model, filter out the background noise, and extract the voiceprint signal of the target sound;
[0111] The voiceprint matching judgment module 500 is used to embed an audio tag in the voiceprint signal, perform similarity matching between the voiceprint signal embedded with the audio tag and the voiceprint fingerprint in the voiceprint library, and judge whether the target voice is the user's voice.
[0112] It should be understood that, although each step in the flow chart of each embodiment of the present invention is shown in sequence according to the indication of the arrow, these steps are not necessarily performed in sequence according to the order indicated by the arrow. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be performed in other orders. Moreover, at least a portion of the steps in each embodiment may include a plurality of sub-steps or a plurality of stages, and these sub-steps or stages are not necessarily performed at the same time, but can be performed at different times, and the execution order of these sub-steps or stages is not necessarily performed in sequence, but can be performed in turn or alternately with at least a portion of other steps or sub-steps or stages of other steps.
[0113] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0114] The technical features of the above-described embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0115] The above-mentioned embodiments only express several implementation methods of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the patent of the present invention. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.
[0116] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A voiceprint recognition method for a high noise environment, characterized in that: The method comprises: Pre-capture the user's voice signal, analyze the voice features, extract the voiceprint features, combine the extracted voiceprint features with the user's identity information to form a user-specific voiceprint fingerprint, and establish a voiceprint library for storage; In a high-noise environment, the positions of all array microphones and the sound information collected by each microphone in different directions are obtained, and the direction of the target sound is determined by calculating the phase difference and delay between the signals received by each microphone; Enhance the identified target sound and suppress the noise from other directions. Perform spectrum analysis on the enhanced sound information, identify the characteristic frequency of the noise, and generate an environmental noise characteristic model. According to the environmental noise feature model, the noise frequency in the current environment is identified, and the background noise is filtered out to extract the voiceprint signal of the target sound; An audio tag is embedded in the voiceprint signal, and the voiceprint signal embedded with the audio tag is matched with the voiceprint fingerprint in the voiceprint library for similarity to determine whether the target voice is the user's voice.
2. The method according to claim 1, characterized in that The analyzing of the sound features, extracting the voiceprint features, combining the extracted voiceprint features with the user's identity information to form a user-specific voiceprint fingerprint, and establishing a voiceprint library for storage, specifically includes: Pre-recording the user's voice information, and pre-processing the recorded voice information, dividing the pre-processed voice information into a number of short time frame audio segments; Simulate the auditory characteristics of the human ear, obtain voiceprint features, and combine the extracted voiceprint features with the user's identity information to form a complete set of voiceprint fingerprint data; Create a voiceprint library structure to store the voiceprint fingerprints of all users, including user IDs and corresponding voiceprint feature sets.
3. The method according to claim 2, characterized in that The obtaining of the positions of all array microphones and the sound information collected by each microphone in different directions, and determining the direction of the target sound by calculating the phase difference and delay between the signals received by each microphone, specifically includes: Read the audio signals from different directions collected by each microphone and convert them into digital signals; Perform short-time Fourier transform on the signals collected by different microphones to obtain the time-frequency characteristics of each signal; The direction angle of the target sound is calculated based on the time delay from the target sound to each microphone and the microphone position relationship, and the signals collected by different microphones in the same time window are compared to calculate the phase difference: Among them, Δt i represents the signal delay received by the i-th microphone, d represents the distance between adjacent microphones in the array, θ represents the angle between the sound source and the microphone array, and c represents the propagation speed of sound waves in the air. represents the phase difference between the i-th microphone and the reference microphone, and f represents the frequency of the sound source signal.
4. The method according to claim 3, characterized in that The step of enhancing the identified target sound and suppressing noise from other directions, performing spectrum analysis on the enhanced sound information, identifying the characteristic frequency of the noise, and generating an environmental noise characteristic model specifically includes: Based on the calculated directional angle of the target sound, a directional beam is formed for the microphone array, the sensitivity of the array is concentrated in the direction of the target sound, the target sound is enhanced and interference signals from other directions are suppressed; Perform fast Fourier transform on the enhanced target sound, convert the target sound from the time domain to the frequency domain, extract frequency domain features, and identify the existing noise frequency components; Based on the recognition results, the frequency components and power spectral density distribution of the noise are analyzed, and an environmental noise model is constructed according to the analysis results.
5. The method according to claim 4, characterized in that The method further comprises: identifying the noise frequency existing in the current environment according to the environmental noise feature model, filtering out the background noise, and extracting the voiceprint signal of the target sound; The audio signal collected by the microphone in real time is divided into several short-time frame audio segments for analysis, and the spectrum characteristics of each frame are extracted; Based on the environmental noise model, the spectrum extracted in real time is analyzed to determine the noise components in the current signal and identify the characteristic frequency of the noise; Set a frequency threshold, filter and mark the noise characteristic frequencies that exceed the frequency threshold, subtract the marked noise spectrum from the spectrum of the current audio signal to obtain the denoised signal spectrum, and use a filter to separate the noise; The filtered frequency domain signal is converted back to the time domain, reorganized into the target sound signal, and the voiceprint features are extracted.
6. The method according to claim 5, characterized in that The embedding of the audio tag in the voiceprint signal, performing similarity matching between the voiceprint signal embedded with the audio tag and the voiceprint fingerprint in the voiceprint library, and determining whether the target sound is the user's voice specifically includes: Based on the extracted voiceprint features, a preset audio tag is embedded in the corresponding voiceprint signal; Confirm the identity information of the current user to be matched, and extract the voiceprint fingerprint related to the target user from the voiceprint database; Compare the voiceprint signal embedded in the audio tag with the voiceprint fingerprint in the voiceprint library, calculate the similarity, and determine whether the sound is the user's voice: Among them, D represents the comparison similarity, x i is the feature of the embedded voiceprint signal, y i Represents the characteristics of the voiceprint fingerprint, n is the total number of characteristics; A decision threshold is set, and the calculated similarity is compared with the preset decision threshold. If the similarity is higher than the decision threshold, the target voice is determined to be the user's voice.
7. Voiceprint recognition system for high noise environment, characterized in that: The system comprises: The user voiceprint feature extraction module is used to pre-capture the user's voice signal, analyze the voice features, extract the voiceprint features, combine the extracted voiceprint features with the user's identity information to form a user-specific voiceprint fingerprint, and establish a voiceprint library for storage; The array microphone sound information acquisition module is used to obtain the positions of all array microphones and the sound information collected by each microphone in different directions in a high-noise environment, and determine the direction of the target sound by calculating the phase difference and time delay between the signals received by each microphone; The target sound enhancement module is used to enhance the identified target sound and suppress the noise from other directions, perform spectrum analysis on the enhanced sound information, identify the characteristic frequency of the noise, and generate an environmental noise characteristic model; The target sound extraction module is used to identify the noise frequency existing in the current environment according to the environmental noise feature model, filter out the background noise, and extract the voiceprint signal of the target sound; The voiceprint matching judgment module is used to embed an audio tag in the voiceprint signal, perform similarity matching between the voiceprint signal embedded with the audio tag and the voiceprint fingerprint in the voiceprint library, and judge whether the target voice is the user's voice.
Citation Information
Cited By
Voiceprint recognition method and system for vehicle-mounted child safety seat
CN122050367A