Speaker identification method and system based on time-frequency domain dynamic characteristic matrix
By introducing the method of time-frequency domain dynamic feature matrix and multi-model fusion in speaker recognition technology, the problem of single feature extraction in the prior art is solved, and a more accurate and robust speaker recognition effect is achieved.
Patent Information
- Application Number
- CN202510118255.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-24
AI Technical Summary
The existing speaker recognition technology is single in feature extraction, ignoring the phase relationship between the voice signal in the time domain and the frequency domain, resulting in the inability to fully utilize the speaker-specific information.
A speaker recognition method based on the time-frequency domain dynamic feature matrix is proposed. By mapping the time-dynamic feature sequence of the original speech into a two-dimensional image and calculating the similarity matrix, the time-domain dynamic feature matrix is obtained. At the same time, the frequency-domain dynamic feature matrix is calculated by short-time Fourier transform, and feature extraction and fusion are combined with a convolutional neural network and a conformer model.
The time-frequency domain dynamic features of speech signals are extracted more comprehensively, which enhances the accuracy and robustness of speaker recognition. It is adapted to a variety of feature extraction solutions to effectively identify speakers in complex scenarios.
Smart Images

Figure CN119943058A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speaker recognition, and in particular to a speaker recognition method and system based on a dynamic feature matrix in time-frequency domain. Background Art
[0002] At present, speaker recognition is a technology that authenticates or identifies an individual by analyzing their voice characteristics. Compared with other biometric recognition technologies such as fingerprints and irises, voiceprints have the advantages of being unique, easy to obtain and non-invasive.
[0003] Common speaker recognition systems include acoustic feature extraction, model training, and back-end speaker discrimination. In general, speaker features are represented by spectral space vectors, such as spectrograms, FBank, and MFCC. The original speech file is input, and the speech signal is pre-processed by pre-emphasis, windowing, and other pre-processing operations. Based on the human vocalization mechanism or the human ear perception mechanism, different algorithms are used to extract features from the signal spectrum, and finally the acoustic features in vector form are obtained; in order to increase the discrimination ability of the model, the speaker features are processed by LDA, de-meaning, length regularization, and score regularization at the back end, and finally the Cosine scoring and PLDA and other discrimination models are used for similarity discrimination.
[0004] At the same time, with the development of deep learning, many works use raw waveforms as input to train speaker recognition models, and have achieved comparable performance to feature extraction methods such as FBank. Some works use pre-trained models as front-end feature extraction modules to replace traditional acoustic features. However, simply replacing traditional feature extraction methods with the general representation of pre-trained models may result in the speaker-specific information contained in the original speech not being fully utilized.
[0005] Commonly used speech features focus on extracting features related to the spectral envelope of speech signals, essentially ignoring the phase relationship between frequency components. Although Khadar Nawas et al. used a recurrence plot-based method to extract features and extracted the cyclic pattern of the vocal cord vibration system through phase space reconstruction, this method only processes speech in the time domain and ignores the key spectral information that distinguishes the characteristics of speakers. Therefore, finding a feature extraction scheme that can comprehensively model the dynamic characteristics of speech in both time and frequency domains requires further research.
[0006] In response to the above problems, the present application proposes a speaker recognition method and system based on a dynamic feature matrix in the time-frequency domain. Summary of the invention
[0007] This application proposes the following technical solutions to address one or more technical deficiencies in the above-mentioned prior art.
[0008] Based on the first aspect of the present application, a speaker recognition method based on a time-frequency domain dynamic feature matrix is proposed, comprising:
[0009] S1: Map the temporal dynamic feature sequence of the original speech into a two-dimensional image, and calculate the similarity of each frame of the original speech signal through the similarity matrix, and use the adaptive weighting method to enhance the time domain dynamic features in the temporal dynamic feature sequence to obtain the time domain dynamic feature matrix;
[0010] The time domain dynamic feature matrix is expressed as:
[0011] R ij =w(i,j)·θ(∈-||X||);
[0012]
[0013] Among them, R ij represents the value of the time domain dynamic feature matrix, ∈ represents the preset time domain dynamic feature threshold; ||X|| represents the distance between time points i and j, θ represents the step function, w(i, j) is the Gaussian weight of the time position, and σ represents the standard deviation used to adjust the weight decay rate;
[0014] S2: performing short-time Fourier transform on the original speech to obtain the spectrum value of each frame of speech signal, calculating the frequency domain dynamic features of each frame of speech signal of the original speech, and dynamically adjusting the similarity threshold to obtain a frequency domain dynamic feature matrix;
[0015] The frequency domain dynamic feature matrix is expressed as:
[0016] R f (i, j) = θ(ε f -‖‖S(i)-S(j)||);
[0017] Among them, R f (i, j) represents the frequency domain dynamic feature matrix, ε f represents the similarity threshold of the frequency domain dynamic feature matrix, ||S(i)-S(j)|| represents the distance between the spectral energy of frame i and frame j, and θ represents the step function;
[0018] S3: Input the frequency domain dynamic features and the time domain dynamic features into a convolutional neural network (CNN) for training, extract the acoustic features of the original speech signal by a traditional feature extraction method and input them into a conformer model for processing, so as to obtain the initial features of the speaker of the original speech;
[0019] S4: weighted adaptive fusion of the trained time domain dynamic features, the frequency domain dynamic features and the speaker initial features, and transmitting the fused feature vector to a fully connected layer and mapping it to a low-dimensional space;
[0020] S5: The feature fusion classifier calculates the speaker category probability distribution according to the feature vector output by the fully connected layer, and takes the category with the highest probability as the final speaker recognition result.
[0021] The time-domain dynamic feature matrix can extract the nonlinear dynamic changes of speech signals in time series, identify diagonal and block structures in images, and better extract the similarity of speech signals and the characteristics of speakers.
[0022] Furthermore, before S1, the process also includes extracting a temporal dynamic feature sequence from the original speech and dividing the original speech into frames of 25 ms.
[0023] Furthermore, the calculation formula of the similarity threshold of the frequency domain dynamic feature matrix is:
[0024] εf=α·mean(||S(i)-S(j)||);
[0025] Among them, α represents the adjustment factor, S(i) and S(j) represent the spectrum energy; mean represents the average value.
[0026] The frequency domain dynamic feature matrix can reflect the changing pattern of the speech signal in the spectrum.
[0027] Furthermore, the calculation formula of the spectrum energy is:
[0028] S(t)=[|X(t, f1)|, |X(t, f2)|, ..., |X(t, f F )|];
[0029]
[0030] Among them, X(t, f F ) represents the spectrum value of time t and frequency f, ω[nt] represents the short-time window function, and N represents the number of Fourier transform points.
[0031] Furthermore, the calculation formula for weighted adaptive fusion of the time domain dynamic features, the frequency domain dynamic features and the speaker initial features is:
[0032] e fusion =β1·e time +β2·e freq +β3·e feature ;
[0033] Among them, e fusion represents the fused feature vector, e time represents the time domain dynamic characteristics, e freq represents the frequency domain dynamic characteristics, e feature represents the initial speaker feature, β1, β2 and β3 represent feature fusion weight parameters.
[0034] By learning the weight ratios of the three groups of features and assigning different weight ratios, the contribution of different features is dynamically adjusted, so that the model can fully extract the dynamic and static features in the speaker's speech, integrate multi-dimensional information in the time domain and frequency domain, and enhance the expressiveness and discrimination capabilities of the features.
[0035] Based on the second aspect of the present application, a speaker recognition system based on a time-frequency domain dynamic feature matrix is also proposed, comprising:
[0036] Time domain module: maps the time dynamic feature sequence of the original speech into a two-dimensional image, calculates the similarity of each frame of the original speech signal through the similarity matrix, and uses the adaptive weighting method to enhance the time domain dynamic features in the time dynamic feature sequence to obtain the time domain dynamic feature matrix;
[0037] The time domain dynamic feature matrix is expressed as:
[0038] R ij =w(i,j)·θ(∈-||X||);
[0039]
[0040] Among them, R ij represents the value of the time domain dynamic feature matrix, ∈ represents the preset time domain dynamic feature threshold; ||X|| represents the distance between time points i and j, θ represents the step function, w(i, j) is the Gaussian weight of the time position, and σ represents the standard deviation used to adjust the weight decay rate;
[0041] Frequency domain module: performing short-time Fourier transform on the original speech to obtain the spectrum value of each frame of speech signal, calculating the frequency domain dynamic features of each frame of speech signal of the original speech, and dynamically adjusting the similarity threshold to obtain the frequency domain dynamic feature matrix;
[0042] The frequency domain dynamic feature matrix is expressed as:
[0043] R f (i, j) = θ(ε f -||S(i)-S(j)||);
[0044] Among them, R f(i, j) represents the frequency domain dynamic feature matrix, ε f represents the similarity threshold of the frequency domain dynamic feature matrix, ||S(i)-S(j)|| represents the distance between the spectral energy of frame i and frame j, and θ represents the step function;
[0045] Feature module: input the frequency domain dynamic features and the time domain dynamic features into the convolutional neural network CNN for training, extract the acoustic features of the original speech signal by traditional feature extraction method and input them into the conformer model for processing to obtain the initial features of the speaker of the original speech;
[0046] Adaptive fusion module: weighted adaptive fusion of the trained time domain dynamic features, the frequency domain dynamic features and the speaker initial features, and transmitting the fused feature vector to the fully connected layer and mapping it to a low-dimensional space;
[0047] Recognition module: The feature fusion classifier calculates the speaker category probability distribution according to the feature vector output by the fully connected layer, and takes the category with the highest probability as the final speaker recognition result.
[0048] Based on the third aspect of the present application, a computer program product is also proposed, which has one or more computer programs thereon, and when the computer program is executed by a computer processor, implements any of the methods described above.
[0049] The technical effect of the present application is that the present application solves the problem of the current single feature extraction method by adding and replacing feature extraction methods and models, and combining the dynamic information and time-frequency domain features of the speech signal, more fully retains the information for distinguishing the identity of the speaker in the speech signal, and uses different models to model the extracted features, more comprehensively extracts the global and local information of the speaker, can adapt to a variety of feature extraction schemes, and enhances the accuracy and robustness of speaker recognition in complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Other features, objects and advantages of the present application will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings.
[0051] Figure 1 It is an overall flow chart of a speaker recognition method based on a time-frequency domain dynamic feature matrix provided according to an embodiment of the present application.
[0052] Figure 2 It is a framework diagram of a speaker recognition system based on a time-frequency domain dynamic feature matrix provided according to an embodiment of the present application.
[0053] Figure 3It is a structural diagram of a computer system of an electronic device applicable to an embodiment of the present application. DETAILED DESCRIPTION
[0054] The present application will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the relevant invention, rather than to limit the invention. It should also be noted that, for ease of description, only the parts related to the relevant invention are shown in the accompanying drawings.
[0055] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0056] Figure 1 A speaker recognition method based on a time-frequency domain dynamic feature matrix of the present application is shown, comprising:
[0057] S1: Map the time dynamic feature sequence of the original speech into a two-dimensional image, and calculate the similarity of each frame of the speech signal of the original speech through the similarity matrix, and use the adaptive weighting method to enhance the time domain dynamic features in the time dynamic feature sequence to obtain the time domain dynamic feature matrix;
[0058] The time domain dynamic feature matrix is expressed as:
[0059] R ij =w(i,j)·θ(∈-||X||);
[0060]
[0061] Among them, R ij represents the value of the time domain dynamic feature matrix, ∈ represents the preset time domain dynamic feature threshold; ||X|| represents the distance between time points i and j, θ represents the step function, w(i, j) is the Gaussian weight of the time position, and σ represents the standard deviation used to adjust the weight decay rate;
[0062] S2: performing short-time Fourier transform on the original speech to obtain the spectrum value of each frame of speech signal, calculating the frequency domain dynamic features of each frame of speech signal of the original speech, and dynamically adjusting the similarity threshold to obtain a frequency domain dynamic feature matrix;
[0063] The frequency domain dynamic feature matrix is expressed as:
[0064] R f (i, j) = θ(ε f -||S(i)-S(j)||);
[0065] Among them, R f (i, j) represents the frequency domain dynamic feature matrix, εf represents the similarity threshold of the frequency domain dynamic feature matrix, ||S(i)-S(j)|| represents the distance between the spectral energy of frame i and frame j, and θ represents the step function;
[0066] S3: Input the frequency domain dynamic features and the time domain dynamic features into a convolutional neural network (CNN) for training, extract the acoustic features of the original speech signal by a traditional feature extraction method and input them into a conformer model for processing, so as to obtain the initial features of the speaker of the original speech;
[0067] S4: weighted adaptive fusion of the trained time domain dynamic features, the frequency domain dynamic features and the speaker initial features, and transmitting the fused feature vector to a fully connected layer and mapping it to a low-dimensional space;
[0068] S5: The feature fusion classifier calculates the speaker category probability distribution according to the feature vector output by the fully connected layer, and takes the category with the highest probability as the final speaker recognition result.
[0069] It should be noted that the calculation formula for the similarity threshold of the frequency domain dynamic feature matrix is:
[0070] ε f =α·mean(||S(i)-S(j)||);
[0071] Among them, α represents the adjustment factor, S(i) and S(j) represent the spectrum energy; mean represents the average value.
[0072] It should be noted that the calculation formula of the spectrum energy is:
[0073] S(t)=[|X(t, f1)|, |X(t, f2)|, ..., |X(t, f F )|];
[0074]
[0075] Among them, X(t, f F ) represents the spectrum value of time t and frequency f, ω[nt] represents the short-time window function, and N represents the number of Fourier transform points.
[0076] It should be noted that the calculation formula for weighted adaptive fusion of the time domain dynamic features, the frequency domain dynamic features and the speaker initial features is:
[0077] e fusion =β1·e time +β2·e freq +β3·e feature ;
[0078] Among them, e fusion represents the fused feature vector, e time represents the time domain dynamic characteristics, e freq represents the frequency domain dynamic characteristics, e feature represents the initial speaker feature, β1, β2 and β3 represent feature fusion weight parameters.
[0079] It should be noted that the convolutional neural network model CNN was trained for 50 iterations using the training framework pytorch lightning, and optimized using the optimizer adam W.
[0080] It should be noted that optimizing the convolutional neural network model CNN through the optimizer adam W is not only simple in implementation, but also computationally efficient. It can perform adaptive learning rate adjustment to ensure the stability of the training process and enable the model to converge better.
[0081] It should be noted that the time domain dynamic features and frequency domain features jointly provide a comprehensive description of the original speech signal. The time domain dynamic features describe the waveform of the speech signal that changes over time, including the amplitude, period and pulse width of the original language. The frequency domain dynamic features describe the distribution of the speech signal in frequency, including frequency distribution, frequency density and frequency components.
[0082] It should be noted that, preferably, the batch size during training is set to 256, the initial learning rate is set to 5e-4, and a linear decay learning rate strategy is adopted during the training process.
[0083] It should be noted that by learning the weight ratios of time domain dynamic features, frequency domain dynamic features and traditional acoustic features, giving the three features different weight ratios, and dynamically adjusting the contributions of different features, the model can fully extract the dynamic and static features of the speaker's voice, while integrating multi-dimensional information in the time domain and frequency domain to enhance the expressiveness and distinguishing capabilities of the features.
[0084] It should be noted that the fused features are input into the fully connected layer for feature compression and nonlinear mapping, so that the fused features are adapted to the classification task and classified in the classifier (such as AAM-Softmax). By introducing angular boundaries in the feature space, the intervals between different types of features are widened, which can effectively improve the accuracy and robustness of speaker classification and enhance the system's adaptability to different speech scenarios.
[0085] In a specific embodiment, through a time dynamic feature sequence with a value of 0 or 1, the hidden pattern, periodicity and nonlinear dynamics of the time series can be analyzed. The time dynamic feature sequence can be used to analyze the hidden pattern, periodicity and nonlinear dynamics of the time series. The time domain dynamic feature matrix can extract the nonlinear dynamic changes of the speech signal in the time series, identify the diagonal and block structures in the image, and better extract the similarity of the speech signal and the characteristics of the speaker.
[0086] In a specific embodiment, the value range of β is [0, 1], and the sum of β1, β2 and β3 is 1.
[0087] It should be noted that the present application solves the problem of the current single feature extraction method by adding and replacing feature extraction methods and models, and combining the dynamic information and time-frequency domain features of the speech signal, more fully retains the information in the speech signal that distinguishes the speaker's identity, and uses different models to model the extracted features, more comprehensively extracting the speaker's global and local information, and can adapt to a variety of feature extraction schemes to enhance the accuracy and robustness of speaker recognition in complex scenarios.
[0088] Reference below Figure 2 , Figure 2 A speaker recognition system based on a dynamic feature matrix in time-frequency domain is shown, comprising a time-domain module a, a frequency-domain module b, a feature module c, an adaptive fusion module d and a recognition module e.
[0089] In a specific embodiment, the time domain module a is configured to: map the time dynamic feature sequence of the original speech into a two-dimensional image, calculate the similarity of each frame of the speech signal of the original speech through a similarity matrix, and enhance the time domain dynamic features in the time dynamic feature sequence by an adaptive weighting method to obtain a time domain dynamic feature matrix;
[0090] The time domain dynamic feature matrix is expressed as:
[0091] R ij =w(i,j)·θ(∈-||X||);
[0092]
[0093] Among them, R ij represents the value of the time domain dynamic feature matrix, ∈ represents the preset time domain dynamic feature threshold; ||X|| represents the distance between time points i and j, θ represents the step function, w(i, j) is the Gaussian weight of the time position, and σ represents the standard deviation used to adjust the weight decay rate.
[0094] In a specific embodiment, the frequency domain module b is configured to: perform short-time Fourier transform on the original speech to obtain the spectrum value of each frame of speech signal, calculate the frequency domain dynamic features of each frame of speech signal of the original speech, and dynamically adjust the similarity threshold to obtain a frequency domain dynamic feature matrix;
[0095] The frequency domain dynamic feature matrix is expressed as:
[0096] R f (i, j) = θ(ε f -||S(i)-S(j)||);
[0097] Among them, R f (i, j) represents the frequency domain dynamic feature matrix, ε f represents the similarity threshold of the frequency domain dynamic feature matrix, ||S(i)-S(j)|| represents the distance between the spectral energy of frame i and frame j, and θ represents the step function.
[0098] In a specific embodiment, the feature module C is configured to: input the frequency domain dynamic features and the time domain dynamic features into a convolutional neural network CNN for training, extract the acoustic features of the original speech signal by a traditional feature extraction method and input them into a conformer model for processing to obtain the initial features of the speaker of the original speech.
[0099] In a specific embodiment, the adaptive fusion module d is configured to: perform weighted adaptive fusion on the trained time domain dynamic features, the frequency domain dynamic features and the speaker initial features, and transmit the fused feature vector to the fully connected layer and map it to a low-dimensional space.
[0100] In a specific embodiment, the recognition module e is configured as follows: a feature fusion classifier calculates the speaker category probability distribution according to the feature vector output by the fully connected layer, and takes the category with the highest probability as the final speaker recognition result.
[0101] It should be noted that the convolutional neural network model CNN was trained for 50 iterations using the training framework pytorch lighting, and optimized using the optimizer adam W.
[0102] It should be noted that, preferably, the batch size during training is set to 256, the initial learning rate is set to 5e-4, and a linear decay learning rate strategy is adopted during the training process.
[0103] It should be noted that the present application solves the problem of the current single feature extraction method by adding and replacing feature extraction methods and models, and combining the dynamic information and time-frequency domain features of the speech signal, more fully retains the information in the speech signal that distinguishes the speaker's identity, and uses different models to model the extracted features, more comprehensively extracting the speaker's global and local information, and can adapt to a variety of feature extraction schemes to enhance the accuracy and robustness of speaker recognition in complex scenarios.
[0104] Reference below Figure 3 , which shows a schematic diagram of the structure of a computer system suitable for implementing an electronic device of an embodiment of the present application. Figure 3 The electronic device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0105] like Figure 3 As shown, the computer system includes a central processing unit (CPU) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage part 308 into a random access memory (RAM) 303. In RAM 303, various programs and data required for system operation are also stored. CPU 301, ROM 302 and RAM 303 are connected to each other via a bus 304. An input / output (I / O) interface 305 is also connected to bus 304.
[0106] The following components are connected to the I / O interface 305: an input section 306 including a keyboard, a mouse, etc.; an output section 307 including a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 308 including a hard disk, etc.; and a communication section 309 including a network interface card such as a LAN card, a modem, etc. The communication section 309 performs communication processing via a network such as the Internet. A drive 310 is also connected to the I / O interface 305 as needed. A removable medium 311, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 310 as needed, so that a computer program read therefrom is installed into the storage section 308 as needed.
[0107] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable storage medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 309, and / or installed from the removable medium 311. When the computer program is executed by the central processing unit (CPU) 301, the above functions defined in the method of the present application are executed. It should be noted that the computer-readable storage medium of the present application can be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium can be, for example, - but not limited to - an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection with one or more conductors, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device, or device. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable storage medium other than a computer-readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, device, or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wireless, wireline, optical cable, RF, etc., or any suitable combination of the foregoing.
[0108] Computer program code for performing the operations of the present application may be written in one or more programming languages or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0109] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0110] The modules involved in the embodiments of the present application may be implemented by software or by hardware.
[0111] As another aspect, the present application also provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiment; or it may exist independently and not be assembled into the electronic device. The above computer-readable storage medium carries one or more programs. When the above one or more programs are executed by the electronic device, the electronic device: maps the time dynamic feature sequence of the original speech into a two-dimensional image, and calculates the similarity of each frame of speech signal of the original speech through a similarity matrix, and uses an adaptive weighting method to enhance the time domain dynamic features in the time dynamic feature sequence to obtain a time domain dynamic feature matrix; performs a short-time Fourier transform on the original speech to obtain the spectrum value of each frame of speech signal, calculates the frequency domain dynamic features of each frame of speech signal of the original speech, and dynamically adjusts the similarity threshold to obtain a frequency domain dynamic feature matrix; The frequency domain dynamic features and the time domain dynamic features are input into a convolutional neural network (CNN) for training; the acoustic features of the original speech signal are extracted by a traditional feature extraction method and input into a conformer model for processing to obtain the speaker's initial features of the original speech; the trained time domain dynamic features, the frequency domain dynamic features and the speaker's initial features are weighted and adaptively fused; the fused feature vector is transmitted to a fully connected layer and mapped to a low-dimensional space; a feature fusion classifier calculates the speaker's category probability distribution according to the feature vector output by the fully connected layer, and takes the category with the largest probability as the final speaker recognition result.
[0112] Finally, it should be noted that the above description is only a preferred embodiment of the present application and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the present application is not limited to the technical solution formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above invention concept. For example, the above features are replaced with the technical features with similar functions disclosed in this application (but not limited to) by each other to form a technical solution.
Claims
1. A speaker recognition method based on a dynamic feature matrix in time-frequency domain, characterized in that: include: S1: Map the time dynamic feature sequence of the original speech into a two-dimensional image, and calculate the similarity of each frame of the speech signal of the original speech through the similarity matrix, and use the adaptive weighting method to enhance the time domain dynamic features in the time dynamic feature sequence to obtain the time domain dynamic feature matrix; The time domain dynamic feature matrix is expressed as: R ij =w(i,j)·θ(∈-||X||); Among them, R ij represents the value of the time domain dynamic feature matrix, ∈ represents the preset time domain dynamic feature threshold; ||X|| represents the distance between time points i and j, θ represents the step function, w(i,j) is the Gaussian weight of the time position, and σ represents the standard deviation used to adjust the weight decay rate; S2: performing short-time Fourier transform on the original speech to obtain the spectrum value of each frame of speech signal, calculating the frequency domain dynamic features of each frame of speech signal of the original speech, and dynamically adjusting the similarity threshold to obtain the frequency domain dynamic feature matrix; The frequency domain dynamic feature matrix is expressed as: R f (i,j)=θ(ε f -||S(i)-S(j)||); Among them, R f (i, j) represents the frequency domain dynamic feature matrix, ε f represents the similarity threshold of the frequency domain dynamic feature matrix, ||S(i)-S(j)|| represents the distance between the spectral energy of frame i and frame j, and θ represents the step function; S3: Inputting the frequency domain dynamic features and the time domain dynamic features into a convolutional neural network (CNN) for training, extracting the acoustic features of the original speech signal by a traditional feature extraction method and inputting them into a conformer model for processing, thereby obtaining the initial features of the speaker of the original speech; S4: weighted adaptive fusion of the trained time domain dynamic features, the frequency domain dynamic features and the speaker initial features, and transmitting the fused feature vector to a fully connected layer and mapping it to a low-dimensional space; S5: The feature fusion classifier calculates the speaker category probability distribution according to the feature vector output by the fully connected layer, and takes the category with the highest probability as the final speaker recognition result.
2. The method according to claim 1, characterized in that Before S1, the process also includes extracting a temporal dynamic feature sequence from the original speech and dividing the original speech into frames of 25 ms.
3. The method according to claim 1, characterized in that The calculation formula of the similarity threshold of the frequency domain dynamic feature matrix is: ε f =α·mean(||S(i)-S(j)||); Among them, α represents the adjustment factor, S(i) and S(j) represent the spectrum energy; mean represents the average value.
4. The method according to claim 3, characterized in that The calculation formula of the spectrum energy is: S(t)=[|X(t,f1)|,|X(t,f2)|,…,|X(t,f F )|]; Among them, X(t,f F ) represents the spectrum value of time t and frequency f, ω[nt] represents the short-time window function, and N represents the number of Fourier transform points.
5. The method according to claim 1, characterized in that The calculation formula for weighted adaptive fusion of the time domain dynamic features, the frequency domain dynamic features and the speaker initial features is: and fusion =β1 e time +β2 e freq +β3 e featue ; Among them, e fusion represents the fused feature vector, e time represents the time domain dynamic characteristics, e freq represents the frequency domain dynamic characteristics, e feature represents the initial speaker feature, β1, β2 and β3 represent feature fusion weight parameters.
6. A speaker recognition system based on a dynamic feature matrix in the time-frequency domain, characterized in that: include: Time domain module: maps the time dynamic feature sequence of the original speech into a two-dimensional image, calculates the similarity of each frame of the original speech signal through the similarity matrix, and uses the adaptive weighting method to enhance the time domain dynamic features in the time dynamic feature sequence to obtain the time domain dynamic feature matrix; The time domain dynamic feature matrix is expressed as: R ij =w(i,j)·θ(∈-||X||); Among them, R ij represents the value of the time domain dynamic feature matrix, ∈ represents the preset time domain dynamic feature threshold; ||X|| represents the distance between time points i and j, θ represents the step function, w(i,j) is the Gaussian weight of the time position, and σ represents the standard deviation used to adjust the weight decay rate; Frequency domain module: performing short-time Fourier transform on the original speech to obtain the spectrum value of each frame of speech signal, calculating the frequency domain dynamic features of each frame of speech signal of the original speech, and dynamically adjusting the similarity threshold to obtain the frequency domain dynamic feature matrix; The frequency domain dynamic feature matrix is expressed as: R f (i,j)=θ(ε f -||S(i)-S(j)||); Among them, R f (i, j) represents the frequency domain dynamic feature matrix, ε f represents the similarity threshold of the frequency domain dynamic feature matrix, ||S(i)-S(j)|| represents the distance between the spectral energy of frame i and frame j, and θ represents the step function; Feature module: input the frequency domain dynamic features and the time domain dynamic features into the convolutional neural network CNN for training, extract the acoustic features of the original speech signal by traditional feature extraction method and input them into the conformer model for processing to obtain the initial features of the speaker of the original speech; Adaptive fusion module: weighted adaptive fusion of the trained time domain dynamic features, the frequency domain dynamic features and the speaker initial features, and transmitting the fused feature vector to the fully connected layer and mapping it to a low-dimensional space; Recognition module: The feature fusion classifier calculates the speaker category probability distribution according to the feature vector output by the fully connected layer, and takes the category with the highest probability as the final speaker recognition result.
7. A computer program product having one or more computer programs thereon, characterized in that: When the computer program is executed by a computer processor, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Anti-counterfeiting speaker identification method and system based on embedded feature fusion
CN117789727A
Voiceprint collection and analysis method and system based on artificial intelligence
CN118609574A
Voiceprint recognition method based on spatial cross learning multi-scale attention feature module
CN119091887A
Method, device and program for extracting audio feature amount
JP2003044077A
Real-time cumulative data processing method and device for flow-oriented integrated analysis
KR102690827B1
Cited By
Industrial field high-frequency voice recognition method and storage medium
CN120496577A