A short-time speech speaker recognition system and method based on deep learning
By adopting a deep learning-based spatiotemporal Transformer model in the short-term speech speaker recognition system, the problem of insufficient recognition efficiency and accuracy in the prior art is solved, and more efficient and accurate speaker recognition is achieved.
Patent Information
- Application Number
- CN202210464168.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-29
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-04-29
AI Technical Summary
The prior art lacks efficiency and accuracy in short-term voice speaker recognition, especially when faced with a large number of registrations and verifications, there are defects and deviations in manual judgments.
A short-term voice speaker recognition system based on deep learning is adopted, including a speaker voice collection module, a speech recognition processing module, a sample database, a short-term voice speaker recognition model based on space-time Transformer, a judgment scoring module and a result output module. The system generates the speaker's deep embedding through spectral segmentation and fusion of space-time features, and makes similarity comparison and scoring judgment.
It significantly improves the efficiency and accuracy of short-term voice speaker recognition, and can accurately judge the speaker's identity and output evaluation indicators when voice input is input in real time.
Smart Images

Figure CN114822559B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a short-time speech speaker recognition system based on deep learning and a method based on the system, and belongs to the field of computer and security. Background Art
[0002] Nowadays, with the rapid development of artificial intelligence, the field of identity authentication has also developed rapidly. Among them, identity authentication technology based on deep learning has been widely used in national defense security, monitoring and tracking, network identity authentication, etc.
[0003] Speaker recognition has developed rapidly in the past decade. With the speaker's identity information, text-independent speaker recognition can be used for national defense security, monitoring and tracking, network authentication, etc. In general, the speaker verification system filters out all non-target speakers from a number of speakers and finds the speaker with the same identity as the registered voice. The process of a basic speaker recognition system is as follows: In the training phase, the acoustic features of the original speech are used to obtain features for distinguishing speakers using machine learning methods, and the embedding vectors of all speakers are stored in the feature library of the model. In the verification phase, the system performs the same feature extraction process on the speech and obtains the speaker's identity by finding the speaker with the highest discrimination score from the feature library. Although authentication through voice does not require a complex sampling process, in actual scenarios, most speech lasts only 1-5 seconds. Therefore, short-term speech speaker recognition plays a very critical role and significance. In the past decade, traditional methods such as the i-vector method have been the main methods for speaker recognition. Generally speaking, the x-vector method has achieved better performance than the i-vector method in text-independent speaker recognition. For short-term speech speaker recognition, different scholars use data mining technology to extract different acoustic features to fit the model, and some scholars optimize the typical speaker recognition model from the model perspective. However, in real scenarios, the above methods do not provide substantial performance improvements.
[0004] At present, for example, important identity verification links require not only verification of facial information, but also verification of the person being tested's voice information. Faced with a large number of registrations and verifications, manual judgment of the results has certain defects and deviations. Therefore, it is necessary to use a short-term speech speaker recognition system and method based on deep learning. Summary of the invention
[0005] The purpose of the present invention is to greatly improve the efficiency and accuracy of short-term speech speaker recognition.
[0006] To achieve the above object, a technical solution of the present invention is to provide a short-time speech speaker recognition system based on deep learning, which is characterized by including a speaker speech acquisition module, a speech recognition processing module, a sample database, a short-time speech speaker recognition model based on spatio-temporal Transformer, a determination scoring module, and a result output module, wherein:
[0007] Speaker speech acquisition module: used to collect the speaker's speech and obtain the original audio data;
[0008] Speech recognition processing module: used to perform recognition processing on the original audio data by using an acoustic feature extraction method to obtain the speaker's speech spectrogram, and perform normalization to obtain a standard speaker's speech spectrogram;
[0009] Sample database: used to train the short-time speech speaker recognition model based on spatio-temporal Transformer. The sample database stores standard speaker sample speech spectrograms and corresponding sample labels;
[0010] The short-time speech speaker recognition model based on spatio-temporal Transformer further includes a spectrogram segmentation module and a spatio-temporal feature fusion module. Among them, the spectrogram segmentation module: used to slice a standard speaker's speech spectrogram to be recognized to obtain a series of time feature patches and space feature patches; the spatio-temporal feature fusion module: used to fuse the time feature patches and space feature patches to obtain the deep embedding of the speaker's verification speech or the deep embedding of the speaker's registration speech;
[0011] Determination scoring module: used to compare the deep embedding of the speaker's verification speech with the deep embedding of the target speaker's registration speech and perform scoring judgment;
[0012] Result output module: used to output the identity of the speaker who inputs speech in real time by using the speaker speech acquisition module, and output the corresponding target speaker and the scoring result.
[0013] Preferably, the result output module is signal-connected to a display screen and / or a printer, and the result is output by using the display screen and / or the printer.
[0014] Another technical solution of the present invention is to provide a short-time speech speaker recognition method based on deep learning, which uses the aforementioned short-time speech speaker recognition system, and is characterized by including the following steps:
[0015] Step 1. When the user registers, use the speaker voice acquisition module, the speech recognition processing module, and the short-time speech speaker recognition model based on spatio-temporal Transformer to obtain the deep embedding of the user's registered voice, and bind the deep embedding of the registered voice to the user identity;
[0016] Step 2. When identifying the current speaker, collect the original audio data of the current speaker through the speaker voice acquisition module;
[0017] Step 3. The speech recognition processing module extracts the standard speaker speech spectrogram based on the original audio data;
[0018] Step 4. Input the standard speaker speech spectrogram into the trained short-time speech speaker recognition model based on spatio-temporal Transformer. The short-time speech speaker recognition model processes the input speaker speech spectrogram using the following steps:
[0019] Step 401. Divide the input speaker speech spectrogram into square subset images of size M×N;
[0020] Step 402. Take the directly segmented subset images as time feature patches, and transpose the time feature patches to obtain spatial feature patches;
[0021] Step 403. Input the time feature patches and spatial feature patches into the spatio-temporal feature fusion module. The processing process of the spatio-temporal feature fusion module for the time feature patches and spatial feature patches specifically includes the following steps
[0022] Step 4031. Extract the tokens of the time feature patches and spatial feature patches through the dynamic convolution layer;
[0023] Step 4032. Add position encoding to each token;
[0024] Step 4033. Input the tokens with added position encoding into the N-layer Transformer module, and obtain the time deep embedding and spatial deep embedding through the normalization layer, multi-head attention layer, and feed-forward layer respectively;
[0025] Step 4034. Perform embedding fusion on the obtained time deep embedding and spatial deep embedding to obtain the deep embedding of the verification voice of the speaker or obtain the deep embedding of the registered voice of the speaker;
[0026] Step 5. Compare the similarity between the deep embedding of the verification voice of the current speaker and the deep embedding of the registered voice of the target speaker, and score;
[0027] Step 6: Output the identity of the target speaker and the corresponding scoring result.
[0028] Preferably, in Step 3, the speech recognition processing module extracts the standard speaker speech spectrogram using the following steps:
[0029] Step 301: Input the original audio data into a high-pass filter to enhance the high-frequency part of the original audio data and improve the signal-to-noise ratio, thereby realizing the pre-emphasis of the original audio data;
[0030] Step 302: Frame the pre-emphasized original audio data;
[0031] Step 303: Multiply each audio frame by a window function to increase the continuity at the left and right ends of each audio frame;
[0032] Step 304: Use the fast Fourier transform to calculate the power spectrum of each audio frame;
[0033] Step 305: Use 40 Mel-scale triangular filters to extract the frequency bands of the power spectrum. The triangular filters are sparsely distributed in the high-frequency region and densely distributed in the low-frequency region, thereby simulating the effect that the human ear has better resolution for low-frequency sounds;
[0034] Step 306: Combine the features of the 40 Mel-scale triangular filters to obtain the acoustic features of a segment of original audio data, that is, the speaker's speech spectrogram;
[0035] Step 307: Normalize the speaker's speech spectrogram to obtain the standard speaker's speech spectrogram.
[0036] Preferably, in Step 4033, the token added with positional encoding is sent into the N-layer Transformer module, and the temporal depth embedding and spatial depth embedding of the j-th segment of speech of the i-th speaker are obtained through the normalization layer, multi-head attention layer, and feed-forward layer respectively and spatial depth embedding
[0037] Then in Step 4034, the obtained temporal depth embedding and spatial depth embedding are fused in the following way to obtain the depth embedding of the j-th segment of speech of the i-th speaker
[0038]
[0039] where represents the concatenation operation.
[0040] Preferably, when training the short-time speech speaker recognition model based on spatio-temporal Transformer, after obtaining the speaker speech spectrogram of the sample language data by using the speech recognition processing module, the speaker speech spectrogram is input into the sample database, and then the sample database is used to train the short-time speech speaker recognition model based on spatio-temporal Transformer.
[0041] Preferably, when training the short-time speech speaker recognition model based on spatio-temporal Transformer, after the step 4034, the following steps are further included:
[0042] Input the deep embedding obtained in step S4034 into the linear classification layer, and use the cross-entropy loss function to train the model; when using the cross-entropy loss function, the Softmax function is used to convert the output of the model into probability values, as shown in the following formula:
[0043]
[0044] In the formula, and respectively represent the output value of the linear classification layer of the j-th segment of speech belonging to the c-th class and the probability value of predicting that it belongs to the c-th class. C is the number of classes in the training set, and the cross-entropy loss function is expressed as:
[0045]
[0046] In the formula, is the label of the i-th sample belonging to the c-th class, and B is the value of the batch in the training process.
[0047] Preferably, in step 6, the identity of the target speaker and the corresponding scoring result are output by using a display screen and / or a printer.
[0048] Through the above technical solutions, the present invention has the following effects:
[0049] The structure of the present invention is reasonably designed. Supervised learning is carried out by using the sample database, and the segmented spectrogram sequence is sent into the deep neural network based on spatio-temporal Transformer. Finally, the identity of the speaker of the given speech is accurately judged, and evaluation indexes are output to evaluate the output result, and the efficiency and accuracy in the short-time speech speaker recognition process are greatly improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 is a structural block diagram of the short-time speech speaker recognition system based on deep learning of the present invention;
[0051] Figure 2This is the flowchart of the short-time speech speaker recognition method based on deep learning according to the present invention;
[0052] Figure 3 This is a schematic diagram of the feature fusion module based on spatio-temporal Transformer in the short-time speech speaker recognition method based on deep learning according to the present invention. Detailed implementation manners
[0053] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only for helping those of ordinary skill in the art to understand the principles and knowledge of the present invention, and are not used to limit the scope of the present invention, and should not be considered as limiting the application scenarios of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, but the deformations, changes and conversions made to the embodiments based on the principles and purposes of the present invention also fall within the scope defined by the appended claims of this application. And it is obvious that this specification only takes the preferred implementation manners as examples and does not need to elaborate all the implementation manners.
[0054] As Figure 1 shown, an embodiment of the present invention proposes a short-time speech speaker recognition system based on deep learning, including a speaker speech acquisition module, a speech recognition processing module, a sample database, a short-time speech speaker recognition model based on spatio-temporal Transformer, a determination scoring module and a result output module, wherein:
[0055] The speaker speech acquisition module: used to acquire the speaker's speech and obtain the original audio data;
[0056] The speech recognition processing module: used to perform recognition processing on the original audio data by using an acoustic feature extraction method to obtain the speaker's speech spectrogram, and perform normalization to obtain a standard speaker's speech spectrogram;
[0057] The sample database: used to train the short-time speech speaker recognition model based on spatio-temporal Transformer. The sample database stores standard speaker sample speech spectrograms and corresponding sample labels;
[0058] The short-time speech speaker recognition model based on spatio-temporal Transformer further includes a spectrogram segmentation module and a spatio-temporal feature fusion module. Among them, the spectrogram segmentation module: used to slice a standard speaker's speech spectrogram to be recognized to obtain a series of time feature patches and space feature patches; the spatio-temporal feature fusion module: used to fuse the time feature patches and space feature patches to obtain the deep embedding of the speaker's verification speech or the deep embedding of the speaker's registration speech;
[0059] Determination and scoring module: used to compare the deep embedding of the verification speech of the speaker with the deep embedding of the registered speech of the target speaker, and perform scoring and judgment;
[0060] Result output module: used to output the result of the speaker identity of the input speech.
[0061] In this embodiment, when collecting the speaker's voice, the original audio data of the speaker is obtained through a microphone. The microphone collects audio by using the electromagnetic induction phenomenon. When the sound wave vibrates the diaphragm, the coil (called the voice coil) connected to the diaphragm vibrates together. The voice coil vibrates in the magnetic field, and an induced current (electrical signal) is generated therein. The magnitude and direction of the induced current both change, and the changing amplitude and frequency are determined by the sound wave. Therefore, the microphone can convert the voice signal into an electrical signal. After being processed by an electronic computer to obtain a spectrogram, it is very easy to collect audio with a microphone. Therefore, speaker recognition has great potential advantages for identity authentication.
[0062] In this embodiment, the result output module is signal-connected to a display screen and a printer. By setting the result output module to be signal-connected to a display screen and a printer, the screen display and document printing of the determination result are realized, which is convenient for result analysis.
[0063] The following lists the preferred embodiments of the short-time speech speaker recognition system and method based on deep learning to clearly illustrate the content of the present invention. It should be clear that the content of the present invention is not limited to the following embodiments, and other improvements by the conventional technical means of those of ordinary skill in the art are also within the scope of the idea of the present invention.
[0064] As Figure 2 shown, the embodiment of the present invention proposes a short-time speech speaker recognition method based on deep learning, including the following steps:
[0065] S100. Original data collection: Collect the original audio data of the speaker through a microphone, extract the acoustic features of the original audio data, and obtain a standard speaker speech spectrogram. When training the short-time speech speaker recognition model based on spatio-temporal Transformer, input the extracted standard speaker speech spectrogram into the sample database, and use the sample database to train the short-time speech speaker recognition model based on spatio-temporal Transformer.
[0066] Specifically, the extraction of acoustic features includes the following steps:
[0067] S101. By inputting the original audio data into a high-pass filter, enhance the high-frequency part of the original audio data and improve the signal-to-noise ratio, so as to realize the pre-emphasis of the original audio data.
[0068] S102. Frame the pre-emphasized original audio data.
[0069] Typically, 25 ms is aggregated into one frame. At the same time, in order to avoid excessive changes between adjacent frames, there needs to be an overlapping area of 10 ms between adjacent frames.
[0070] S103. Windowing: Multiply each audio frame by a window function, aiming to increase the continuity at the left and right ends of each audio frame.
[0071] S104. Use the fast Fourier transform to calculate the power spectrum of each audio frame to obtain the spectral line energy of the original audio data.
[0072] S105. Use 40 Mel-scale triangular filters to extract the frequency bands of the above power spectrum. The triangular filters are sparsely distributed in the high-frequency region and densely distributed in the low-frequency region, thus simulating the effect that the human ear has better resolution for low-frequency sounds.
[0073] S106. Combine the features of 40 Mel-scale triangular filters to obtain the acoustic features of a segment of original audio data, that is, the speaker voice spectrogram.
[0074] S107. Normalize the speaker voice spectrogram. Subtract its mean from the feature itself and then divide by its variance to obtain the standard speaker voice spectrogram. When training the short-time speech speaker recognition model based on spatio-temporal Transformer, the standard speaker voice spectrogram is in the sample database.
[0075] S200. Input the standard speaker voice spectrogram into the trained short-time speech speaker recognition model based on spatio-temporal Transformer, and the short-time speech speaker recognition model based on spatio-temporal Transformer extracts the deep embedding of the verification speech of the speaker.
[0076] Use the sample database to train the short-time speech speaker recognition model based on spatio-temporal Transformer. The processing process of the short-time speech speaker recognition model based on spatio-temporal Transformer for the input speaker voice spectrogram specifically includes the following steps:
[0077] S201. Divide the input speaker voice spectrogram into square subset images of size M×N.
[0078] S202. Use the directly segmented subset images as time feature patches, and transpose the time feature patches to obtain spatial feature patches;
[0079] S203. Input the time feature patches and spatial feature patches into the spatio-temporal feature fusion module, such as Figure 3As shown, the processing process of the spatio-temporal feature fusion module for the time feature patch and the spatial feature patch specifically includes the following steps
[0080] S2031. Extract the tokens of the time feature patch and the spatial feature patch through the dynamic convolution layer;
[0081] S2032. Add position encoding to each token;
[0082] S2033. Feed the tokens with added position encoding into the Transformer module of N layers, and obtain the time-depth embedding of the j-th segment of speech of the i-th speaker through the normalization layer, the multi-head attention layer and the feed-forward layer and the spatial-depth embedding
[0083] S2034. Fuse the obtained time-depth embedding and the spatial-depth embedding in the following way to obtain the depth embedding of the j-th segment of speech of the i-th speaker
[0084]
[0085] In the formula, represents the concatenation operation.
[0086] When training the short-time speech speaker recognition model based on spatio-temporal Transformer, the following steps are further included after step 2034:
[0087] Step S2035. Input the depth embedding obtained in step S2034 into the linear classification layer, and use the cross-entropy loss function to train the model. When using the cross-entropy loss function, the Softmax function is used to convert the output (logits) of the model into probability values, as shown in the following formula:
[0088]
[0089] In the formula, and respectively represent the output value of the linear classification layer of the j-th segment of speech belonging to the c-th class and the probability value of predicting that it belongs to the c-th class. C is the number of classes in the training set. Therefore, the cross-entropy loss function is expressed as:
[0090]
[0091] In the formula, is the label of the i-th sample belonging to the c-th category, and B is the value of the batch during the training process. The derivative of the cross-entropy loss for multiple classes is simpler, which is good at learning information between classes. Moreover, the gradient of the weights in the last layer of the network is not related to the derivative of the activation function, resulting in an accelerated update speed of the weight matrix and faster convergence during training. Through the above training process, after the short-time speech speaker recognition model based on the spatio-temporal Transformer converges, the output of the layer before the linear classification layer is used as the deep embedding of the speaker. The deep embeddings of the registered speech and the verification speech of the speaker can be obtained during registration and verification.
[0092] S300. Compare the deep embedding of the verification speech of the speaker with the deep embedding of the registered speech input by the target speaker during registration and score.
[0093] S400. Analysis result display:
[0094] Output the identity of the target speaker and the corresponding scoring result. When outputting, the result can be displayed on the display screen or the output report result can be printed by a printer.
[0095] The specific operation process of the present invention corresponding to a computer is as follows:
[0096] (1) Turn on the recording device, enter the speaker's name, click the registration button, and click start recording;
[0097] (2) According to the text appearing on the screen, the speaker can perform shadowing to ensure a registration audio of about 5 seconds.
[0098] (3) The speaker for testing turns on the recording device, clicks the test button, clicks start recording, and records the speech according to the prompt.
[0099] (4) Click the image analysis button, and the tool will automatically analyze and evaluate the speech.
[0100] (5) The tool pops up a dialog box indicating that the analysis is completed, and displays the determination result of the test speech, such as "The speaker is xx, and the probability is 85%".
[0101] In summary, compared with the prior art, the present invention has the following advantages:
[0102] The present invention uses a sample database for supervised learning, sends the segmented spectrogram sequence into a computing system, improves the information content of short-time speech acoustic features by fusing the spatio-temporal information of acoustic features, uses the Transformer results to greatly enhance the ability to extract speaker discriminative features, finally calculates the evaluation score of the speech to be measured, assists speech workers and identity authentication systems to carry out further research and authentication, and can evaluate the output results through evaluation indicators, greatly improving the efficiency and accuracy in the process of short-time speech speaker recognition compared with the judgments of manual and traditional machine learning methods.
Claims
1. A short-time speech speaker recognition system based on deep learning, characterized in that, it includes a speaker speech acquisition module, a speech recognition processing module, a sample database, a short-time speech speaker recognition model based on spatio-temporal Transformer, a determination scoring module, and a result output module, where: Speaker speech acquisition module: used to acquire speaker speech and obtain original audio data; Speech recognition processing module: used to perform recognition processing on the original audio data by using an acoustic feature extraction method to obtain a speaker speech spectrogram, and perform normalization to obtain a standard speaker speech spectrogram; Sample database: used to train the short-time speech speaker recognition model based on spatio-temporal Transformer. The sample database stores standard speaker sample speech spectrograms and corresponding sample labels; The short-time speech speaker recognition model based on spatio-temporal Transformer further includes a spectrogram segmentation module and a spatio-temporal feature fusion module. Among them, the spectrogram segmentation module: used to slice a standard speaker speech spectrogram to be recognized to obtain a series of time feature patches and space feature patches; the spatio-temporal feature fusion module: used to fuse the time feature patches and space feature patches to obtain the deep embedding of the verification speech of the speaker or the deep embedding of the registration speech of the speaker; Determination scoring module: used to compare the deep embedding of the verification speech of the speaker with the deep embedding of the registration speech of the target speaker and perform scoring judgment; Result output module: used to output the result of the identity of the speaker who inputs the speech in real time by using the speaker speech acquisition module, and output the corresponding target speaker and the scoring result.
2. The short-time speech speaker recognition system based on deep learning according to claim 1, characterized in that, the result output module is signal-connected to a display screen and / or a printer, and the result is output by using the display screen and / or the printer.
3. A short-time speech speaker recognition method based on deep learning, using the short-time speech speaker recognition system according to any one of claims 1 to 2, characterized in that, it includes the following steps: Step 1: When the user registers, use the speaker speech acquisition module, the speech recognition processing module, and the short-time speech speaker recognition model based on spatio-temporal Transformer to obtain the deep embedding of the registration speech of the user, and bind the deep embedding of the registration speech to the user identity; Step 2: When recognizing the current speaker, collect the original audio data of the current speaker through the speaker speech acquisition module; Step 3: The speech recognition processing module extracts a standard speaker speech spectrogram based on the original audio data; Step 4: Input the standard speaker speech spectrogram into the trained short-time speech speaker recognition model based on spatio-temporal Transformer. The short-time speech speaker recognition model processes the input speaker speech spectrogram by the following steps: Step 401: Divide the input speaker voice spectrogram into square subset images of size M×N; Step 402: Use the directly segmented subset images as temporal feature patches, and transpose the temporal feature patches to obtain spatial feature patches; Step 403: Input the temporal feature patches and spatial feature patches into the spatio-temporal feature fusion module. The processing of the temporal feature patches and spatial feature patches by the spatio-temporal feature fusion module specifically includes the following steps: Step 4031: Extract tokens of the temporal feature patches and spatial feature patches through a dynamic convolutional layer; Step 4032: Add positional encoding to each token; Step 4033: Send the tokens with positional encoding into the N-layer Transformer module, and obtain temporal depth embedding and spatial depth embedding through a normalization layer, a multi-head attention layer, and a feed-forward layer respectively; Step 4034: Perform embedding fusion on the obtained temporal depth embedding and spatial depth embedding to obtain the depth embedding of the verification speech of the speaker or the depth embedding of the registered speech of the speaker; Step 5: Compare the depth embedding of the verification speech of the current speaker with the depth embedding of the registered speech of the target speaker, and score; Step 6: Output the identity of the target speaker and the corresponding scoring result.
4. A method for short-time speech speaker recognition based on deep learning as described in claim 3, characterized in that, in step 3, the speech recognition processing module extracts the standard speaker voice spectrogram by the following steps: Step 301: Input the original audio data into a high-pass filter to enhance the high-frequency part of the original audio data and improve the signal-to-noise ratio, so as to realize pre-emphasis of the original audio data; Step 302: Frame the pre-emphasized original audio data; Step 303: Multiply each audio frame by a window function to increase the continuity of the left and right ends of each audio frame; Step 304: Use the fast Fourier transform to calculate the power spectrum of each audio frame; Step 305: Use 40 Mel-scale triangular filters to extract the frequency bands of the power spectrum. The triangular filters are sparsely distributed in the high-frequency region and densely distributed in the low-frequency region, so as to simulate the effect that the human ear has better resolution for low-frequency sounds; Step 306: Combine the features of 40 Mel-scale triangular filters to obtain the acoustic features of a segment of original audio data, that is, the speaker voice spectrogram; Step 307: Normalize the speaker voice spectrogram to obtain the standard speaker voice spectrogram.
5. A method for short-time speech speaker recognition based on deep learning as described in claim 3, characterized in that, In step 4033, the tokens with added positional encoding are fed into the Transformer module of N layers, and the temporal depth embedding of the j-th segment of speech of the i-th speaker is obtained through a normalization layer, a multi-head attention layer, and a feed-forward layer respectively and the spatial depth embedding Then in step 4034, a temporal depth embedding is obtained and a spatial depth embedding are fused in the following manner to obtain the depth embedding of the j-th segment of speech of the i-th speaker In the formula, represents the splicing operation.
6. A method for short-time speech speaker recognition based on deep learning as described in claim 3, characterized in that, When training the short-time speech speaker recognition model based on spatio-temporal Transformer, after obtaining the speaker speech spectrogram of the sample language data by using the speech recognition processing module, the speaker speech spectrogram is input into the sample database, and then the sample database is used to train the short-time speech speaker recognition model based on spatio-temporal Transformer.
7. A short-time speech speaker recognition method based on deep learning as claimed in claim 6, characterized in that, when training the short-time speech speaker recognition model based on spatio-temporal Transformer, the following steps are further included after the step 4034: Input the deep embedding obtained in step S4034 into a linear classification layer, and use the cross-entropy loss function to train the model; when using the cross-entropy loss function, the Softmax function is used to convert the output of the model into probability values, as shown in the following formula: wherein, and respectively represent the output value of the linear classification layer of the j-th segment of speech belonging to the c-th class and the probability value of predicting that it belongs to the c-th class. C is the number of classes in the training set, and the cross-entropy loss function is expressed as: wherein, is the label of the i-th sample belonging to the c-th class, and B is the value of the batch during the training process.
8. A short-time speech speaker recognition method based on deep learning as claimed in claim 3, characterized in that, in step 6, the identity of the target speaker and the corresponding scoring result are output by using a display screen and / or a printer.
Citation Information
Patent Citations
Techniques for variable resolution encoding and decoding of digital video
CN101507278A
Speech emotion recognition based on slice convolution
CN107358946A