Single-channel speech transcription method and device, electronic equipment and storage medium

Through the single-channel speech transcription method, matrix decomposition and neural network model are used to perform speech separation and sensitive word masking, which solves the transcription accuracy and data security problems when multiple people speak, and realizes efficient speech transcription and information protection in scenarios such as banks.

CN120496536APending Publication Date: 2025-08-15AGRICULTURAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510986182.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-17
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The prior art cannot effectively ensure the accuracy of speech transcription and data security in scenarios such as banks, especially when multiple people speak at the same time, and sensitive information is easily leaked.

Method used

The single-channel speech transcription method is adopted, and the non-negative matrix decomposition and time-frequency mask calculation are used to perform non-negative matrix decomposition and time-frequency mask calculation through time-frequency domain preprocessing, combined with the neural network model to perform speech separation, and sensitive words are classified and masked during the speech recognition process.

Benefits of technology

It realizes that the speech of multiple speakers is accurately separated without relying on microphone arrays, and effectively block sensitive information during the transcription process, improving the accuracy of transcription and data security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120496536A_ABST
    Figure CN120496536A_ABST
Patent Text Reader

Abstract

The invention discloses a single-channel voice transcription method and device, electronic equipment and a storage medium, and the method comprises the steps: collecting a current original mixed voice signal of a single channel, and carrying out the time-frequency domain preprocessing of the current original mixed voice signal, and obtaining an amplitude spectrum and a phase spectrum of the current original mixed voice signal; performing non-negative matrix factorization on the amplitude spectrum through a matrix factorization model, calculating a time-frequency mask by using a factorization result, performing signal reconstruction by using the time-frequency mask and the amplitude spectrum, and outputting a plurality of current separated voice amplitude spectrums; wherein the matrix decomposition model and the neural network model are subjected to joint optimization training; converting each current separated voice amplitude spectrum by using the phase spectrum of the current original mixed voice signal to obtain each current separated voice; and performing speech recognition and sensitive word classification on each current separated speech through the speech transcription model, and after fusing a speech recognition text and a sensitive word classification result, outputting a transcription text of each current separated speech shielding the sensitive words.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of speech transcription technology, and in particular to a single-channel speech transcription method and device, electronic device, and storage medium. Background Art

[0002] With the advancement of digital transformation in financial services, banks are increasingly using remote conferencing to facilitate multi-party communication and record meeting minutes during business operations and internal office processes. Because manual recording is prone to errors, voice transcription is currently the primary method used. This involves capturing and transcribing voice into text, then masking sensitive information to generate meeting minutes.

[0003] Because voice signals collected in live scenarios cannot avoid the aliasing of surrounding noise and other voice signals, current methods mainly use microphone arrays for multi-channel voice collection. By analyzing the voice collected by each channel, a primary voice, namely the current speaker's voice, can be identified, thereby filtering out other audio. Then, based on the isolated current speaker's voice, traditional speech recognition models are used for speech recognition and converted into text. Subsequently, sensitive information in the text is masked to obtain the meeting minutes.

[0004] However, this method relies on a microphone array, placing high demands on the equipment and failing to meet these requirements in many scenarios. This results in poor speech separation and, consequently, low transcription accuracy. Furthermore, it is primarily designed for scenarios where a single person is speaking. When multiple people are speaking simultaneously, the inability to accurately identify the voices of individual speakers also affects the accuracy of the transcription results. Furthermore, existing methods further mask sensitive information after transcribing to text. Therefore, the leakage of the transcribed text can lead to the leakage of sensitive information, making it impossible to guarantee information security. Summary of the Invention

[0005] Based on the above-mentioned deficiencies of the existing technology, the present application provides a single-channel voice transcription method and device, electronic device, and storage medium to solve the problem that the existing technology cannot guarantee transcription accuracy and cannot guarantee data security.

[0006] In order to achieve the above objectives, this application provides the following technical solutions:

[0007] The first aspect of the present application provides a single-channel speech transcription method, comprising:

[0008] Collect the current original mixed voice signal of a single channel;

[0009] Performing time-frequency domain preprocessing on the current original mixed speech signal to obtain an amplitude spectrum and a phase spectrum of the current original mixed speech signal;

[0010] After performing non-negative matrix decomposition on the amplitude spectrum of the current original mixed speech signal by optimizing the trained matrix decomposition model, the decomposition result is used to calculate the time-frequency mask, and the calculated time-frequency mask and the amplitude spectrum of the current original mixed speech signal are used to reconstruct the signal, and output a plurality of current separated speech amplitude spectra; wherein, the separated speech amplitude spectra obtained by separating the mixed amplitude spectrum by the matrix decomposition model, the predicted signal group composed of the mixed amplitude spectrum, and the real signal group are input into the neural network model for classification, thereby realizing the joint optimization training of the matrix decomposition model and the neural network model; the real signal group is composed of the real separated speech amplitude spectrum and the mixed amplitude spectrum;

[0011] Using the phase spectrum of the current original mixed speech signal, respectively performing time domain signal conversion on the amplitude spectrum of each current separated speech to obtain each current separated speech;

[0012] For each of the currently separated speech, speech recognition and sensitive word classification are performed using a trained speech transcription model, and after fusing the speech recognition text with the sensitive word classification result, the transcription text of each of the currently separated speech with sensitive words masked is output.

[0013] Optionally, in the above-mentioned single-channel speech transcription method, after performing non-negative matrix decomposition on the amplitude spectrum of the current original mixed speech signal by optimizing the trained matrix decomposition model, a time-frequency mask is calculated using the decomposition result, and the calculated time-frequency mask and the amplitude spectrum of the current original mixed speech signal are used to perform signal reconstruction, and multiple current separated speech amplitude spectra are output, including:

[0014] Inputting the amplitude spectrum of the current original mixed speech signal into the optimized and trained matrix decomposition model;

[0015] Decomposing the amplitude spectrum of the current original mixed speech signal into a plurality of non-negative matrices by a non-negative matrix decomposition algorithm in the matrix decomposition model;

[0016] Calculating a time-frequency mask using each of the non-negative matrices;

[0017] The time-frequency mask is used to perform signal reconstruction calculation on the amplitude spectrum of the current original mixed speech signal to obtain multiple current separated speech amplitude spectra.

[0018] Optionally, in the above-mentioned single-channel speech transcription method, the optimization training method of the matrix decomposition model includes:

[0019] Training the matrix decomposition model using a first source speech to obtain a first speech dictionary, and training the matrix decomposition model using a second source speech to obtain a second speech dictionary;

[0020] connecting the first phonetic dictionary and the second phonetic dictionary into an overall training dictionary;

[0021] Fixing the first phonetic dictionary in the overall training dictionary, and updating the second phonetic dictionary and the overall coefficient matrix in the overall training dictionary by training the matrix decomposition model;

[0022] cyclically utilizing the matrix decomposition model to separate the mixed amplitude spectrum, and inputting the separated speech amplitude spectra, the predicted signal group composed of the mixed amplitude spectrum, and the real signal group into the neural network model for classification;

[0023] Based on the classification result of the neural network model, the matrix decomposition model and the neural network model are alternately updated through a batch gradient descent algorithm.

[0024] Optionally, in the above-mentioned single-channel speech transcription method, for each of the currently separated speech, speech recognition and sensitive word classification are performed using a trained speech transcription model, and after fusing the speech recognition text with the sensitive word classification result, the transcription text of each currently separated speech with sensitive words masked is output, including:

[0025] For each of the currently separated speech, input the currently separated speech into the speech transcription model;

[0026] Preprocessing the current separated speech through a preprocessing layer in the speech transcription model;

[0027] Encoding the preprocessed current separated speech through an encoder in the speech transcription model to obtain encoded data of the current separated speech;

[0028] Decoding the encoded data of the currently separated speech using a text decoder in the speech transcription model to obtain speech recognition text, and performing sensitive word recognition on the encoded data of the currently separated speech using a sensitive word classifier in the speech transcription model to obtain a sensitive word classification result;

[0029] The words indicated by the sensitive word classification result in the speech recognition text are shielded, and a transcription text of the current separated speech is output.

[0030] Optionally, in the above-mentioned single-channel speech transcription method, the training method of the speech transcription model includes:

[0031] Iteratively training a speech recognition model in the speech transcription model; wherein the speech recognition model includes the preprocessing layer, the encoder, and the decoder;

[0032] Iteratively training the sensitive word classifier in the speech transcription model by fixing the parameters of the encoder;

[0033] All parts of the entire speech transcription model are trained jointly.

[0034] A second aspect of the present application provides a single-channel speech transcription device, comprising:

[0035] A voice acquisition unit, used to acquire the current original mixed voice signal of a single channel;

[0036] A time-frequency domain preprocessing unit, configured to perform time-frequency domain preprocessing on the current original mixed speech signal to obtain an amplitude spectrum and a phase spectrum of the current original mixed speech signal;

[0037] An amplitude spectrum separation unit is used to perform non-negative matrix decomposition on the amplitude spectrum of the current original mixed speech signal by optimizing a trained matrix decomposition model, calculate a time-frequency mask using the decomposition result, and reconstruct a signal using the calculated time-frequency mask and the amplitude spectrum of the current original mixed speech signal to output a plurality of current separated speech amplitude spectra; wherein, the separated speech amplitude spectra obtained by separating the mixed amplitude spectrum by the matrix decomposition model, the predicted signal group composed of the mixed amplitude spectrum, and the real signal group are input into a neural network model for classification, thereby realizing joint optimization training of the matrix decomposition model and the neural network model; the real signal group is composed of the real separated speech amplitude spectrum and the mixed amplitude spectrum;

[0038] A speech conversion unit, configured to perform time domain signal conversion on the amplitude spectrum of each current separated speech using the phase spectrum of the current original mixed speech signal to obtain each current separated speech;

[0039] The speech recognition unit is used to perform speech recognition and sensitive word classification for each of the currently separated speech using a trained speech transcription model, and after fusing the speech recognition text with the sensitive word classification result, output the transcription text of each of the currently separated speech with sensitive words masked.

[0040] Optionally, in the above-mentioned single-channel speech transcription device, the amplitude spectrum separation unit includes:

[0041] An amplitude spectrum input unit, configured to input the amplitude spectrum of the current original mixed speech signal into the optimized and trained matrix decomposition model;

[0042] A matrix decomposition unit, configured to decompose the amplitude spectrum of the current original mixed speech signal into a plurality of non-negative matrices by using a non-negative matrix decomposition algorithm in the matrix decomposition model;

[0043] A time-frequency mask calculation unit, configured to calculate a time-frequency mask using each of the non-negative matrices;

[0044] The signal reconstruction unit is used to perform signal reconstruction calculation on the amplitude spectrum of the current original mixed speech signal using the time-frequency mask to obtain multiple current separated speech amplitude spectra.

[0045] Optionally, in the above-mentioned single-channel speech transcription device, the device further includes:

[0046] a dictionary training unit, configured to train the matrix decomposition model using a first source speech to obtain a first speech dictionary, and to train the matrix decomposition model using a second source speech to obtain a second speech dictionary;

[0047] A connection unit, configured to connect the first phonetic dictionary and the second phonetic dictionary into an overall training dictionary;

[0048] a matrix updating unit, configured to fix the first phonetic dictionary in the overall training dictionary and update the second phonetic dictionary and the overall coefficient matrix in the overall training dictionary by training the matrix decomposition model;

[0049] a cyclic training unit, configured to cyclically utilize the matrix decomposition model to separate the mixed amplitude spectrum, and input the separated speech amplitude spectra, the predicted signal group composed of the mixed amplitude spectrum, and the real signal group into the neural network model for classification;

[0050] An alternating updating unit is used to alternately update the matrix decomposition model and the neural network model through a batch gradient descent algorithm based on the classification result of the neural network model.

[0051] Optionally, in the above-mentioned single-channel speech transcription device, the speech recognition unit includes:

[0052] a speech input unit, configured to input each of the currently separated speech into the speech transcription model;

[0053] A speech preprocessing unit, configured to preprocess the current separated speech through a preprocessing layer in the speech transcription model;

[0054] an encoding unit, configured to encode the preprocessed current separated speech through an encoder in the speech transcription model to obtain encoded data of the current separated speech;

[0055] A decoding unit, configured to decode the encoded data of the currently separated speech through a text decoder in the speech transcription model to obtain speech recognition text;

[0056] a vocabulary classification unit, configured to perform sensitive word recognition on the encoded data of the currently separated speech using a sensitive word classifier in the speech transcription model to obtain a sensitive word classification result;

[0057] The shielding unit is used to shield the words indicated by the sensitive word classification result in the speech recognition text and output the transcription text of the current separated speech.

[0058] Optionally, in the above-mentioned single-channel speech transcription device, the device further includes:

[0059] A speech recognition model training unit, configured to iteratively train the speech recognition model in the speech transcription model; wherein the speech recognition model includes the preprocessing layer, the encoder, and the decoder;

[0060] A sensitive word classifier training unit, configured to iteratively train the sensitive word classifier in the speech transcription model by fixing the parameters of the encoder;

[0061] The joint training unit is used to jointly train all parts of the entire speech transcription model.

[0062] A third aspect of the present application provides an electronic device, including:

[0063] memory and processor;

[0064] Wherein, the memory is used to store programs;

[0065] The processor is used to execute the program, and when the program is executed, it is specifically used to implement the single-channel speech transcription method as described in any one of the above.

[0066] In a fourth aspect, the present application provides a computer storage medium for storing a computer program, wherein when the computer program is executed by a processor, the computer storage medium is used to implement the single-channel speech transcription method as described in any one of the above.

[0067] The present application provides a single-channel speech transcription method, which collects the current original mixed speech signal of a single channel. The current original mixed speech signal is then preprocessed in the time-frequency domain to obtain the amplitude spectrum and phase spectrum of the current original mixed speech signal, so as to perform speech separation through time-frequency domain analysis. Then, the amplitude spectrum of the current original mixed speech signal is subjected to non-negative matrix decomposition by optimizing the trained matrix decomposition model. The decomposition result is used to calculate the time-frequency mask, and the calculated time-frequency mask and the amplitude spectrum of the current original mixed speech signal are used to reconstruct the signal, outputting multiple current separated speech amplitude spectra. Among them, the individual separated speech amplitude spectra obtained by separating the mixed amplitude spectrum by the matrix decomposition model, the predicted signal group composed of the mixed amplitude spectrum, and the real signal group are input into the neural network model for classification, thereby achieving joint optimization training of the matrix decomposition model and the neural network model. The real signal group consists of the real separated speech amplitude spectrum and the mixed amplitude spectrum, so that the joint optimization training of the neural network model and the matrix decomposition model can effectively ensure that the matrix decomposition model can accurately separate the amplitude spectrum. In addition, the mixed amplitude spectrum is also considered in the neural network, which can effectively alleviate the over-smoothing effect on the predicted speech, thereby improving the separation effect. The phase spectrum of the original mixed speech signal is then used to perform time-domain signal conversion on the amplitude spectrum of each currently separated speech, yielding the currently separated speech. The separated amplitude spectrum is then converted back into speech, achieving speech separation. Finally, for each currently separated speech, speech recognition and sensitive word classification are performed using a trained speech transcription model. After fusing the speech recognition text with the sensitive word classification results, the transcribed text for each currently separated speech is output, masking sensitive words. By adding a compliance check design, sensitive information is masked during the speech recognition process, effectively ensuring its security. This results in a single-channel, accurate speech separation method that does not rely on a microphone array and can separate the speech of multiple speakers, effectively ensuring transcription accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without any creative work.

[0069] Figure 1 A flowchart of a single-channel speech transcription method provided in an embodiment of the present application;

[0070] Figure 2 A schematic diagram of the architecture of a matrix decomposition model provided in an embodiment of the present application;

[0071] Figure 3 A schematic diagram of a method for separating speech amplitude spectrum provided in an embodiment of the present application;

[0072] Figure 4 A flowchart of an optimization training method for a matrix decomposition model provided in an embodiment of the present application;

[0073] Figure 5 A schematic diagram of the architecture of a neural network model provided in an embodiment of the present application;

[0074] Figure 6 A schematic diagram of the architecture of a speech transcription model provided in an embodiment of the present application;

[0075] Figure 7 A flowchart of a method for speech recognition provided in an embodiment of the present application;

[0076] Figure 8 A flowchart of a method for training a speech transcription model provided in an embodiment of the present application;

[0077] Figure 9 A schematic diagram of the architecture of a single-channel speech transcription device provided in an embodiment of the present application;

[0078] Figure 10 A schematic diagram of the architecture of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0079] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0080] In this application, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus comprising the element.

[0081] The present application embodiment provides a single-channel speech transcription method, such as Figure 1 As shown, the following steps are included:

[0082] S101: Collect the current original mixed voice signal of a single channel.

[0083] It should be noted that in the embodiment of the present application, it does not rely on a multi-microphone array, so it is only necessary to collect single-channel voice information through a voice collection device, and the directly collected signal is usually a mixed voice signal.

[0084] Optionally, the current original mixed speech signal of a single channel may be collected in real time for transcription, or the current original mixed speech signal of a single channel may be collected throughout the entire process and then transcribed.

[0085] S102 : performing time-frequency domain preprocessing on the current original mixed speech signal to obtain an amplitude spectrum and a phase spectrum of the current original mixed speech signal.

[0086] It should be noted that, due to the short-term stability of the speech signal, the time-frequency domain analysis of the speech signal can well combine the advantages of the time domain and the frequency domain for analysis, and in the case of single-channel input, not all source signals are aliased at each time-frequency point, so at certain time points there is almost only one or a few signal distributions, so that the source signal can be easily separated based on the time-frequency domain. Therefore, in an embodiment of the present application, the current original mixed speech signal is preprocessed in the time-frequency domain to obtain the amplitude spectrum and phase spectrum of the current original mixed speech signal, so that subsequent separation based on the amplitude spectrum can be performed.

[0087] Optionally, a short-time Fourier transform (STFT) can be used to perform time-frequency domain preprocessing on the current original mixed speech signal. Specifically, the input current original mixed speech signal, that is, the input time-frequency domain signal, is segmented into time frames. Then, a short-time Fourier transform is performed on each frame of the signal to obtain STFT coefficients. Finally, a modulo operation is performed on the STFT coefficients to obtain the STFT amplitude spectrum and phase spectrum.

[0088] S103. After performing non-negative matrix decomposition on the amplitude spectrum of the current original mixed speech signal by optimizing the trained matrix decomposition model, the decomposition result is used to calculate the time-frequency mask, and the calculated time-frequency mask and the amplitude spectrum of the current original mixed speech signal are used to reconstruct the signal, and multiple current separated speech amplitude spectra are output.

[0089] It should be noted that, in the embodiment of the present application, the matrix of the amplitude spectrum of the mixed speech signal is decomposed by non-negative matrix decomposition to achieve separation of the mixed speech signal.

[0090] Specifically, the amplitude spectrum of the original mixed speech signal is subjected to non-negative matrix factorization (NMF) by optimizing a trained matrix factorization model, decomposing it into the product of multiple non-negative matrices. For example, given a non-negative matrix V (m×n), the NMF algorithm is used to find two non-negative matrices W (m×k) and H (k×n) such that V ≈ WH. W is called the speech dictionary.

[0091] After obtaining the decomposition results, a corresponding time-frequency mask is calculated using the decomposed non-negative matrix to separate the amplitude spectra of each speech signal in the mixed speech signal. This time-frequency mask is a sentence with the same characteristics as the time-frequency of the signal, and its element values are generally between 0 and 1. Using this time-frequency mask, the amplitude spectrum of the current original mixed speech signal can be reconstructed, outputting multiple amplitude spectra of the current separated speech signals.

[0092] So specifically, the architecture of the matrix decomposition model can be as follows Figure 2 As shown, it includes non-negative matrix factorization, time-frequency masking and signal reconstruction.

[0093] It should also be noted that in order to ensure the accuracy of the matrix decomposition model in separating the mixed amplitude spectrum, in an embodiment of the present application, a neural network model and a matrix decomposition model are jointly optimized and trained. The neural network is mainly used to identify whether the input speech group is a real speech group or a speech group of the matrix decomposition model. Therefore, specifically in the training process, the separated speech amplitude spectra obtained by separating the mixed amplitude spectrum by the matrix decomposition model, the predicted signal group composed of the mixed amplitude spectrum, and the real signal group are input into the neural network model for classification, thereby realizing the joint optimization training of the matrix decomposition model and the neural network model.

[0094] The mixed amplitude spectrum refers to the amplitude spectrum of the mixed speech. The real signal group consists of the real separated speech amplitude spectrum and the mixed amplitude spectrum. The real separated speech amplitude spectrum can be the amplitude spectrum of an independent speech, or the amplitude spectrum of a speech that can be accurately separated.

[0095] Specifically, through training, the matrix decomposition model separates the mixed amplitude spectrum, resulting in separate speech amplitude spectra, and the predicted signal group composed of the mixed amplitude spectra. This input is then fed into the neural network model, preventing the neural network from analyzing the predicted signal group as a function of the matrix decomposition model. However, training the neural network model allows it to accurately analyze the speech group composed of the real signal group and the data separated by the matrix decomposition model. Therefore, by jointly optimizing and training the two, the divergence between the distribution differences between the matrix decomposition model-generated and the real speech quality inspection data can be minimized, thereby ensuring the accuracy of the matrix decomposition model's decomposition results.

[0096] Optionally, in another embodiment of the present application, a specific implementation of step S103 is as follows: Figure 3 Shown, including:

[0097] S301: Input the amplitude spectrum of the current original mixed speech signal into the optimized and trained matrix decomposition model.

[0098] S302 : Decompose the amplitude spectrum of the current original mixed speech signal into multiple non-negative matrices using a non-negative matrix decomposition algorithm in a matrix decomposition model.

[0099] Specifically, the amplitude spectrum of the current original mixed speech signal is decomposed into multiple non-negative matrices whose multiplication is equal to the matrix of the amplitude spectrum of the current original mixed speech signal by the non-negative matrix decomposition algorithm. For example, the amplitude spectrum X of the current original mixed speech signal is input, and the non-negative matrix decomposition algorithm is used to decompose it to obtain two non-negative matrices and .

[0100] S303: Calculate a time-frequency mask using each non-negative matrix.

[0101] Specifically, the time-frequency mask is calculated as follows:

[0102]

[0103] Where f=1,2,...,F, represents different frequencies.

[0104] S304: Use the time-frequency mask to perform signal reconstruction calculation on the amplitude spectrum of the current original mixed speech signal to obtain multiple current separated speech amplitude spectra.

[0105] Specifically, by using the time-frequency mask to predict and separate the amplitude spectrum of the current original mixed speech signal, two current separated speech amplitude spectra can be obtained. Specifically, it can be expressed as:

[0106]

[0107]

[0108] Optionally, in another embodiment of the present application, a matrix decomposition model optimization training method is provided, such as Figure 4 As shown, the following steps are included:

[0109] S401: Using a first source speech to train a matrix decomposition model to obtain a first speech dictionary, and using a second source speech to train the matrix decomposition model to obtain a second speech dictionary.

[0110] In order to enable the matrix decomposition model to have basic decomposition capabilities, it is necessary to first initialize the speech dictionary W and the overall coefficient matrix H by randomly generating non-negative numbers. Then, using the clean first source speech 1 dataset, that is, non-mixed speech, separate training is performed to obtain the speech dictionary W(1), and the clean second source speech 2 training set is used to separate training to obtain the speech dictionary W(2).

[0111] S402: Connect the first phonetic dictionary and the second phonetic dictionary into an overall training dictionary.

[0112] S403 , fixing the first phonetic dictionary in the overall training dictionary, and updating the second phonetic dictionary and the overall coefficient matrix in the overall training dictionary by training the matrix decomposition model.

[0113] S404, cyclically utilizing the matrix decomposition model to separate the mixed amplitude spectrum, and inputting the separated speech amplitude spectra, the prediction signal group consisting of the mixed amplitude spectrum, and the real signal group into the neural network model for classification.

[0114] Specifically, a matrix decomposition model is used to separate the mixed amplitude spectrum to obtain separate speech amplitude spectra. The separate speech amplitude spectra and the mixed amplitude spectrum are then combined to form a prediction signal group. The predicted signal group and the true signal group are then input into a neural network model for classification.

[0115] Optionally, when the input signal is a real signal group, the neural network model outputs 1, and when the current input signal is a predicted signal group, the neural network model outputs 0.

[0116] Specifically, such as Figure 5 As shown in the figure, the neural network model includes an input layer, a convolutional layer, a feature layer, and an output layer.

[0117] S405. Based on the classification result of the neural network model, the matrix decomposition model and the neural network model are alternately updated through a batch gradient descent algorithm.

[0118] Optionally, the model may be replaced and updated every time a set number of trainings are performed. When one model is updated, the parameters of the other model remain unchanged.

[0119] It should be noted that the goal of the neural network unit is to minimize the divergence of the distribution difference between the real and predicted speech, and the present invention adds an additional mixed speech signal spectrum X to the neural network unit during training. Therefore, the method proposed in the present invention can effectively alleviate the over-smoothing effect on the predicted speech, thereby improving the separation effect.

[0120] S104 , using the phase spectrum of the current original mixed speech signal, respectively perform time domain signal conversion on the amplitude spectrum of each current separated speech to obtain each current separated speech.

[0121] Since the separation obtained at this time is the amplitude spectrum, and what is ultimately needed is the separation of the separation, it is necessary to use the phase spectrum of the current original mixed speech signal to restore the amplitude spectrum of each current separated speech to time domain information, so as to restore it to speech.

[0122] Optionally, the phase spectrum may be specifically used to convert it into a time domain signal through inverse short-time Fourier transform, thereby obtaining each current separated speech.

[0123] S105. For each currently separated speech, speech recognition and sensitive word classification are performed using a trained speech transcription model. After fusing the speech recognition text with the sensitive word classification result, the transcription text of each currently separated speech with sensitive words masked is output.

[0124] It should be noted that in order to be able to detect and shield sensitive words, the process is moved from the post-transcription stage to the recognition stage, thereby effectively ensuring the security of sensitive information. Therefore, the speech transcription model in the embodiment of the present application can not only transcribe the speaker's speech signal in real time, but also increase compliance detection and involve shielding sensitive information when generating text. Therefore, speech recognition and sensitive word classification are performed specifically through the trained speech transcription model, and after the speech recognition text is fused with the sensitive word classification result, the transcribed text of each currently separated speech with the sensitive words shielded is output. Moreover, since the input speech signal is a pure speech signal after separation, that is, the signal removes the interference of noise and other aliased speech, thereby ensuring that the accuracy of speech recognition is higher.

[0125] Optionally, in another embodiment of the present application, a sensitive word classifier is added to the traditional automatic speech recognition (ASR) model to form a speech transcription model. Figure 6 As shown, a speech transcription model provided in an embodiment of the present application includes a speech preprocessing layer, an encoder, a text decoder, and a sensitive word classifier. Accordingly, in another embodiment of the present application, a specific implementation of step S105 is as follows: Figure 7 Shown, including:

[0126] S701: For each currently separated speech, input the currently separated speech into a speech transcription model.

[0127] S702: Preprocess the current separated speech through the preprocessing layer in the speech transcription model.

[0128] S703 : Encode the pre-processed current separated speech through an encoder in the speech transcription model to obtain encoded data of the current separated speech.

[0129] S704. Decode the encoded data of the current separated speech through the text decoder in the speech transcription model to obtain speech recognition text, and perform sensitive word recognition on the encoded data of the current separated speech through the sensitive word classifier in the speech transcription model to obtain a sensitive word classification result.

[0130] S705: Shield the words indicated by the sensitive word classification result in the speech recognition text, and output the transcribed text of the current separated speech.

[0131] Optionally, in another embodiment of the present application, a method for training a speech transcription model is provided, such as Figure 8 Shown, including:

[0132] S801. Iteratively train the speech recognition model in the speech transcription model.

[0133] Among them, the speech recognition model includes a preprocessing layer, an encoder and a decoder.

[0134] Specifically, in order to enable the speech recognition model to accurately perform speech recognition, the sensitive word classifier is first removed and the speech recognition model is iteratively trained.

[0135] S802: Iteratively train the sensitive word classifier in the speech transcription model by fixing the parameters of the encoder.

[0136] Similarly, in order to enable the sensitive word classifier to accurately classify vocabulary and identify sensitive words, it is necessary to perform separate iterative training on the sensitive word classifier in the speech transcription model.

[0137] S803: Jointly train all parts of the entire speech transcription model.

[0138] In order for the speech recognition model and the sensitive word classifier to work effectively together, it is necessary to jointly train all parts of the entire speech transcription model and fine-tune all parameters.

[0139] The present application provides a single-channel speech transcription method that collects a single-channel current original mixed speech signal. The current original mixed speech signal is then preprocessed in the time-frequency domain to obtain the amplitude spectrum and phase spectrum of the current original mixed speech signal, so as to perform speech separation through time-frequency domain analysis. The amplitude spectrum of the current original mixed speech signal is then subjected to non-negative matrix decomposition by optimizing a trained matrix decomposition model. The decomposition result is used to calculate a time-frequency mask, and the calculated time-frequency mask and the amplitude spectrum of the current original mixed speech signal are used to reconstruct the signal, outputting multiple current separated speech amplitude spectra. The matrix decomposition model is then used to separate the mixed amplitude spectrum, and the predicted signal group consisting of the mixed amplitude spectrum and the real signal group are input into a neural network model for classification, thereby achieving joint optimization training of the matrix decomposition model and the neural network model. The real signal group consists of the real separated speech amplitude spectrum and the mixed amplitude spectrum, so that the neural network model and the matrix decomposition model are jointly optimized and trained, which can effectively ensure that the matrix decomposition model can accurately separate the amplitude spectrum. In addition, the mixed amplitude spectrum is also considered in the neural network, which can effectively alleviate the over-smoothing effect on the predicted speech, thereby improving the separation effect. The phase spectrum of the original mixed speech signal is then used to perform time-domain signal conversion on the amplitude spectrum of each currently separated speech, yielding the currently separated speech. The separated amplitude spectrum is then converted back into speech, achieving speech separation. Finally, for each currently separated speech, speech recognition and sensitive word classification are performed using a trained speech transcription model. After fusing the speech recognition text with the sensitive word classification results, the transcribed text for each currently separated speech is output, masking sensitive words. By adding a compliance check design, sensitive information is masked during the speech recognition process, effectively ensuring its security. This results in a single-channel, accurate speech separation method that does not rely on a microphone array and can separate the speech of multiple speakers, effectively ensuring transcription accuracy.

[0140] Another embodiment of the present application provides a single-channel speech transcription device, such as Figure 9 Shown, including:

[0141] The voice collection unit 901 is used to collect the current original mixed voice signal of a single channel.

[0142] The time-frequency domain preprocessing unit 902 is configured to perform time-frequency domain preprocessing on the current original mixed speech signal to obtain an amplitude spectrum and a phase spectrum of the current original mixed speech signal.

[0143] The amplitude spectrum separation unit 903 is used to perform non-negative matrix decomposition on the amplitude spectrum of the current original mixed speech signal by optimizing the trained matrix decomposition model, use the decomposition result to calculate the time-frequency mask, and use the calculated time-frequency mask and the amplitude spectrum of the current original mixed speech signal to reconstruct the signal, and output multiple current separated speech amplitude spectra.

[0144] The matrix decomposition model is used to separate the mixed amplitude spectrum, resulting in separate speech amplitude spectra. The predicted signal group, composed of the mixed amplitude spectrum, and the real signal group are then input into the neural network model for classification, achieving joint optimization training of the matrix decomposition model and the neural network model. The real signal group consists of the real separated speech amplitude spectra and the mixed amplitude spectrum.

[0145] The speech conversion unit 904 is configured to perform time domain signal conversion on the amplitude spectrum of each current separated speech using the phase spectrum of the current original mixed speech signal to obtain each current separated speech.

[0146] The speech recognition unit 905 is used to perform speech recognition and sensitive word classification for each currently separated speech through a trained speech transcription model, and after fusing the speech recognition text with the sensitive word classification result, output the transcription text of each currently separated speech with sensitive words masked.

[0147] Optionally, in the single-channel speech transcription device provided in another embodiment of the present application, the amplitude spectrum separation unit includes:

[0148] The amplitude spectrum input unit is used to input the amplitude spectrum of the current original mixed speech signal into the optimized and trained matrix decomposition model.

[0149] The matrix decomposition unit is used to decompose the amplitude spectrum of the current original mixed speech signal into multiple non-negative matrices through a non-negative matrix decomposition algorithm in a matrix decomposition model.

[0150] The time-frequency mask calculation unit is used to calculate the time-frequency mask using each non-negative matrix.

[0151] The signal reconstruction unit is used to perform signal reconstruction calculation on the amplitude spectrum of the current original mixed speech signal by using the time-frequency mask to obtain multiple current separated speech amplitude spectra.

[0152] Optionally, in another embodiment of the present application, the single-channel speech transcription device further includes:

[0153] The dictionary training unit is used to train the matrix decomposition model using the first source speech to obtain a first speech dictionary, and to train the matrix decomposition model using the second source speech to obtain a second speech dictionary.

[0154] The connection unit is used to connect the first phonetic dictionary and the second phonetic dictionary into an overall training dictionary.

[0155] The matrix updating unit is used to fix the first phonetic dictionary in the overall training dictionary and update the second phonetic dictionary and the overall coefficient matrix in the overall training dictionary by training the matrix decomposition model.

[0156] The cyclic training unit is used to cyclically use the matrix decomposition model to separate the mixed amplitude spectrum, and input the separated speech amplitude spectra, the prediction signal group composed of the mixed amplitude spectrum, and the real signal group into the neural network model for classification.

[0157] The alternating update unit is used to alternately update the matrix decomposition model and the neural network model through a batch gradient descent algorithm based on the classification results of the neural network model.

[0158] Optionally, in a single-channel speech transcription device provided in another embodiment of the present application, the speech recognition unit includes:

[0159] The speech input unit is used to input the current separated speech into the speech transcription model for each current separated speech.

[0160] The speech preprocessing unit is used to preprocess the current separated speech through the preprocessing layer in the speech transcription model.

[0161] The encoding unit is used to encode the preprocessed current separated speech through the encoder in the speech transcription model to obtain encoded data of the current separated speech.

[0162] The decoding unit is used to decode the encoded data of the current separated speech through the text decoder in the speech transcription model to obtain the speech recognition text.

[0163] The vocabulary classification unit is used to identify sensitive words in the encoded data of the current separated speech through the sensitive word classifier in the speech transcription model to obtain a sensitive word classification result.

[0164] The shielding unit is used to shield the words indicated by the sensitive word classification results in the speech recognition text and output the transcribed text of the current separated speech.

[0165] Optionally, in another embodiment of the present application, the single-channel speech transcription device further includes:

[0166] The speech recognition model training unit is used to iteratively train the speech recognition model in the speech transcription model. The speech recognition model includes a preprocessing layer, an encoder, and a decoder.

[0167] The sensitive word classifier training unit is used to iteratively train the sensitive word classifier in the speech transcription model by fixing the parameters of the encoder.

[0168] The joint training unit is used to jointly train all parts of the entire speech transcription model.

[0169] It should be noted that the specific working process of each unit provided in the above embodiments of the present application can refer to the implementation process of the corresponding steps in the above method embodiments, and will not be repeated here.

[0170] Another embodiment of the present application provides an electronic device, such as Figure 10 Shown, including:

[0171] Memory 1001 and processor 1002 .

[0172] The memory 1001 is used to store programs.

[0173] The processor 1002 is used to execute the program stored in the memory 1001. When the program is executed, it is specifically used to implement the single-channel speech transcription method provided by any one of the above embodiments.

[0174] Another embodiment of the present application provides a computer storage medium for storing a computer program. When the computer program is executed by a processor, it is used to implement the single-channel speech transcription method as described in any of the above embodiments.

[0175] Computer storage media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0176] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0177] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A single-channel speech transcription method, characterized in that: include: Collect the current original mixed voice signal of a single channel; Performing time-frequency domain preprocessing on the current original mixed speech signal to obtain an amplitude spectrum and a phase spectrum of the current original mixed speech signal; After performing non-negative matrix decomposition on the amplitude spectrum of the current original mixed speech signal by optimizing the trained matrix decomposition model, the decomposition result is used to calculate the time-frequency mask, and the calculated time-frequency mask and the amplitude spectrum of the current original mixed speech signal are used to reconstruct the signal, and output a plurality of current separated speech amplitude spectra; wherein, the separated speech amplitude spectra obtained by separating the mixed amplitude spectrum by the matrix decomposition model, the predicted signal group composed of the mixed amplitude spectrum, and the real signal group are input into the neural network model for classification, thereby realizing the joint optimization training of the matrix decomposition model and the neural network model; the real signal group is composed of the real separated speech amplitude spectrum and the mixed amplitude spectrum; Using the phase spectrum of the current original mixed speech signal, respectively performing time domain signal conversion on the amplitude spectrum of each current separated speech to obtain each current separated speech; For each of the currently separated speech, speech recognition and sensitive word classification are performed using a trained speech transcription model, and after fusing the speech recognition text with the sensitive word classification result, the transcription text of each of the currently separated speech with sensitive words masked is output.

2. The method according to claim 1, characterized in that After performing non-negative matrix decomposition on the amplitude spectrum of the current original mixed speech signal by optimizing the trained matrix decomposition model, a time-frequency mask is calculated using the decomposition result, and the calculated time-frequency mask and the amplitude spectrum of the current original mixed speech signal are used to perform signal reconstruction, and multiple current separated speech amplitude spectra are output, including: Inputting the amplitude spectrum of the current original mixed speech signal into the optimized and trained matrix decomposition model; Decomposing the amplitude spectrum of the current original mixed speech signal into a plurality of non-negative matrices by a non-negative matrix decomposition algorithm in the matrix decomposition model; Calculating a time-frequency mask using each of the non-negative matrices; The time-frequency mask is used to perform signal reconstruction calculation on the amplitude spectrum of the current original mixed speech signal to obtain multiple current separated speech amplitude spectra.

3. The method according to claim 1, characterized in that The optimization training method of the matrix decomposition model includes: Training the matrix decomposition model using a first source speech to obtain a first speech dictionary, and training the matrix decomposition model using a second source speech to obtain a second speech dictionary; connecting the first phonetic dictionary and the second phonetic dictionary into an overall training dictionary; Fixing the first phonetic dictionary in the overall training dictionary, and updating the second phonetic dictionary and the overall coefficient matrix in the overall training dictionary by training the matrix decomposition model; cyclically utilizing the matrix decomposition model to separate the mixed amplitude spectrum, and inputting the separated speech amplitude spectra, the predicted signal group composed of the mixed amplitude spectrum, and the real signal group into the neural network model for classification; Based on the classification result of the neural network model, the matrix decomposition model and the neural network model are alternately updated through a batch gradient descent algorithm.

4. The method according to claim 1, wherein The method of performing speech recognition and sensitive word classification on each of the currently separated speech using a trained speech transcription model, fusing the speech recognition text with the sensitive word classification results, and outputting the transcription text of each currently separated speech with sensitive words masked includes: For each of the currently separated speech, input the currently separated speech into the speech transcription model; Preprocessing the current separated speech through a preprocessing layer in the speech transcription model; Encoding the preprocessed current separated speech through an encoder in the speech transcription model to obtain encoded data of the current separated speech; Decoding the encoded data of the currently separated speech using a text decoder in the speech transcription model to obtain speech recognition text, and performing sensitive word recognition on the encoded data of the currently separated speech using a sensitive word classifier in the speech transcription model to obtain a sensitive word classification result; The words indicated by the sensitive word classification result in the speech recognition text are shielded, and a transcription text of the current separated speech is output.

5. The method according to claim 4, characterized in that The training method of the speech transcription model includes: Iteratively training a speech recognition model in the speech transcription model; wherein the speech recognition model includes the preprocessing layer, the encoder, and the decoder; Iteratively training the sensitive word classifier in the speech transcription model by fixing the parameters of the encoder; All parts of the entire speech transcription model are trained jointly.

6. A single-channel speech transcription device, characterized in that: include: A voice acquisition unit, used to acquire the current original mixed voice signal of a single channel; A time-frequency domain preprocessing unit, configured to perform time-frequency domain preprocessing on the current original mixed speech signal to obtain an amplitude spectrum and a phase spectrum of the current original mixed speech signal; An amplitude spectrum separation unit is used to perform non-negative matrix decomposition on the amplitude spectrum of the current original mixed speech signal by optimizing a trained matrix decomposition model, calculate a time-frequency mask using the decomposition result, and reconstruct a signal using the calculated time-frequency mask and the amplitude spectrum of the current original mixed speech signal to output a plurality of current separated speech amplitude spectra; wherein, the separated speech amplitude spectra obtained by separating the mixed amplitude spectrum by the matrix decomposition model, the predicted signal group composed of the mixed amplitude spectrum, and the real signal group are input into a neural network model for classification, thereby realizing joint optimization training of the matrix decomposition model and the neural network model; the real signal group is composed of the real separated speech amplitude spectrum and the mixed amplitude spectrum; A speech conversion unit is used to perform time domain signal conversion on the amplitude spectrum of each current separated speech by using the phase spectrum of the current original mixed speech signal to obtain each current separated speech; The speech recognition unit is used to perform speech recognition and sensitive word classification for each of the currently separated speech using a trained speech transcription model, and after fusing the speech recognition text with the sensitive word classification result, output the transcription text of each of the currently separated speech with sensitive words masked.

7. The device according to claim 6, characterized in that The amplitude spectrum separation unit comprises: An amplitude spectrum input unit, configured to input the amplitude spectrum of the current original mixed speech signal into the optimized and trained matrix decomposition model; A matrix decomposition unit, configured to decompose the amplitude spectrum of the current original mixed speech signal into a plurality of non-negative matrices by using a non-negative matrix decomposition algorithm in the matrix decomposition model; A time-frequency mask calculation unit, configured to calculate a time-frequency mask using each of the non-negative matrices; The signal reconstruction unit is used to perform signal reconstruction calculation on the amplitude spectrum of the current original mixed speech signal using the time-frequency mask to obtain multiple current separated speech amplitude spectra.

8. The device according to claim 6, characterized in that Also includes: a dictionary training unit, configured to train the matrix decomposition model using a first source speech to obtain a first speech dictionary, and to train the matrix decomposition model using a second source speech to obtain a second speech dictionary; A connection unit, configured to connect the first phonetic dictionary and the second phonetic dictionary into an overall training dictionary; a matrix updating unit, configured to fix the first phonetic dictionary in the overall training dictionary and update the second phonetic dictionary and the overall coefficient matrix in the overall training dictionary by training the matrix decomposition model; a cyclic training unit, configured to cyclically utilize the matrix decomposition model to separate the mixed amplitude spectrum, and input the separated speech amplitude spectra, the predicted signal group composed of the mixed amplitude spectrum, and the real signal group into the neural network model for classification; An alternating updating unit is used to alternately update the matrix decomposition model and the neural network model through a batch gradient descent algorithm based on the classification result of the neural network model.

9. An electronic device, characterized in that: include: memory and processor; Wherein, the memory is used to store programs; The processor is used to execute the program, and when the program is executed, it is specifically used to implement the single-channel speech transcription method according to any one of claims 1 to 5.

10. A computer storage medium, characterized in that Used to store a computer program, which, when executed by a processor, is used to implement the single-channel speech transcription method according to any one of claims 1 to 5.