Speech recognition method, system and device

By combining an explicit music harmonic convolutional encoder and an implicit music-aware encoder, the problem of background music noise interfering with the speech recognition system is solved, improving the accuracy and robustness of speech recognition and adapting to different music interference scenarios.

CN121583261BActive Publication Date: 2026-04-10TIANJIN UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TIANJIN UNIV
Filing Date
2026-01-27
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Background music noise reduces the accuracy of speech recognition systems, especially when background music is mixed with target speech. Traditional methods struggle to accurately distinguish harmonic frequency characteristics from the spectral distribution of human voices, increasing the difficulty of speech separation and recognition. Furthermore, lyrics embedded in background music may be misidentified as valid speech signals, interfering with semantic understanding and text decoding.

Method used

An explicit music harmonic convolutional encoder is used to extract the scale and harmonic structure features of background music noise. This is combined with an implicit music-aware encoder for speech enhancement and noise suppression. A gated attention fusion processor is used to generate a speech separation mask. Finally, a noise-aware attention mechanism and a convolutional self-attention hybrid encoder are used to adaptively suppress residual music noise, thereby improving recognition accuracy.

Benefits of technology

It effectively separates background music noise, improves the accuracy of speech recognition, reduces the error rate, enhances scene adaptability, adapts to different music interference scenarios, prevents over-suppression or under-suppression, and improves the robustness of the speech recognition system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121583261B_ABST
    Figure CN121583261B_ABST
Patent Text Reader

Abstract

The application provides a speech recognition method, system and device, which can be applied to the technical field of speech processing. The speech recognition method comprises the following steps: in response to a speech recognition instruction, obtaining original audio data, wherein the original audio data comprises a speech signal with background music noise; performing time-frequency conversion on the speech signal to obtain initial spectral features; inputting the initial spectral features into a pre-trained speech separation module for background music suppression to output target spectral features; and inputting the target spectral features and the initial spectral features into a pre-trained speech recognition module to output recognized text; wherein the speech separation module comprises an explicit music harmonic convolutional encoder, an implicit music perception encoder and a gated attention fusion processor, and the speech recognition module comprises a noise perception attention mechanism and a convolutional self-attention hybrid encoder.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of speech processing, and more particularly to a speech recognition method, system and device. BACKGROUND

[0002] In a speech recognition task, the ubiquitous existence of background music noise poses a severe challenge to the application performance of a speech recognition system in an actual scene. Background music is a highly structured interference signal, which contains not only complex harmonic components, but also human voice lyrics with semantic information. This uniqueness makes its interference on a speech recognition system far beyond general noise, resulting in a low recognition accuracy of the speech recognition system. SUMMARY

[0003] In view of the above problems, the present application provides a speech recognition method, system and device for improving recognition accuracy.

[0004] According to a first aspect of the present application, a speech recognition method is provided, the speech recognition method comprising: in response to a speech recognition instruction, obtaining original audio data, the original audio data comprising a speech signal with background music noise; performing time-frequency transformation on the speech signal to obtain initial spectral features; inputting the initial spectral features into a pre-trained speech separation module for background music suppression to output target spectral features; inputting the target spectral features and the initial spectral features into a pre-trained speech recognition module to output a recognized text; wherein the speech separation module comprises an explicit music harmonic convolutional encoder, an implicit music perception encoder and a gated attention fusion processor, the explicit music harmonic convolutional encoder is configured to extract scale and harmonic structure features of the background music noise in the initial spectral features to determine frame-level music features, the implicit music perception encoder is configured to perform speech enhancement and noise suppression on the initial spectral features to output frame-level speech features with the same dimension as the frame-level music features, and the gated attention fusion processor is configured to perform gated fusion on the frame-level music features and the frame-level speech features to output gated fusion features; the gated fusion features are used to generate a speech separation mask, and the speech separation mask is used to separate the initial spectral features to obtain the target spectral features; the speech recognition module comprises a noise perception attention mechanism and a convolutional self-attention hybrid encoder, the noise perception attention mechanism fuses the target spectral features and the initial spectral features through learnable gating weights to adaptively suppress residual music noise; and the convolutional self-attention hybrid encoder is configured to perform sequence modeling on the fused features to obtain the recognized text.

[0005] The second aspect of the present application provides a speech recognition system, comprising: an acquisition module configured to acquire original audio data in response to a speech recognition instruction, the original audio data comprising a speech signal with background music noise; a time-frequency conversion module configured to perform time-frequency conversion on the speech signal to obtain initial spectral features; a target spectral feature determination module configured to input the initial spectral features into a pre-trained speech separation module to suppress the background music, so as to output target spectral features; and a text recognition module configured to input the target spectral features and the initial spectral features into a pre-trained speech recognition module to output recognized text; wherein the speech separation module comprises an explicit music harmonic convolutional encoder, an implicit music perception encoder, and a gated attention fusion processor, the explicit music harmonic convolutional encoder is configured to extract scale and harmonic structure features of the background music noise in the initial spectral features to determine frame-level music features, the implicit music perception encoder is configured to perform speech enhancement and noise suppression on the initial spectral features to output frame-level speech features with the same dimension as the frame-level music features, and the gated attention fusion processor is configured to perform gated fusion on the frame-level music features and the frame-level speech features to output gated fusion features; the gated fusion features are used to generate a speech separation mask, and the speech separation mask is used to separate the initial spectral features to obtain the target spectral features; the speech recognition module comprises a noise perception attention mechanism and a convolutional self-attention hybrid encoder, the noise perception attention mechanism fuses the target spectral features and the initial spectral features through learnable gating weights to adaptively suppress residual music noise, and the convolutional self-attention hybrid encoder is configured to perform sequence modeling on the fused features to obtain the recognized text.

[0006] The third aspect of the present application provides an electronic device, comprising: one or more processors; and a memory configured to store one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the speech recognition method.

[0007] The fourth aspect of the present application further provides a computer-readable storage medium having a computer program or instructions stored thereon, wherein the computer program or instructions are executed by a processor to implement the steps of the speech recognition method.

[0008] The fifth aspect of the present application further provides a computer program product comprising a computer program or instructions, wherein the computer program or instructions are executed by a processor to implement the steps of the speech recognition method.

[0009] According to the embodiments of the present application, the pitch and harmonic structure features of the background music noise are extracted by the explicit music harmonic convolution encoder, and the initial spectral features are subjected to speech enhancement and noise suppression by the implicit music perception encoder, so that the background music noise of the speech signal can be separated, the high-purity speech features can be extracted, and the recognition accuracy of the recognized text can be improved. By online correcting the frequency response of the filter bank, the speech dominant frequency band can be made more regular, and the music dominant frequency band can be made more controllable, further improving the recognition accuracy of the recognized text. By adaptively suppressing the residual music noise and performing sequence modeling on the fused features, the residual noise can be accurately suppressed without affecting the spectral features of the target speech, further improving the recognition accuracy of the recognized text, reducing the error rate, and enhancing the scene adaptability, which can adapt to different music interference scenes to prevent "over suppression" or "insufficient suppression". BRIEF DESCRIPTION OF DRAWINGS

[0010] The above and other objects, features and advantages of the present application will become more apparent from the following description when taken in conjunction with the accompanying drawings, in which:

[0011] Figure 1 A scene diagram of application of the speech recognition method according to the embodiments of the present application is shown.

[0012] Figure 2 A flowchart of the speech recognition method according to the embodiments of the present application is shown.

[0013] Figure 3 A schematic diagram of data processing of the speech recognition method according to the embodiments of the present application is shown.

[0014] Figure 4 A schematic diagram of data processing of the pre-trained speech recognition module according to the embodiments of the present application is shown.

[0015] Figure 5 A schematic diagram of data processing of the noise perception attention mechanism according to the embodiments of the present application is shown.

[0016] Figure 6 A schematic diagram of pre-training data processing according to the embodiments of the present application is shown.

[0017] Figure 7 A structural block diagram of the speech recognition system according to the embodiments of the present application is shown.

[0018] Figure 8 A block diagram of an electronic device suitable for implementing the speech recognition method according to the embodiments of the present application is shown. DETAILED DESCRIPTION

[0019] Embodiments of the present application will be described herein below with reference to the drawings. It should be understood, however, that the description is merely exemplary and is not intended to limit the scope of the present application. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the embodiments of the present application. It will be apparent, however, that one or more embodiments can be practiced without these specific details. In other instances, well-known structures and techniques have not been described in detail in order to avoid unnecessarily obscuring the concepts of the present application.

[0020] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the present application. As used herein, the term "includes" and tautological expressions thereof, such as "including," "includes," "include," "contains," "containing," and so on, mean the term "comprises."

[0021] All terms used herein, including technical and scientific terms, have the meanings commonly understood by one of ordinary skill in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having meanings consistent with the context of the specification, and should not be interpreted in an idealized or overly formal manner.

[0022] In the case of using expressions similar to "at least one of A, B, and C, etc.", it should be generally interpreted that the meaning of the expression is at least one of the items listed before the conjunction, and pluralities thereof (e.g., "a system having at least one of A, B, and C" should include but not be limited to a system having A alone, a system having B alone, a system having C alone, a system having both A and B, a system having both A and C, a system having both B and C, and / or a system having A, B, and C together, etc.).

[0023] In the technical solutions of the present application, the user information (including but not limited to user personal information, user image information, user equipment information, such as location information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved are information and data authorized by the user or authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of related data comply with relevant laws, regulations and standards, necessary security measures are taken, do not violate public order and good customs, and provide corresponding operation portal for user to choose authorization or refusal.

[0024] The speech recognition system has the problems of harmonic confusion and human voice lyric interference at the signal level. When the background music is mixed with the target speech, the harmonic frequency characteristics of the accompaniment part of the background music often overlap with the spectral distribution of the human voice target speech, making it difficult for traditional speech enhancement and separation methods to accurately distinguish between the two. At the same time, the human voice lyric component embedded in the background music has a high degree of coupling with the target speech in the time and frequency dimensions, further increasing the difficulty of speech separation and recognition.

[0025] The voice recognition system has the problem of decoder confusion caused by residual lyrics at the semantic level. Even after noise reduction or signal separation processing at the front end, the residual lyrics in the background music may still be recognized as valid speech signals by the speech model or decoder at the back end, interfering with semantic understanding and text decoding, and thus reducing the recognition accuracy.

[0026] Currently, the problem of the influence of background music noise on the voice recognition model has been widely concerned. By introducing noisy speech data and clean speech data for joint training, the recognition ability of the voice recognition model in real scenarios is improved. However, this method faces the problem of domain difference, that is, the background music noise data and the clean speech frequency data have differences in feature distribution, which makes it difficult for the voice recognition model to learn robust parameters with good generalization ability when processing clean speech and noisy speech scenes at the same time.

[0027] Therefore, there is an urgent need for a technical solution that can effectively separate and recognize speech in a high background music noise environment.

[0028] Embodiments of the present application provide a voice recognition method that can effectively separate and recognize speech in a high background music noise environment, in order to solve at least one of the above technical problems.

[0029] Figure 1 An application scenario diagram of the voice recognition method according to an embodiment of the present application is shown.

[0030] As shown in Figure 1 The application scenario 100 according to this embodiment can include a voice recognition application scenario. The network 104 is a medium for providing communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103 and the server 105. The network 104 can include various connection types, such as wired, wireless communication links or optical fiber cables, etc.

[0031] The user can use the first terminal device 101, the second terminal device 102, the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as an example).

[0032] The first terminal device 101, the second terminal device 102, the third terminal device 103 can be various electronic devices with display screens and supporting web browsing, including but not limited to smartphones, tablet computers, laptop computers and desktop computers, etc.

[0033] The server 105 can be a server providing various services, for example, a background management server providing support for a website browsed by a user using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (only as an example). The background management server can perform analysis and the like on received user requests and the like, and feed back a processing result (for example, a web page, information, or data, or the like, obtained or generated according to a user request) to a terminal device.

[0034] It should be noted that the voice recognition method provided by the embodiments of the present application can generally be executed by the server 105. Correspondingly, the voice recognition system provided by the embodiments of the present application can generally be arranged in the server 105. The voice recognition method provided by the embodiments of the present application can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105. Correspondingly, the voice recognition system provided by the embodiments of the present application can also be arranged in a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105.

[0035] It should be understood that the number of terminal devices, networks, and servers in the system 100 is only illustrative. According to the needs of implementation, there can be any number of terminal devices, networks, and servers. Figure 1

[0036] The voice recognition method according to the embodiments of the present application will be described in detail below based on the scenario described above, by Figure 1 Figures 2-8

[0037] Figure 2 A flowchart of the voice recognition method according to the embodiments of the present application is shown.

[0038] As shown in Figure 2 , the voice recognition method includes operations S210-S240.

[0039] In operation S210, in response to a voice recognition instruction, original audio data is acquired, the original audio data including a voice signal with background music noise.

[0040] The original audio data can be a real-time collected voice stream or a pre-recorded audio file. The voice signal is a mixed audio with background music noise. The original audio data can be preprocessed, for example, sampling rate conversion, volume normalization, and the like.

[0041] In operation S220, the voice signal is subjected to time-frequency conversion to obtain initial spectral features. ​​​

[0042] A short-time Fourier transform can be employed for time-frequency transform. The speech signal is processed in frames, and the length of each frame can be in the range of 20-30 milliseconds, and there can be some overlap between adjacent frames. A window function (e.g., a Hamming window or a Hanning window) is applied to each frame of signal, and then a fast Fourier transform is performed to obtain a complex spectrum. The magnitude spectrum of the complex spectrum is extracted, and the magnitude spectrum is taken as an initial spectral feature.

[0043] In operation S230, the initial spectral feature is input into a pre-trained speech separation module for background music suppression to output a target spectral feature.

[0044] In operation S240, the target spectral feature and the initial spectral feature are input into a pre-trained speech recognition module to output a recognized text. The recognized text can be displayed through a display interface, or passed to a subsequent application for processing.

[0045] Figure 3 A schematic diagram of data processing of a speech recognition method according to an embodiment of the present application is shown.

[0046] As shown in Figure 3 , the original audio data includes a speech signal with background music noise, and the speech signal is subjected to time-frequency transform to obtain an initial spectral feature.

[0047] As shown in Figure 3 , the speech separation module includes an explicit music harmonic convolutional encoder, an implicit music perception encoder, and a gated attention fusion processor. The explicit music harmonic convolutional encoder is used to extract the scale and harmonic structure features of the background music noise in the initial spectral feature to determine frame-level music features. The implicit music perception encoder is used to perform speech enhancement and noise suppression on the initial spectral feature to output frame-level speech features of the same dimension as the frame-level music features. The gated attention fusion processor is used to perform gated fusion on the frame-level music features and the frame-level speech features to output gated fusion features; the gated fusion features are used to generate a speech separation mask, and the speech separation mask is used to separate the initial spectral feature to obtain a target spectral feature.

[0048] The scale and harmonic structure features of the background music noise can be summarized as three high-level categories: pitch (scale), harmonic kernel, and band gating function. The "key" of the background music noise is determined by the most prominent fundamental frequency, corresponding to the scale; the "thickness" and "brightness" are determined by the concentration intensity envelope of the integer harmonic of the fundamental frequency, i.e., the harmonic kernel; and which harmonics can be released in real time and which are suppressed are determined by a 0-1 switch curve that changes over time, i.e., the band gating function, which can successively limit the access rights of the pitch, timbre, and audible segments. The gated fusion feature refers to a multi-scale high-dimensional feature representation obtained by weighted fusion of the explicit music feature and the implicit speech feature through a gated attention mechanism.

[0049] As shown in Figure 3 The voice recognition module includes a noise-aware attention mechanism and a convolutional self-attention hybrid encoder. The noise-aware attention mechanism fuses the target spectral features and the initial spectral features through learnable gating weights to adaptively suppress residual music noise. The convolutional self-attention hybrid encoder is used for sequence modeling of the fused features to obtain the recognized text. Sequence modeling can convert the fused high-dimensional features into more compact and effective low-dimensional sequence representations, generate more recognizable feature representations, maximize the information value of the fused features, and help the model output more accurate results.

[0050] According to embodiments of the present application, the explicit music harmonic convolutional encoder extracts the pitch and harmonic structure features of the background music noise, and the implicit music-aware encoder performs speech enhancement and noise suppression on the initial spectral features, which can separate the background music noise of the speech signal and extract high-purity speech features, thereby improving the recognition accuracy of the recognized text. By online correcting the frequency response of the filter bank, the speech dominant frequency band can be made more regular, and the music dominant frequency band can be made more controllable, further improving the recognition accuracy of the recognized text. By adaptively suppressing residual music noise and sequence modeling of the fused features, residual noise can be accurately suppressed without affecting the spectral features of the target speech, further improving the recognition accuracy of the recognized text, while reducing the error rate and enhancing the scene adaptability, which can adapt to different music interference scenes to prevent "over suppression" or "insufficient suppression".

[0051] The speech dominant frequency band represents a specific frequency range in the spectrum of the speech signal with the most concentrated amplitude and the most significant contribution to the overall speech signal. The music dominant frequency band represents a specific frequency range in the spectrum of the background music with the most concentrated amplitude and the most significant contribution to the overall background music. The music dominant frequency band can include the dominant frequency band of the vocal lyrics in the background music and the dominant frequency band of the instrumental accompaniment in the background music. In the voice recognition scene, the dominant frequency band of the vocal lyrics in the speech signal often overlaps with the dominant frequency band of the vocal lyrics in the background music, and the dominant frequency band of the instrumental accompaniment in the background music can also intersect with the dominant frequency band of the speech signal.

[0052] The dominant frequency band of the instrumental accompaniment in the background music is a "harmonic set frequency band" based on the instrument fundamental frequency and the overtone as the main component. The instrument fundamental frequency is the lowest frequency generated when the instrument vibrates, and is also the "base frequency" of the sound. The overtone frequency band is the "harmonic component" superimposed on the instrument fundamental frequency, which is a high-frequency vibration generated synchronously when the instrument fundamental frequency vibrates. The instrument fundamental frequency and the overtone frequency band form a continuous harmonic spectrum, and when the instrument fundamental frequency and the overtone frequency band of the background music overlap with the spectrum range of the speech signal, "harmonic confusion" occurs. The spectrum of the speech signal also contains the instrument fundamental frequency and the overtone frequency band.

[0053] The structure of the "instrument fundamental overtone frequency band" is explicitly introduced into the speech separation module by using an explicit music harmonic convolution encoder, a human ear-like filter bank is constructed by using a deep separable convolution, the speech dominant frequency band is enhanced and the music dominant frequency band is suppressed, the instrument fundamental frequency and the overtone frequency band can be accurately extracted, the problem of harmonic overlap is solved, and the extraction accuracy can be improved in the same data set of clean speech and background music with high instrument accompaniment ratio.

[0054] The convolutional self-attention hybrid encoder can be a fusion of CNN (Convolutional Neural Network) and Transformer (Transformer model), which can capture local spectral features of target spectral features, model long-time dependence, capture the relationship between acoustic features and semantic information of target spectral features at different time points, and distinguish the long-time association of lyrics fragments in speech signals and background music. Speech signals are continuous time series, and the context of speech signals (such as previous and subsequent sentences) and the context of lyrics in background music have different association patterns, and modeling long-time dependence can clearly define the boundaries of the two. Even if there are lyrics left after the front-end processing, the convolutional self-attention hybrid encoder can determine the semantic association of the lyrics left fragments and the speech signals through long-time dependence, and exclude invalid interference.

[0055] According to the embodiments of the present application, the initial spectral features are input into the explicit music harmonic convolution encoder to obtain frame-level music features; the initial spectral features are input into the implicit music perception encoder to obtain frame-level speech features; the frame-level music features and the frame-level speech features are input into the gated attention fusion processor to generate gated fusion features; the gated fusion features are input into the mask generator to predict the speech separation mask; and the speech separation mask is used to weight the initial spectral features in the frequency domain to separate the target spectral features.

[0056] Frame-level music features The frame-level music features represent the pitch energy and harmonic structure of the background music effectively extracted by the explicit music harmonic convolution encoder, and can provide context information of music distribution for the speech separation task. Frame-level speech features The frame-level speech features represent the clean speech features after suppressing background music noise extracted by the implicit music perception encoder. The frame-level music features and the frame-level speech features are high-dimensional representations at the frame level, which can be understood as that the output result is a high-dimensional feature representation, highlighting the features of the speech signal at each time frame, and providing clearer speech information for subsequent processing steps. By converting each frame of frequency domain spectrum of the initial spectral features into a high-dimensional vector, single-frame detail information can be accurately captured, and high-quality basis can be provided for subsequent time series modeling, cross-modal fusion, etc.

[0057] The frame-level music features obtained by the explicit path are spliced along the channel dimension Frame-level speech features obtained by implicit path , to obtain spliced features :

[0058] (1).

[0059] Generate dynamic gating weights by lightweight convolution :

[0060] (2);

[0061] wherein, is an activation function . represents a one-dimensional convolution operation.

[0062] Frame-level music features and frame-level speech features are spliced after weighting to obtain gating fusion features . The gating fusion features (512 channels, T frames) can be represented as:

[0063] (3).

[0064] Decode the gating fusion features , and use transposed convolution to upsample the gating fusion features to the original spectral dimension 513xTx2, to predict the speech separation mask . The speech separation mask is a complex mask that can match the noisy spectral dimension. The speech separation mask can be represented as:

[0065] (4)

[0066] wherein, represents transposed convolution, and the transposed convolution is a carrier of trainable parameters, the weight of which is continuously optimized in the back propagation, rather than a fixed algorithm. When the loss function is back propagated, the gradient directly corrects the mask prediction parameters (transposed convolution kernel, gating weight, etc.), which is beneficial. Effect, can realize the dynamic prediction optimization of the mask under multi-task supervision.

[0067] The speech separation mask is used to weight the initial spectral features in the frequency domain, which can separate out clean speech to obtain target spectral features .

[0068] (5).

[0069] According to an embodiment of the present application, based on the twelve equal temperament, the initial spectral features are frequency band divided by scale and mapped to a standard music scale energy matrix; for each scale of the standard music scale energy matrix, a harmonic kernel corresponding to each scale is constructed, and the effective frequency range of the harmonic kernel is constrained based on the frequency band gating function of the twelve equal temperament, the harmonic kernel including learnable harmonic amplitude and decay envelope parameters; based on the constrained harmonic kernel, a convolution operation is performed on the target scale corresponding to the harmonic kernel to obtain the time sequence energy value of the target scale, and after traversing all scales, a target scale time sequence energy matrix is generated; and the target scale time sequence energy matrix is normalized and channel expanded to output frame-level music features.

[0070] According to the twelve equal temperament, an octave can be divided into twelve logarithmically uniform semitones. A standard scale encoder is used to align the initial spectral features to the 88 scales of a standard piano to obtain a standard music scale energy matrix. For each scale of the standard music scale energy matrix, a harmonic kernel corresponding to each scale is constructed, and the harmonic kernel of each scale is a "frequency-amplitude" two-dimensional vector group.

[0071] Time-domain harmonic kernel which can be expressed as:

[0072] (6).

[0073] (7).

[0074] wherein, is the center frequency, represents the harmonic order, represents the decay envelope, the harmonic order amplitude and the decay envelope coefficient are learnable parameters. represents the scale index from 1 to 88 scales, represents the harmonic order, corresponds to the typical value of the piano timbre. is a discrete time point from 0 to , and is the length of the time-domain harmonic kernel. is a natural constant.

[0075] The time-domain harmonic kernel is subjected to Fourier transform to obtain the frequency-domain harmonic kernel . A frequency band gating function based on the twelve equal temperament is used for frequency band constraint.

[0076] The frequency band gating function can be expressed as:

[0077] (8);

[0078] (9).

[0079] Based on the constrained harmonic kernel, a convolution operation is performed on the target pitch corresponding to the harmonic kernel to obtain the temporal energy value of the target pitch. After traversing all musical scales, the temporal energy matrix of the target musical scale is generated.

[0080] Temporal energy value of the target scale It can be represented as:

[0081] (10);

[0082] in, This represents the convolution operation. The initial spectral features in time frames and frequency The plural representation of the place, It is the first The set of frequency bands corresponding to each piano key.

[0083] Temporal energy value of the target scale Normalization is performed to obtain the processed temporal energy value. Processed time-series energy value It can be represented as:

[0084] (11);

[0085] in, This represents the temporal energy value of the maximum value in the target scale temporal energy matrix. The processed target scale temporal energy matrix is ​​essentially a frame-level musical feature. .

[0086] According to embodiments of this application, the harmonic response is concentrated only in the actual musical scale frequency band by constraining the band gate function. The attenuation envelope can be used to model the natural attenuation of different harmonics in the time dimension, more closely reflecting the acoustic characteristics of real musical instruments. By using an explicit musical harmonic convolutional encoder to learn musical features using prior knowledge, separation accuracy can be improved in background music noise with significant overtone and non-harmonic interference.

[0087] An implicit music-aware encoder can be used to learn to suppress background music noise.

[0088] According to the embodiment of the present application, the initial spectral feature is divided into a plurality of subbands, a depth separable convolution is independently performed on each subband, and a subband feature of each subband is extracted; the spliced subband features are input into a bandwidth controller to extract a timing feature of each subband, a channel attention mechanism is used to calculate a subband scaling factor based on the timing feature of each subband, the spliced subband features are weighted and enhanced and inhibited according to the subband scaling factor, and an enhanced subband feature corresponding to each subband is output through a residual connection; the enhanced subband features are input into a frequency domain attention gate, attention weights of each subband are independently generated and weighted after the channel dimension is expanded, and each subband feature is aggregated along the frequency dimension after the residual connection, and a frame-level speech feature is output.

[0089] The speech signal is divided into a plurality of time frames, and each frame is processed independently. On each time frame, a one-dimensional convolution operation is used to encode the spectral feature. The one-dimensional convolution is performed along the frequency axis (frequency dimension), and can extract local frequency patterns. A plurality of one-dimensional convolution layers can be used to extract timing features from the initial spectral features frame by frame. Each one-dimensional convolution layer is followed by batch normalization and a ReLU (Rectified Linear Unit) activation function, which converts each time frame into a high-dimensional vector, and collectively forms a frame-level feature to obtain a compressed audio feature representation.

[0090] A set of filter banks is constructed to independently perform depth separable convolution on each subband. The filter banks can be calculated according to the twelve equal temperament, and the center frequency of each subband . For example, the boundaries of each subband are defined as the geometric mean of the center frequencies of adjacent scales, and 513 frequency points can be divided into 24 frequency subbands according to the boundaries, . A one-dimensional convolution is independently performed on each subband, the input channel is 2, the output channel is 128, and feature splicing is performed after convolution.

[0091] The filter banks are constructed using depth separable convolution, which can simulate the perceptual characteristics of the human ear for different frequency sounds. Depth separable convolution decomposes standard convolution into depth convolution and point-by-point convolution, which can reduce the number of parameters and the amount of calculation compared with standard convolution.

[0092] The timing features of each subband are extracted along the time dimension using depth separable convolution . Depth separable convolution is independently used for each subband , which can further reduce the number of parameters and enhance the real-time performance in the subband dimension. The timing features of each subband can be represented as:

[0093] (12) ;

[0094] ​wherein, represents a depthwise separable convolution operation. Each subband is an input spectrum of 24xTx128.

[0095] The time sequence features are globally averaged pooled, and a scaling factor of each subband (a 24-dimensional vector) is generated by a fully connected layer. The scaling factor of each subband can be expressed as:

[0096] (13).

[0097] wherein, represents parameters of the first fully connected layer; represents parameters of the second fully connected layer. represents Global Average Pooling; represents an activation function; represents an activation function.

[0098] The scaling factor of each subband is applied to each subband and a residual connection is performed to output an enhanced subband feature . The enhanced subband feature can be expressed as:

[0099] (14).

[0100] Each subband is channel-wise weighted by the scaling factor of each subband, which can enhance a speech-dominant subband and suppress a noise-dominant subband, and through a residual connection, original information can be preserved.

[0101] The enhanced subband feature is input into a frequency domain attention gate. After expanding the channel dimension, attention weights are independently generated for each subband and are used for weighted correction, and after a residual connection, the features of each subband are aggregated along the frequency dimension to output a frame-level speech feature. The weights of each subband can describe the importance of each subband and reflect the relative strength of the speech signal in each subband. The weight value can also reflect the importance of the subband to the speech recognition task. All weighted subband features are concatenated along the frequency axis to form a complete spectral feature. Using attention weights to weight each frequency subband can enhance the features of a speech-dominant frequency band and suppress the features of a music-dominant frequency band. A one-dimensional convolution is used to integrate the channels of the concatenated subband features, converting the features into a high-dimensional vector.

[0102] For example, a global average pooling and linear layer can be used to construct a frequency domain attention mechanism, which can also be referred to as a bandwidth controller. The subband features of each subband are globally averaged pooled to obtain statistical information of each subband, and then the statistical information of each subband is input into two linear layers (with ReLU activation in between) and a Sigmoid activation function to generate weight coefficients of each subband. The weight coefficients of each subband are applied to the subband features of each subband. The channel number can be expanded by one-dimensional convolution, for example, from 128 to 256.

[0103] The input channel number is expanded from 128 to 256, and the spatial and temporal dimensions (24xT) are retained. The dimension of each subband after channel expansion is 24xTx256, and each subband after channel expansion can be represented as:

[0104] (15) ;

[0105] wherein each subband after channel expansion represents that the channel number of each subband is expanded from 128 to 256.

[0106] The attention weight of each subband is independently modeled as:

[0107] (16) ;

[0108] wherein ; ; ; wherein represents an indexing operation on each subband after channel expansion , the “:” in is a standard symbol for indexing, which means taking all elements in the corresponding dimension.

[0109] The attention weight is applied to each subband, and a residual connection is used to retain the original features after expansion to avoid information loss. The attention weight is independently generated for each subband after expansion of the channel dimension, and the final output has a dimension of 24xTx256. After the residual connection, the features of each subband are aggregated along the frequency dimension, and the frame-level speech feature is output.

[0110] The frequency domain attention gate works with the filter bank and bandwidth controller of the implicit music perception encoder to provide high-discriminative spatio-temporal frequency features for subsequent feature fusion. Then, one-dimensional convolution is used to fuse the features of the 24 subbands to output the gated fusion features, which represent the clean speech features extracted by the implicit music perception encoder after suppressing background music noise. ​​

[0111] According to the embodiments of the present application, the frequency response characteristics of the filter bank are dynamically adjusted by using the attention weight mechanism, which can enhance the feature response of the speech dominant subband and suppress the noise dominant subband. The spliced modified subband features and the frame-level high-dimensional representation are fused through the residual connection to obtain the final separated speech features. By suppressing the interference of background music and retaining the main information of the target speech, high-purity speech spectrum features can be output, further improving the feature extraction capability.

[0112] According to the embodiments of the present application, the implicit music perception encoder adds a "linear bottleneck-attention weight" bypass inside the network of the speech separation module, which can online correct the frequency response of the filter bank without additional supervision. Through the joint action, the implicit music perception encoder can dynamically adjust and optimize the feature representation, realize "music scene adaptation", and overcome the problem that the traditional static filter cannot cope with various accompaniments. Further improve the accuracy and robustness of speech recognition.

[0113] Figure 4 A schematic diagram of data processing of a pre-trained speech recognition module according to an embodiment of the present application is shown.

[0114] As shown in Figure 4 , the target spectrum feature and the initial spectrum feature are input into the noise perception attention mechanism to generate a fusion feature.

[0115] Figure 5 A schematic diagram of data processing of a noise perception attention mechanism according to an embodiment of the present application is shown.

[0116] As shown in Figure 5 , the target spectrum feature and the initial spectrum feature are spliced to generate a fusion input spectrum feature. One-dimensional convolution compression is performed on the fusion input spectrum feature to extract the encoded feature representation of the noisy speech, so as to obtain a compressed feature. Based on the compressed feature, a gating weight is generated through a convolution layer and an activation function, and the gating weight is used to represent the relative proportion of speech and noise in the original audio. The activation function can be a Sigmoid activation function. The initial spectrum feature is dynamically weighted by using the gating weight to obtain a weighted spectrum feature in which the noise is suppressed. The weighted spectrum feature is a more semantically clean feature representation. An element-level dynamic weighting method can be used to guide the filtering of the initial spectrum feature, which can suppress the spectrum response of the music dominant frequency band. The weighted spectrum feature and the target spectrum feature are combined through an attention fusion mechanism to output a fusion feature, which is a high-dimensional semantic feature.

[0117] According to an embodiment of the present application, the noise-aware attention mechanism introduces learnable gating weights at the recognition end to estimate the "noise proportion" of the target spectral features and the initial spectral features, i.e., the ratio of speech and noise in the original audio, to adaptively identify and retain speech features while suppressing background music features, and to determine in real time whether to "trust the separation result" or "retain the original clues", thereby reducing decoding confusion caused by residual lyrics.

[0118] According to an embodiment of the present application, by fusing the target spectral features and the initial spectral features, and dynamically adjusting the contribution of the target spectral features and the initial spectral features according to the relative proportion of speech and noise in the original audio, the word error rate and the word error rate can be reduced in an English background music mixed dataset with more noisy background music, and the robustness of the speech recognition system can be further improved.

[0119] As shown in Figure 4 The fusion features are input into a convolutional self-attention hybrid encoder, sequentially pass through a feedforward network, multi-head self-attention, deep separable convolution, and residual normalization, and output high-order hidden representations H.

[0120] The convolutional self-attention hybrid encoder adopts a multi-layer hybrid coding structure, which can effectively extract local and global features. Each layer of the hybrid coding structure includes a feedforward network, multi-head self-attention, deep separable convolution, and residual normalization. The feedforward network can use a feedforward neural network to perform nonlinear feature transformation on the fusion features, enhancing the feature expression capability. The multi-head self-attention can capture long-distance relationships in the speech segment, generate context-related feature representations by capturing relationships between different position features. The deep separable convolution can extract local temporal features, enhance the modeling capability of the local structure of the speech, and reduce the computational complexity. The residual normalization can use residual connection and layer normalization, and each sub-module uses residual connection to prevent gradient disappearance, accelerate network training, and enhance the perceptual ability and training stability of the convolutional self-attention hybrid encoder.

[0121] For example, the fusion features are mapped to the hidden dimension of the convolutional self-attention hybrid encoder, for example, 256 dimensions, through a linear layer, and processed through a multi-layer, for example, 12 layers, hybrid coding structure. The feedforward network can use two layers of linear transformation and GELU (Gaussian Error Linear Unit) activation function for nonlinear feature transformation. The multi-head self-attention can use a multi-head self-attention mechanism to capture long-distance dependencies in the speech sequence.

[0122] As shown in Figure 4As shown, the high-order hidden representation H is input into a connectionist temporal classification decoder, and after linear classification and beam search, the recognized text is obtained. The high-order hidden representation H contains rich speech information and can be directly used for speech recognition. The connectionist temporal classification decoder can directly classify the high-order hidden representation H through linear classification to obtain the recognized text. The beam search can obtain more accurate recognition results while reducing the amount of calculation.

[0123] The connectionist temporal classification decoder is a CTC (Connectionist Temporal Classification) decoder or an attention-based decoder. The CTC decoder maps the hidden features to the output of the size of the vocabulary table through a linear layer, which can be trained in combination with the CTC loss function. CTC can handle cases where the lengths of the input sequence and the output sequence are inconsistent. The attention-based decoder uses an autoregressive method to generate the output text sequence step by step. The decoder calculates the attention weight according to the encoder output and the previously generated text at each time step, and predicts the next character or word, and finally outputs the recognized text sequence, i.e., the content of the target speech.

[0124] In order to realize the collaborative optimization of the speech separation module and the speech recognition module, the application provides a training method for end-to-end training using a multi-task learning framework.

[0125] Figure 6 A schematic diagram of pre-training of data according to an embodiment of the application is shown.

[0126] As shown Figure 6 The pre-training of the speech separation module and the speech recognition module includes the following operations.

[0127] A noisy speech sample, a clean speech sample, and a text label corresponding to the clean speech sample are obtained. The noisy speech sample and the clean speech sample are respectively subjected to time-frequency transformation to generate noisy spectrum features and clean spectrum features. The noisy spectrum features are input into an initial speech separation module to generate predicted spectrum features after noise reduction. The predicted spectrum features and the noisy spectrum features are jointly input into an initial speech recognition module to obtain predicted text. A joint loss of the speech separation module and the speech recognition module is determined based on the clean spectrum features, the predicted spectrum features, the predicted text, and the text label. The initial speech separation module and the initial speech recognition module are jointly trained based on the joint loss to update the module parameters until convergence.

[0128] By balancing the training of the speech separation module and the speech recognition module through a dynamic loss weighting strategy, the speech separation loss and the speech recognition loss can be calculated by constructing a joint loss function.

[0129] The speech separation loss L_separation is calculated according to the predicted spectral feature and the clean spectral feature. In one example, a spectral distance loss such as Mean Squared Error (MSE) or Scale-Invariant Signal-to-Distortion Ratio (SI-SDR) can be employed as the speech separation loss. The speech separation loss L_separation can measure the difference between the separated speech and the clean speech reference.

[0130] The speech recognition loss L_recognition is calculated according to the predicted text and the text label. In one example, a CTC loss or a cross-entropy loss can be used. The speech recognition loss L_recognition can measure the difference between the predicted text and the true labeled text.

[0131] The joint loss L_total is determined according to the speech separation loss L_separation and the speech recognition loss L_recognition. The joint loss L_total can be expressed as:

[0132] L_total = λ1 × L_separation + λ2 × L_recognition (17);

[0133] where λ1 is the weight coefficient of the speech separation loss L_separation, and λ2 is the weight coefficient of the speech recognition loss L_recognition. Through the dynamic loss weighting strategy balance, the training of the speech separation task and the speech recognition task can be balanced.

[0134] In the early stage of training, the weight of the speech separation task is larger, and the speech separation loss is given a higher weight, so that the model can prioritize establishing separation capability and learn effective separation capability. As the training proceeds, the weight of the speech recognition task is gradually increased, and the weight of the speech recognition loss is gradually increased, so that the model can optimize both tasks at the same time. The dynamic loss weighting strategy can realize the dynamic transition from separation dominance to recognition dominance, and until the two tasks reach balanced convergence. The dynamic weight adjustment can be based on the performance of the validation set or the change trend of the training loss.

[0135] For example, at the beginning of training, the speech separation loss weight is set to be higher, for example, 0.8, and the speech recognition loss weight is set to be 0.2, so that the model preferentially learns the separation capability; after each training round is completed, the word error rate is calculated on the validation set, if the word error rate does not decrease for two consecutive training rounds, the speech recognition loss weight is increased by 0.05, and the speech separation loss weight is correspondingly reduced by 0.05; the lower limit of the weight adjustment is that the speech separation loss weight is 0.4 and the speech recognition loss weight is 0.6, and the lower limit is kept constant until the joint loss converges.

[0136] The network model parameters of the speech separation task and the speech recognition task are optimized in an end-to-end manner. The gradient of the joint loss L_total with respect to each parameter is calculated. The optimizer such as Adam (Adaptive Moment Estimation) is used to update the network model parameters. The speech separation module and the speech recognition module share part of the underlying feature representation, which can enhance the cross-task feature consistency.

[0137] The training method of the embodiment of the application adopts a multi-task learning framework for end-to-end training, and adopts a dynamic loss weighting strategy to adaptively adjust the speech separation task and the speech recognition task. By jointly optimizing the speech separation task and the speech recognition task, the speech separation and recognition performance can be simultaneously improved in a complex background music noise scene, and the robustness and adaptability of the speech recognition system are improved, so as to be applied to various practical scenes, such as intelligent customer service, voice assistant, conference recording, etc.

[0138] Through the explicit music harmonic convolutional encoder of the speech separation module and the convolutional self-attention hybrid encoder of the speech recognition module, the underlying features are shared, which can enhance the cross-task feature consistency and realize end-to-end multi-task collaborative optimization.

[0139] Figure 7 A structural block diagram of a speech recognition system according to an embodiment of the application is shown.

[0140] As shown in Figure 7 The speech recognition system 700 includes an acquisition module 710, a time-frequency transformation module 720, a target spectral feature determination module 730, and a text recognition module 740.

[0141] The acquisition module 710 is configured to acquire original audio data in response to a speech recognition instruction, the original audio data including a speech signal with background music noise.

[0142] The time-frequency transformation module 720 is configured to perform time-frequency transformation on the speech signal to obtain initial spectral features.

[0143] The target spectrum feature determination module 730 is configured to input the initial spectrum feature into a pre-trained speech separation module for background music suppression, to output a target spectrum feature.

[0144] The text recognition module 740 is configured to input the target spectrum feature and the initial spectrum feature into a pre-trained speech recognition module, to output recognized text.

[0145] The speech separation module includes an explicit music harmonic convolutional encoder, an implicit music perception encoder, and a gated attention fusion processor. The explicit music harmonic convolutional encoder is configured to extract scale and harmonic structure features of background music noise in the initial spectrum feature, to determine frame-level music features. The implicit music perception encoder is configured to perform speech enhancement and noise suppression on the initial spectrum feature, to output frame-level speech features with the same dimension as the frame-level music features. The gated attention fusion processor is configured to perform gated fusion on the frame-level music features and the frame-level speech features, to output gated fusion features. The gated fusion features are used to generate a speech separation mask. The speech separation mask is used to separate the initial spectrum feature to obtain the target spectrum feature.

[0146] The speech recognition module includes a noise perception attention mechanism and a convolutional self-attention hybrid encoder. The noise perception attention mechanism fuses the target spectrum feature and the initial spectrum feature through learnable gating weights, to adaptively suppress residual music noise. The convolutional self-attention hybrid encoder is configured to perform sequence modeling on the fused features, to obtain recognized text.

[0147] According to an embodiment of the present application, the target spectrum feature determination module 730 includes a first frame-level feature extraction submodule, a second frame-level feature extraction submodule, a gated fusion feature extraction submodule, a mask prediction submodule, and a frequency domain weighting submodule.

[0148] The first frame-level feature extraction submodule is configured to input the initial spectrum feature into the explicit music harmonic convolutional encoder, to obtain frame-level music features.

[0149] The second frame-level feature extraction submodule is configured to input the initial spectrum feature into the implicit music perception encoder, to obtain frame-level speech features.

[0150] The gated fusion feature extraction submodule is configured to input the frame-level music features and the frame-level speech features into the gated attention fusion processor, to generate gated fusion features.

[0151] The mask prediction submodule is configured to input the gated fusion features into a mask generator, to predict a speech separation mask.

[0152] The frequency domain weighting submodule is configured to perform frequency domain weighting on the initial spectrum feature by using the speech separation mask, to separate out the target spectrum feature.

[0153] According to an embodiment of the present application, the first frame-level feature extraction submodule includes a frequency band division unit, a harmonic kernel construction unit, an energy matrix generation unit, and a normalization processing unit.

[0154] The frequency band division unit is configured to perform frequency band division on the initial frequency spectrum features according to the twelve equal temperament, map to a standard music scale energy matrix.

[0155] The harmonic kernel construction unit is configured to construct a harmonic kernel corresponding to each scale of the standard music scale energy matrix, and constrain the effective frequency range of the harmonic kernel based on a frequency band gating function of the twelve equal temperament.

[0156] The energy matrix generation unit is configured to perform convolution operation on the target scale corresponding to the harmonic kernel based on the constrained harmonic kernel, to obtain a time sequence energy value of the target scale, and generate a target scale time sequence energy matrix after traversing all scales.

[0157] The normalization processing unit is configured to perform normalization and channel expansion processing on the target scale time sequence energy matrix, and output frame-level music features.

[0158] According to an embodiment of the present application, the second frame-level feature extraction submodule includes a sub-band division unit, an enhanced sub-band feature unit, and a frequency domain attention gate unit.

[0159] The sub-band division unit is configured to divide the initial frequency spectrum features into a plurality of sub-bands, independently perform depth separable convolution on each sub-band, and extract sub-band features of each sub-band.

[0160] The enhanced sub-band feature unit is configured to input the spliced sub-band features into a bandwidth controller to extract time sequence features of each sub-band, calculate a sub-band scaling factor based on the time sequence features of each sub-band through a channel attention mechanism, weight and enhance or suppress the spliced sub-band features according to the sub-band scaling factor, and output enhanced sub-band features corresponding to each sub-band through residual connection.

[0161] The frequency domain attention gate unit is configured to input the enhanced sub-band features into a frequency domain attention gate, independently generate attention weights for each sub-band after expanding the channel dimension, and weight and correct the attention weights, aggregate the sub-band features along the frequency dimension after residual connection, and output frame-level speech features.

[0162] According to an embodiment of the present application, the speech recognition system further includes a pre-training module.

[0163] The pre-training module is configured to pre-train the speech separation module and the speech recognition module.

[0164] The pre-training module includes an acquisition submodule, a transformation submodule, a feature generation submodule, an identification submodule, a determination submodule, and a joint training submodule.

[0165] The acquisition sub-module is configured to acquire noisy speech samples, clean speech samples, and text labels corresponding to the clean speech samples.

[0166] The transformation sub-module is configured to perform time-frequency transformation on the noisy speech samples and the clean speech samples respectively to generate noisy spectrum features and clean spectrum features.

[0167] The generation sub-module is configured to input the noisy spectrum features into an initial speech separation module to generate predicted spectrum features after noise reduction.

[0168] The recognition sub-module is configured to input the predicted spectrum features and the noisy spectrum features into an initial speech recognition module together to obtain predicted text.

[0169] The determination sub-module is configured to determine a joint loss of the speech separation module and the speech recognition module based on the clean spectrum features, the predicted spectrum features, the predicted text, and the text labels.

[0170] The joint training sub-module is configured to perform joint training on the initial speech separation module and the initial speech recognition module based on the joint loss to update module parameters until convergence.

[0171] According to an embodiment of the present application, the determination sub-module includes a first calculation unit, a second calculation unit, and a determination loss unit.

[0172] The first calculation unit is configured to calculate a speech separation loss according to the predicted spectrum features and the clean spectrum features.

[0173] The second calculation unit is configured to calculate a speech recognition loss according to the predicted text and the text labels.

[0174] The determination loss unit is configured to determine a joint loss according to the speech separation loss and the speech recognition loss.

[0175] According to an embodiment of the present application, the text recognition module 740 includes a fusion feature generation sub-module, a hybrid encoding sub-module, and a decoding sub-module.

[0176] The fusion feature generation sub-module is configured to input the target spectrum features and the initial spectrum features into a noise perception attention mechanism to generate fusion features.

[0177] The hybrid encoding sub-module is configured to input the fusion features into a convolutional self-attention hybrid encoder, sequentially through a feedforward network, a multi-head self-attention, a depth separable convolution, and a residual normalization, to output high-order hidden representations.

[0178] The decoding sub-module is configured to input the high-order hidden representations into a connected temporal classification decoder, through linear classification and beam search, to obtain recognized text.

[0179] According to an embodiment of the present application, the fusion feature generation sub-module comprises a feature generation unit, a compression unit, a gating weight generation unit, a dynamic weighting unit and a fusion unit.

[0180] The feature generation unit is configured to splice the target spectrum feature and the initial spectrum feature to generate a fusion input spectrum feature.

[0181] The compression unit is configured to perform one-dimensional convolution compression on the fusion input spectrum feature to obtain a compression feature.

[0182] The gating weight generation unit is configured to generate gating weights based on the compression feature through a convolution layer and an activation function, the gating weights being used to represent the relative proportion of speech and noise in the original audio.

[0183] The dynamic weighting unit is configured to dynamically weight the initial spectrum feature using the gating weights to obtain a weighted spectrum feature in which noise is suppressed.

[0184] The fusion unit is configured to combine the weighted spectrum feature and the target spectrum feature through an attention fusion mechanism to output a fusion feature.

[0185] According to an embodiment of the present application, any multiple modules of the acquisition module 710, the time-frequency conversion module 720, the target spectrum feature determination module 730 and the text recognition module 740 can be combined in one module for implementation, or any one of the modules can be split into multiple modules. Alternatively, at least part of the functions of one or more of the modules can be combined with at least part of the functions of other modules, and implemented in one module. According to an embodiment of the present application, at least one of the acquisition module 710, the time-frequency conversion module 720, the target spectrum feature determination module 730 and the text recognition module 740 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on board, a system on package, an application specific integrated circuit (ASIC), or any other reasonable manner of hardware or firmware that can be integrated or packaged, or implemented in any one of software, hardware and firmware or in a proper combination of any of them. Alternatively, at least one of the acquisition module 710, the time-frequency conversion module 720, the target spectrum feature determination module 730 and the text recognition module 740 can be at least partially implemented as a computer program module that can perform corresponding functions when the computer program module is run.

[0186] Figure 8 A block diagram of an electronic device suitable for implementing the speech recognition method according to an embodiment of the present application is shown.

[0187] As Figure 8As shown, the electronic device 800 according to embodiments of the present application includes a processor 801 which can perform various appropriate actions and processes in accordance with a program stored in a read only memory (ROM) 802 or a program loaded into a random access memory (RAM) 803 from a storage section 808. The processor 801 can include, for example, a general purpose microprocessor (e.g., a CPU), an instruction set processor, and / or a related chip set, and / or a dedicated microprocessor (e.g., an application specific integrated circuit (ASIC)), and so on. The processor 801 can also include an on-board memory for cache use. The processor 801 can include a single processing unit or multiple processing units for executing different actions of the method processes according to embodiments of the present application.

[0188] In the RAM 803, various programs and data required for the operation of the electronic device 800 are stored. The processor 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. The processor 801 performs various operations of the method processes according to embodiments of the present application by executing the programs in the ROM 802 and / or the RAM 803. Note that the programs can also be stored in one or more memories other than the ROM 802 and the RAM 803. The processor 801 can also perform various operations of the method processes according to embodiments of the present application by executing the programs stored in the one or more memories.

[0189] According to embodiments of the present application, the electronic device 800 can also include an input / output (I / O) interface 805 which is also connected to the bus 804. The electronic device 800 can also include one or more of the following components connected to the input / output (I / O) interface 805: an input section 806 including a keyboard, a mouse, etc.; an output section 807 including a display such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 808 including a hard disk, etc.; and a communication section 809 including a network interface card such as a LAN card, a modem, etc. The communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the input / output (I / O) interface 805 as necessary. A removable recording medium 811 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is attached to the drive 810 as necessary, so that a computer program read out therefrom is installed in the storage section 808 as necessary.

[0190] The application further provides a computer readable storage medium, which can be included in the device / apparatus / system described in the above embodiments, or can exist independently without being assembled into the device / apparatus / system. The computer readable storage medium carries one or more programs, which, when executed, implement the method according to the embodiments of the application.

[0191] According to the embodiments of the application, the computer readable storage medium can be a non-volatile computer readable storage medium, which can include, but is not limited to, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any appropriate combination thereof. In this application, a computer readable storage medium can be any tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. For example, in the embodiments of the application, a computer readable storage medium can include the ROM 802 and / or the RAM 803 described above, and / or one or more memory devices other than the ROM 802 and the RAM 803.

[0192] The embodiments of the application also include a computer program product, which includes a computer program containing program codes for executing the methods shown in the flowcharts. When the computer program product is run in a computer system, the program codes are used to make the computer system implement the speech recognition method provided by the embodiments of the application.

[0193] The above functions defined in the system / apparatus of the embodiments of the application are performed when the computer program is executed by the processor 801. According to the embodiments of the application, the system, apparatus, module, unit, etc. described above can be implemented by computer program modules.

[0194] In one embodiment, the computer program can rely on a tangible storage medium such as an optical storage device, a magnetic storage device, etc. In another embodiment, the computer program can also be transmitted, distributed, and downloaded in the form of a signal on a network medium, and be downloaded and installed through the communication part 809, and / or installed from the detachable medium 811. The program codes contained in the computer program can be transmitted by any appropriate network medium, including but not limited to wireless, wired, etc., or any appropriate combination thereof.

[0195] In such embodiments, the computer program can be downloaded and installed from the network via the communication section 809, and / or installed from the removable media 811. When the computer program is executed by the processor 801, the above-described functions defined in the system of the embodiments of the present application are executed. According to the embodiments of the present application, the system, device, apparatus, module, unit, and the like described above can be realized by the computer program module.

[0196] According to the embodiments of the present application, the program code for executing the computer program provided by the embodiments of the present application can be written in any combination of one or more programming languages, and specifically, these computer programs can be implemented using high-level procedural and / or object-oriented programming language, and / or assembly / machine language. The programming language includes, but is not limited to, such as Java, C++, python, "C" language, or similar programming language. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In the case involving a remote computing device, the remote computing device can be connected to the user computing device through any kind of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, connected to the Internet through an Internet service provider).

[0197] The flowcharts and block diagrams in the drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowcharts or block diagrams can represent a module, a program segment, or a portion of code, which contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur in a different order than that shown in the figures. For example, two blocks noted in succession can actually be executed substantially concurrently, or they can sometimes be executed in reverse order, depending on the functionality involved. It should also be noted that each block in the flowcharts or block diagrams, and combinations of blocks in the flowcharts or block diagrams, can be implemented by dedicated hardware-based systems that perform the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0198] Those skilled in the art can understand that the features described in various embodiments of the present application can be combined and / or integrated in various combinations and / or integrations, even if such combinations or integrations are not explicitly described in the present application. In particular, the features described in various embodiments of the present application can be combined and / or integrated in various combinations and / or integrations without departing from the spirit and teachings of the present application. All such combinations and / or integrations fall within the scope of the present application.

Claims

1. A voice recognition method, characterized by, The voice recognition method comprises: In response to a voice recognition instruction, original audio data is acquired, the original audio data comprising a voice signal with background music noise; Time-frequency conversion is performed on the voice signal to obtain initial spectral features; The initial spectral features are input into a pre-trained voice separation module for background music suppression to output target spectral features; The target spectral features and the initial spectral features are input into a pre-trained voice recognition module to output recognized text; The voice separation module comprises an explicit music harmonic convolutional encoder, an implicit music perception encoder and a gated attention fusion processor, the explicit music harmonic convolutional encoder being configured to extract scale and harmonic structure features of background music noise in the initial spectral features to determine frame-level music features, the implicit music perception encoder being configured to perform voice enhancement and noise suppression on the initial spectral features to output frame-level voice features of the same dimension as the frame-level music features, and the gated attention fusion processor being configured to perform gated fusion on the frame-level music features and the frame-level voice features to output gated fusion features; the gated fusion features are used to generate a voice separation mask, and the voice separation mask is used to separate the initial spectral features to obtain the target spectral features; The voice recognition module comprises a noise perception attention mechanism and a convolutional self-attention hybrid encoder, the noise perception attention mechanism being configured to fuse the target spectral features and the initial spectral features through learnable gating weights to adaptively suppress residual music noise, and the convolutional self-attention hybrid encoder being configured to perform sequence modeling on the fused features to obtain the recognized text.

2. The voice recognition method of claim 1, wherein, The initial spectral features are input into the explicit music harmonic convolutional encoder to obtain frame-level music features; The initial spectral features are input into the implicit music perception encoder to obtain frame-level voice features; The frame-level music features and the frame-level voice features are input into a gated attention fusion processor to generate gated fusion features; The gated fusion features are input into a mask generator to predict a voice separation mask; The voice separation mask is used to perform frequency domain weighting on the initial spectral features to separate out the target spectral features. The initial spectral features are input into the explicit music harmonic convolutional encoder to obtain frame-level music features; 3. The voice recognition method of claim 2, wherein, The initial spectral features are input into the implicit music perception encoder to obtain frame-level voice features; The frame-level music features and the frame-level voice features are input into a gated attention fusion processor to generate gated fusion features; The gated fusion features are input into a mask generator to predict a voice separation mask; The voice separation mask is used to perform frequency domain weighting on the initial spectral features to separate out the target spectral features. The initial spectral features are input into the explicit music harmonic convolutional encoder to obtain frame-level music features; The initial spectral features are input into the implicit music perception encoder to obtain frame-level voice features; The frame-level music features and the frame-level voice features are input into a gated attention fusion processor to generate gated fusion features; The gated fusion features are input into a mask generator to predict a voice separation mask; The voice separation mask is used to perform frequency domain weighting on the initial spectral features to separate out the target spectral features. The target scale energy matrix is normalized and channel expanded to output the frame-level music feature.

4. The voice recognition method of claim 2, wherein, The initial spectrum feature is input into the implicit music perception encoder to obtain a frame-level speech feature, including: The initial spectrum feature is divided into a plurality of subbands, and a depth separable convolution is independently performed on each subband to extract a subband feature of each subband; The subband feature after splicing is input into a bandwidth controller to extract a time sequence feature of each subband, a channel attention mechanism is used to calculate a subband scaling factor based on the time sequence feature of each subband, the subband feature after splicing is weighted and enhanced and inhibited according to the subband scaling factor, and an enhanced subband feature corresponding to each subband is output through a residual connection. The enhanced subband feature is input into a frequency domain attention gate, attention weights of each subband are independently generated after the channel dimension is expanded and then weighted and corrected, and each subband feature is aggregated along the frequency dimension after a residual connection, and the frame-level speech feature is output.

5. The voice recognition method of claim 1, wherein, The speech separation module and the speech recognition module are pre-trained, including: obtaining a noisy speech sample, a clean speech sample, and a text label corresponding to the clean speech sample; performing time-frequency transformation on the noisy speech sample and the clean speech sample to generate a noisy spectrum feature and a clean spectrum feature; inputting the noisy spectrum feature into an initial speech separation module to generate a denoised predicted spectrum feature; inputting the predicted spectrum feature and the noisy spectrum feature into an initial speech recognition module to obtain a predicted text; determining a joint loss of the speech separation module and the speech recognition module based on the clean spectrum feature, the predicted spectrum feature, the predicted text, and the text label; and jointly training the initial speech separation module and the initial speech recognition module based on the joint loss to update the module parameters until convergence.

6. The voice recognition method of claim 5, wherein, The joint loss of the speech separation module and the speech recognition module based on the clean spectrum feature, the predicted spectrum feature, the predicted text, and the text label includes: calculating a speech separation loss according to the predicted spectrum feature and the clean spectrum feature; calculating a speech recognition loss according to the predicted text and the text label; and determining a joint loss according to the speech separation loss and the speech recognition loss.

7. The voice recognition method of claim 1, wherein, The input of the target spectrum feature and the initial spectrum feature into the pre-trained speech recognition module to output a recognition text includes: inputting the target spectrum feature and the initial spectrum feature into a noise perception attention mechanism to generate a fusion feature; inputting the fusion feature into a convolutional self-attention hybrid encoder to sequentially pass through a feedforward network, a multi-head self-attention, a depth separable convolution, and a residual normalization to output a high-order hidden representation; inputting the high-order hidden representation into a connection time sequence classification decoder to obtain a recognition text through linear classification and beam search.

8. The voice recognition method of claim 7, wherein, The input of the target spectrum feature and the initial spectrum feature into the noise perception attention mechanism to generate the fusion feature includes: splicing the target spectrum feature and the initial spectrum feature to generate a fusion input spectrum feature; performing one-dimensional convolution compression on the fusion input spectrum feature to obtain a compressed feature; Based on the compression feature, a gating weight is generated through a convolution layer and an activation function, the gating weight being used to represent a relative proportion of speech and noise in the original audio; The initial spectral feature is dynamically weighted using the gating weight to obtain a weighted spectral feature in which noise is suppressed; The weighted spectral feature is combined with a target spectral feature through an attention fusion mechanism to output a fused feature.

9. A speech recognition system characterized by The speech recognition system comprises: An acquisition module configured to acquire original audio data including a speech signal with background music noise in response to a speech recognition instruction; A time-frequency conversion module configured to perform time-frequency conversion on the speech signal to obtain an initial spectral feature; A target spectral feature determination module configured to input the initial spectral feature into a pre-trained speech separation module to suppress background music and output a target spectral feature; A text recognition module configured to input the target spectral feature and the initial spectral feature into a pre-trained speech recognition module to output a recognized text; The speech separation module comprises an explicit music harmonic convolutional encoder, an implicit music perception encoder, and a gated attention fusion processor. The explicit music harmonic convolutional encoder is configured to extract scale and harmonic structure features of background music noise in the initial spectral feature to determine frame-level music features. The implicit music perception encoder is configured to perform speech enhancement and noise suppression on the initial spectral feature to output frame-level speech features with the same dimension as the frame-level music features. The gated attention fusion processor is configured to perform gated fusion on the frame-level music features and the frame-level speech features to output a gated fusion feature. The gated fusion feature is used to generate a speech separation mask, and the speech separation mask is used to separate the initial spectral feature to obtain the target spectral feature. The speech recognition module comprises a noise perception attention mechanism and a convolutional self-attention hybrid encoder. The noise perception attention mechanism fuses the target spectral feature and the initial spectral feature through a learnable gating weight to adaptively suppress residual music noise. The convolutional self-attention hybrid encoder is configured to perform sequence modeling on the fused feature to obtain the recognized text.

10. An electronic device comprising: one or more processors; a memory for storing one or more computer programs, characterized in that the one or more processors execute the one or more computer programs to implement the steps of the speech recognition method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Voice separation method and device

    CN105096961A

  • Noise robustness speech recognition method based on double-gating atlas refining network

    CN118538203A