Speech recognition method and device, storage medium and program product
The audio classification model extracts local and global features and fuses them to generate fused speech features. Combined with the language recognition rules of the multilingual recognition model, the problem of misjudgment of the multilingual recognition model when dealing with similar feature languages is solved, and the accuracy and reliability of speech recognition are improved.
Patent Information
- Application Number
- CN202510315287.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-18
AI Technical Summary
When dealing with languages with similar characteristics, the multilingual recognition model is prone to misjudgment of language types, resulting in a decrease in recognition accuracy.
The local and global features of the audio signal to be identified are extracted through the pre-mounted audio classification model, and long-distance dependence is captured through the self-attention mechanism, fused to generate fused speech features, and voice recognition is performed in combination with the language recognition rules of the multilingual recognition model.
It improves the accuracy and reliability of speech recognition, avoids misjudgment by multilingual recognition models when dealing with similar characteristic languages, and significantly improves the recognition accuracy of specific languages.
Smart Images

Figure CN120340459A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a speech recognition method, device, storage medium, and program product. Background Art
[0002] In the current field of speech recognition technology, especially in multi-language recognition, the multi-language recognition model, as an advanced speech recognition system based on the Transformer architecture, has attracted wide attention in the industry due to its excellent performance and wide support for multiple languages. Without explicit language indication, the multi-language recognition model processes the input speech through its powerful feature learning ability to achieve automatic language recognition and speech-to-text conversion.
[0003] When performing speech recognition, the multi-language recognition model determines the final language recognition result based on the model's calculation results and certain decision rules. When processing languages with similar features, the multi-language recognition model may misjudge the language type, resulting in a decrease in recognition accuracy. Summary of the Invention
[0004] This application provides a speech recognition method, device, storage medium, and program product to at least solve the problem of low recognition accuracy of the multi-language recognition model in processing languages with similar features in the related art.
[0005] This application provides a speech recognition method, including: obtaining a continuous audio signal to be recognized; inputting the audio signal to be recognized into a pre-trained audio classification model to obtain the target language type output by the audio classification model; the audio classification model is used to extract local features and global features of the audio signal to be recognized, fuse the local features and the global features, and determine the language type to which the audio signal to be recognized belongs based on the fused speech features; inputting the audio signal to be recognized and the target language type into a pre-trained multi-language recognition model to obtain the text content output by the multi-language recognition model; the multi-language recognition model is used to recognize the audio signal to be recognized according to the language recognition rule corresponding to the target language type to obtain the text content corresponding to the audio signal to be recognized.
[0006] The present application also provides a voice recognition device, including: a voice acquisition module, configured to obtain continuous audio signals to be recognized; a language type recognition module, configured to input the audio signals to be recognized into a pre-trained audio classification model to obtain a target language type output by the audio classification model; the audio classification model is configured to extract local features and global features of the audio signals to be recognized, fuse the local features and the global features, and determine the language type to which the audio signals to be recognized belong based on the fused speech features; a voice transcription module, configured to input the audio signals to be recognized and the target language type into a pre-trained multi-language recognition model to obtain text content output by the multi-language recognition model; the multi-language recognition model is configured to recognize the audio signals to be recognized according to a language recognition rule corresponding to the target language type to obtain the text content corresponding to the audio signals to be recognized.
[0007] The present application also provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of any of the above voice recognition methods when executing the computer program.
[0008] The present application also provides a computer-readable storage medium, in which a computer program is stored, and wherein the computer program, when executed by a processor, implements the steps of any of the above voice recognition methods.
[0009] The present application also provides a computer program product, including a computer program, and the computer program, when executed by a processor, implements the steps of any of the above voice recognition methods.
[0010] Through this application, the language type of the audio signal to be recognized is first identified by a pre - placed audio classification model. During the process of identifying the language type, the audio classification model uses a deep convolutional architecture to extract local features of the audio signal to be recognized, and captures long - distance dependencies of the audio signal to be recognized through a self - attention mechanism to obtain global features. The outputs of the two are combined through an attention - based fusion method to construct fused speech features. This fusion strategy overcomes the limitations of convolutional neural networks in processing sequence data and also makes up for the possibly overlooked local information, thereby more accurately determining the language type of the audio signal to be recognized, and solving the problem that convolutional neural networks are difficult to capture the dependencies between parts that are far apart in the speech sequence; the target language type of the audio signal to be recognized is first identified by a pre - placed audio classification model, and then speech recognition is carried out specifically. During the recognition process, the target language type determined by the audio classification model is used as one of the input parameters of the multi - language recognition model. This means that the multi - language recognition model can intelligently select the most suitable language recognition rule to process the audio signal to be recognized according to the target language type provided by the audio classification model, avoiding the recognition deviation caused by directly receiving audio signals in an unspecified language, effectively avoiding misjudgment of the model, and solving the problem of low recognition accuracy of the multi - language recognition model when processing languages with similar features, improving the accuracy and reliability of speech recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] To more clearly illustrate the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0012] Figure 1 It is a schematic application diagram of a speech recognition method provided by an embodiment of the present application;
[0013] Figure 2 It is a flowchart of a speech recognition method provided by an embodiment of the present application;
[0014] Figure 3 It is a processing flowchart of an optional language recognition method according to an embodiment of the present application;
[0015] Figure 4 It is a processing flowchart of an optional audio classification model according to an embodiment of the present application;
[0016] Figure 5 It is a processing flowchart of an optional enhanced residual network according to an embodiment of the present application;
[0017] Figure 6It is a flowchart for training an optional audio classification model according to an embodiment of the present application;
[0018] Figure 7 It is a schematic structural diagram of a speech recognition device provided by an embodiment of the present application. Detailed implementation manners
[0019] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0020] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variation thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects, rather than to describe a specific order or sequence.
[0021] To enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0022] In view of the problem that the multi-language recognition model may misjudge the language type when processing languages with similar features, resulting in a decrease in the recognition accuracy, in the embodiments of the present application, the language type of the audio signal to be recognized is first identified by a pre-set audio classification model, and then speech recognition is performed specifically, effectively solving the problem that the multi-language recognition model misjudges similar language types, and significantly improving the recognition accuracy of specific languages.
[0023] According to one aspect of the embodiments of the present application, a speech recognition method is provided. Optionally, in this embodiment, the above speech recognition method may but is not limited to be applied to a hardware environment such as Figure 1 shown in the figure, including a terminal device 102 and a server 104. The server 104 can be connected to the terminal device 102 through a network, and can be used to provide services (such as application services, etc.) for the terminal device 102 or a client installed on the terminal device 102. A database can be set on the server 104 or independently of the server 104 to provide data storage services for the server 104.
[0024] The above network may include, but is not limited to, at least one of the following: a wired network, a wireless network. The above wired network may include, but is not limited to, at least one of the following: a wide area network, a metropolitan area network, a local area network. The above wireless network may include, but is not limited to, at least one of the following: WIFI (Wireless Fidelity), Bluetooth. The terminal device 102 may be, but is not limited to, a PC (Personal Computer), a mobile phone, a tablet computer, etc. The server 104 may be, but is not limited to, a cloud server, a server cluster, or other server types.
[0025] The speech recognition method according to the embodiment of the present application may be executed by the server 104, or may be executed by the terminal device 102, or may also be jointly executed by the server 104 and the terminal device 102. Among them, when the terminal device 102 executes the speech recognition method according to the embodiment of the present application, it may also be executed by a client installed thereon.
[0026] Taking the terminal device 102 executing the speech recognition method in this embodiment as an example, Figure 2 is a schematic flowchart of an optional speech recognition method according to the embodiment of the present application, as Figure 2 shown, the process of this method may include the following steps:
[0027] Step S202, obtain a continuous audio signal to be recognized.
[0028] Among them, the language type of the audio signal to be recognized is not limited to standard languages, but also widely covers various dialects and spoken language variants.
[0029] Step S204, input the audio signal to be recognized into a pre-trained audio classification model to obtain the target language type output by the audio classification model; the audio classification model is used to extract the local features and global features of the audio signal to be recognized, fuse the local features and global features, and determine the language type to which the audio signal to be recognized belongs based on the fused speech features obtained by the fusion.
[0030] Among them, the audio signal to be recognized is a kind of data with time series characteristics, and there are complex dependencies between different parts of the speech. These dependencies have an important impact on accurate audio classification. However, in the field of audio processing and analysis, the convolution operation of convolutional neural networks (CNNs) is essentially local, and it can only process the speech signal within a limited receptive field, making it difficult to capture the dependencies between parts that are far apart in the speech sequence. Therefore, to solve this problem, in the embodiment of the present application, the audio classification model captures and establishes the global features of the audio through the self-attention mechanism, combines the local features and global features of the audio signal to be recognized, so as to better model the global structure of the audio signal to be recognized, and solves the problem of the limited receptive field of CNNs.
[0031] The function of the audio classification model is to receive the audio signal to be recognized, identify and classify the language type of the audio through feature extraction and fusion, and provide accurate language information for subsequent speech recognition. Among them, local features usually refer to the characteristics of the audio signal to be recognized within a short-time window, such as timbre, pitch, etc. Local features are obtained from the audio waveform and spectrogram through the local connection and weight sharing mechanisms of the convolutional layer, providing a fine-grained analysis basis for the speech signal. Global features refer to the features extracted from the entire time series of the audio signal to be recognized, such as context semantic information, long-distance dependencies, etc. Global features calculate the dependencies between any two positions in the audio signal to be recognized through the self-attention mechanism, thereby capturing the global structure of the audio signal to be recognized. The fused speech features are the comprehensive feature representations obtained by the local features and the global features through an attention-based fusion method. It contains both fine-grained local information and comprehensive global context, providing richer and more accurate decision-making basis for the audio classification model to determine the language type of the audio signal to be recognized. Among them, the attention-based fusion method can be to use the local features as the query vector, the global features as the key vector and the value vector, calculate the attention weights, and perform weighted summation on the value vector of the global features to obtain the fused speech features after fusion. This process fuses the local and global feature information and constructs a more comprehensive speech feature description.
[0032] The target language type refers to the specific language or dialect category to which the audio signal to be recognized identified by the audio classification model belongs.
[0033] Optionally, Figure 3 is a processing flowchart of an optional language recognition method according to an embodiment of the present application. As Figure 3 shown, the terminal device inputs the audio signal to be recognized into a pre-trained audio classification model. The audio classification model extracts the local features of the audio signal to be recognized through deep convolutional operations, and captures the global features of the audio signal to be recognized through the self-attention mechanism. Based on the attention-based fusion method, the extracted local features are fused with the extracted global features to obtain the fused speech features. The fused speech features are passed to the output layer of the audio classification model, usually a classification layer, and a probability distribution is calculated using an activation function, thereby outputting the target language type to which the audio signal to be recognized is most likely to belong, such as a dialect or Mandarin.
[0034] Step S206, input the audio signal to be recognized and the target language type into a pre-trained multi-language recognition model to obtain the text content output by the multi-language recognition model; the multi-language recognition model is used to recognize the audio signal to be recognized according to the language recognition rules corresponding to the target language type to obtain the text content corresponding to the audio signal to be recognized.
[0035] Among them, the multilingual recognition model is a deep learning model based on the Transformer architecture, specifically designed to handle multilingual speech recognition tasks. The multilingual recognition model learns the speech feature representations of multiple languages through pre-training on a large-scale multilingual audio dataset, and can recognize and transcribe the speech of multiple languages including dialects and Mandarin, and output the corresponding text content. For example, the multilingual recognition model can be the Whisper model. The Whisper model is a multilingual speech recognition model based on the Transformer architecture. It can automatically recognize and transcribe speech inputs including dialects, Mandarin and other languages without prior language specification. During the full-parameter fine-tuning training process of the Whisper model using dialect and Mandarin speech datasets, the training rounds, learning rate and other parameters of the Whisper model are continuously adjusted, so that the Whisper model can better learn the characteristics and rules of dialect and Mandarin speech, continuously optimize the performance of the model, and specifically improve the speech recognition effect. However, when the multilingual recognition model processes languages with similar characteristics, it is very likely to misjudge the input speech as other similar languages. This misjudgment will not only significantly reduce the accuracy of speech recognition, but also have a serious negative impact on downstream applications that rely on accurate speech recognition results. Therefore, to solve this problem, in the embodiments of the present application, the audio classification model is used to assist the multilingual recognition model to pre-judge the target language type of the input audio signal to be recognized, and then the multilingual recognition model performs speech recognition specifically, effectively reducing the misjudgment rate, solving the problem of misjudging similar language types by the multilingual recognition model, significantly improving the recognition accuracy of specific languages, and thus ensuring the normal operation of downstream applications that rely on accurate speech recognition results.
[0036] The text content refers to the text output obtained by the multilingual recognition model converting the audio signal to be recognized. It can be understood that the text content is the result of the model recognizing the speech information in the audio, usually presented in a human-readable text form. For example, in the example where the multilingual recognition model is the Whisper model, a dialect speech will be transcribed in the form of Chinese text after being processed by the Whisper model for subsequent analysis and application by downstream applications.
[0037] Optionally, as Figure 3 shown, the terminal device inputs the target language type recognized by the audio classification model and the audio signal to be recognized into the pre-trained multilingual recognition model. The multilingual recognition model selects the corresponding language recognition rule to process the audio signal according to the received target language type information, and outputs the text content corresponding to the audio signal to be recognized.
[0038] Through the embodiments of the present application, the language type of the audio signal to be recognized is first identified by a pre-set audio classification model. During the process of identifying the language type, the audio classification model uses a deep convolutional architecture to extract local features of the audio signal to be recognized, and captures long-range dependencies of the audio signal to be recognized through a self-attention mechanism to obtain global features. The outputs of the two are combined through an attention-based fusion method to construct fused speech features. This fusion strategy overcomes the limitations of convolutional neural networks in processing sequence data, and also makes up for the possibly ignored local information, thereby more accurately determining the language type of the audio signal to be recognized, and solving the problem that it is difficult for convolutional neural networks to capture the dependencies between relatively distant parts in the speech sequence; the language type of the audio signal to be recognized is first identified by a pre-set audio classification model, and then speech recognition is carried out specifically. During the recognition process, the target language type determined by the audio classification model is used as one of the input parameters of the multi-language recognition model. This means that the multi-language recognition model can intelligently select the most appropriate language recognition rule to process the audio signal to be recognized according to the target language type provided by the audio classification model, avoiding recognition biases caused by directly receiving audio signals of unspecified languages, effectively avoiding misjudgments of the model, and solving the problem of low recognition accuracy of the multi-language recognition model when processing languages with similar features, and improving the accuracy and reliability of speech recognition.
[0039] In an exemplary embodiment, the text content output by the multi-language recognition model is applied to downstream applications, where the downstream applications can be application products in fields such as healthcare and culture and tourism. Before transmitting the text content output by the multi-language recognition model to the downstream applications, the above speech recognition method further includes:
[0040] Performing semantic verification on the text content output by the multi-language recognition model, and when the text content meets the preset requirements, transmitting the text content output by the multi-language recognition model to the downstream applications.
[0041] Among them, the purpose of performing semantic verification on the text content is to ensure the semantic coherence and correctness of the text content, and avoid problems such as operation errors, decision-making biases, or degradation of user experience caused by downstream applications misunderstanding or being unable to parse the text information.
[0042] During the process of performing semantic verification on the text content, a semantic verification method combined with a knowledge graph can be adopted, and this method specifically includes the following steps:
[0043] The text content output by the multilingual recognition model is input into a pre-trained entity recognition model to obtain multiple target entities in the text content. Query whether there are reference entities in the knowledge graph that match each of the multiple target entities. If so, further verify through the knowledge graph whether the types, attributes, and related relationships of the target entities are consistent with the description in the text content. When the descriptions are consistent, transmit the text content output by the multilingual recognition model to the downstream application.
[0044] Among them, the knowledge graph refers to a structured knowledge database applicable to the downstream application, which stores information in the form of triples of entities, relationships, and attributes. The construction process of the knowledge graph includes: First, collect a large amount of text materials in a specific field, and use natural language processing technologies such as named entity recognition (NER) and dependency analysis to automatically or semi-automatically identify entities (such as drugs, legal provisions, cases, etc.) and the relationships between entities (such as the citation relationship of legal provisions, etc.) from the collected text materials, and further extract attribute information related to the entities, such as the ingredients of drugs, the effective date of legal provisions, etc. Finally, store the identified entities, relationships, and attributes in a graph database in a structured form to form a knowledge graph. Each entity node can store information such as its name, type, and attributes, while the relationship node represents the connection between entities, including the relationship type and attributes.
[0045] The entity recognition model is a model trained based on deep learning technology, which can map the identified text content to the entities and relationships in the knowledge graph. The training process of the entity recognition model includes the following steps: Collect a large amount of labeled text data, where each entity is clearly labeled with its type and boundary. Convert the text into a format that the model can process, such as word embeddings or character embeddings, to capture the semantic information and context features of the vocabulary. Use the labeled data to train the entity recognition model, adjust the weights through backpropagation, and optimize the loss function, such as cross-entropy loss, to improve the accuracy of entity recognition. Evaluate the model performance on the validation set and adjust the hyperparameters; finally evaluate the generalization ability of the model on the test set to ensure its good performance on unseen data. According to the feedback from validation and testing, continuously iterate the model, which may include data augmentation, adjusting the architecture, or optimizing the training strategy, until a satisfactory recognition effect is achieved to obtain a trained entity recognition model.
[0046] Through this embodiment, the text content output by the multilingual recognition model is verified using the entity relationships in the knowledge graph to ensure the logical coherence of the text content output by the multilingual recognition model, avoid errors in context mismatch, provide high-quality text input that has been semantically verified for the downstream application, and reduce processing biases caused by incorrect information.
[0047] In an exemplary embodiment, Figure 4It is a processing flowchart of an optional audio classification model according to an embodiment of the present application. As Figure 4 shown, the audio classification model includes an enhanced residual network, a transformer network, and a feature fusion network.
[0048] Among them, the enhanced residual network (EResNet) is an improvement and enhancement of ResNet (residual network), aiming to optimize the performance of deep convolutional neural networks in audio signal processing. EResNet enhances local information interaction and fine-grained feature extraction by introducing an attention mechanism and local feature fusion (LFF) technology, so as to more effectively process the time-frequency features in audio signals.
[0049] The enhanced residual network uses convolutional neural networks (CNNs). The EResNet model based on CNNs essentially obtains information through convolutional operations. Therefore, it can only process speech signals within a limited receptive field. Although it obtains global information by selecting multi-scale features at different stages, downsampling, and concatenating after expanding channels. However, due to the limitations of convolutional operations, it is still difficult to capture the dependencies between parts that are far apart in the speech sequence. Therefore, to solve this problem, this embodiment uses a transformer network. Through the self-attention mechanism, it can directly calculate the dependencies between any two positions in the audio signal to be recognized, so as to better model the global structure of the audio signal to be recognized.
[0050] The transformer network is a model based on the self-attention mechanism and the Transformer architecture. Compared with traditional convolutional neural networks, the transformer network based on Transformer can perform correlation analysis on each part of the audio signal to be recognized globally to obtain the semantic connection between syllables that are far apart. In the audio classification model, the transformer network captures the long-range dependency information of the speech signal by calculating the dependencies between any two positions in the sequence, making up for the limitations of EResNet in processing sequence data and enhancing the audio classification model's ability to understand global features.
[0051] The feature fusion network in the audio classification model is responsible for fusing different types of features extracted by EResNet and the transformer network. It uses an attention-based mechanism to perform weighted fusion of local features and global features to generate more comprehensive fused speech features, thereby improving the accuracy of speech classification and the decision-making ability of the model.
[0052] In some embodiments, the audio signal to be recognized is input into a pre-trained audio classification model, and the audio classification model performs the following detection operations on the audio signal to be recognized to obtain the target language type output by the audio classification model:
[0053] First, convert the audio signal to be recognized into a two-dimensional spectrogram.
[0054] Among them, the audio signal is essentially a one-dimensional time-domain signal. In order to more effectively extract and represent the features of the audio signal, in this embodiment, the audio signal to be recognized is converted into a two-dimensional spectrogram, which can display the time-domain information and frequency-domain information of the signal at the same time.
[0055] Optionally, as Figure 4 shown, the terminal device performs a fast Fourier transform (FFT) on the audio signal to be recognized, converts the time-domain signal into a frequency-domain signal, and obtains a two-dimensional spectrogram.
[0056] II. Extract features from the two-dimensional spectrogram through an enhanced residual network to obtain the local features of the audio signal to be recognized.
[0057] Optionally, as Figure 4 shown, the terminal device uses an enhanced residual network, uses attention feature fusion in the residual connection, groups and dynamically weights and fuses the two-dimensional spectrogram, then selects multi-scale features at different stages, performs downsampling, expands the channels and then splices the feature representation, and finally calculates and fuses the attention weights to obtain the local features of the audio signal to be recognized, denoted as C.
[0058] III. Extract features from the two-dimensional spectrogram through a transformer network to obtain the global features of the audio signal to be recognized.
[0059] Optionally, as Figure 4 shown, the terminal device uses a transformer network to unfold the two-dimensional spectrogram to obtain a one-dimensional vector sequence, and captures the long-range dependence relationship in the one-dimensional vector sequence through a self-attention mechanism to obtain the global features of the audio signal to be recognized, denoted as T.
[0060] IV. Fuse the local features and global features through a feature fusion network to obtain fused speech features.
[0061] Optionally, as Figure 4 shown, the terminal device uses a feature fusion network, takes the local feature C as the query vector, the global feature T as the provider of the key and the provider of the value, calculates the attention score between the query vector and the provider, normalizes the attention score through an activation function (such as the softmax function) to obtain the attention weight W corresponding to multiple value vectors in the global feature, and weights and sums multiple value vectors in the global feature T according to the attention weight W corresponding to multiple value vectors in the global feature to obtain the fused speech features, denoted as H.
[0062] V. Perform classification probability prediction on the fused speech features to obtain the probability prediction values of the audio signal to be recognized belonging to different language types.
[0063] Optionally, asFigure 4 As shown, the terminal device inputs the fused speech feature H into the activation function to obtain the probability prediction values of the audio signal to be recognized belonging to different language types. Among them, in an example where the language types include dialects and Mandarin, the fused speech feature H can be input into the sigmoid function (S-shaped function, an activation function) to obtain the probability prediction value of the audio signal to be recognized belonging to the dialect, and the probability prediction value of belonging to Mandarin. Among them, the output value range of the sigmoid function is between 0 and 1, which can map any real value to this interval and is often used in binary classification problems to convert the output of the model into a probability form.
[0064] VI. Determine the target language type to which the audio signal to be recognized belongs according to the probability prediction values of the audio signal to be recognized belonging to different language types.
[0065] Optionally, the terminal device determines the language type corresponding to the maximum probability prediction value as the target language type to which the audio signal to be recognized belongs.
[0066] Through this embodiment, by adopting the method of attention fusion between the Transformer-based transducer network and the CNNs-based EResNet model, making full use of the self-attention mechanism of the Transformer-based transducer network, the advantage of being able to directly calculate the dependence relationship between any two positions in the audio signal to be recognized can be utilized, overcoming the limitation that the convolutional operation of the EResNet model can only process speech signals within a limited receptive field and is difficult to capture the dependence relationship of relatively distant parts, so as to better model the global structure of the audio signal to be recognized and improve the processing ability of speech signals.
[0067] In an exemplary embodiment, converting the audio signal to be recognized into a two-dimensional spectrogram includes the following steps:
[0068] Segment the continuous audio signal to be recognized into multiple time frames; the time frames in the multiple time frames contain the time-domain information of the audio signal to be recognized; perform Fourier transform on the multiple time frames to obtain the frequency-domain signals corresponding to the multiple time frames, and generate an initial spectrogram based on the frequency-domain signals corresponding to the multiple time frames; filter the initial spectrogram through a set of Mel filters to obtain the Mel spectrogram, and perform logarithmic processing on the Mel spectrogram to obtain the two-dimensional spectrogram.
[0069] Among them, the process of segmenting the continuous audio signal into a series of short-time segments is called framing, and each segment is a time frame. The setting of the time frame helps subsequent spectral analysis and local feature extraction.
[0070] The initial spectrum refers to the frequency-domain representation of an audio signal obtained after Fourier transform. The initial spectrum is a two-dimensional matrix, where the rows represent frequencies, the columns represent time, and the values of the matrix represent the energy levels at the corresponding frequency and time points.
[0071] To make the perception of the audio classification model closer to that of humans, in this embodiment, a set of Mel filters is used to filter the initial spectrum. Among them, a set of Mel filters is a collection of filters for extracting audio features, designed based on the way the human ear perceives frequencies (i.e., Mel frequencies), used to filter the initial spectrum, extract features related to human ear perception, and generate a Mel spectrum, making the feature extraction closer to human ear perception. The Mel spectrum converts the initial spectrum from a linear frequency scale to a Mel frequency scale, making the spectrogram closer to the perception characteristics of the human ear. The abscissa of the Mel spectrum represents time, and the ordinate represents frequency, but the Mel frequency scale is used instead of the original linear frequency scale.
[0072] The energy distribution of the original audio signal to be recognized may be very wide, from very weak low energy to very strong high energy, which makes it difficult for the audio classification model to handle this wide dynamic range during learning. Therefore, to solve this problem, in this embodiment, the Mel spectrum is logarithmically processed to compress the dynamic range of the audio signal to be recognized and stabilize the features. The logarithmic processing can compress this wide dynamic range into a smaller range, enabling the audio classification model to more stably extract features and classify the audio signal. In addition, the logarithmic transformation can also highlight low-intensity components, facilitating the audio classification model to capture details in speech (such as whispers and intonation changes).
[0073] Optionally, the terminal device divides the continuous audio signal into a series of time frames at a fixed time length (e.g., 25 ms), and sets an overlap (usually 10 - 15 ms) between each frame to avoid information loss and ensure the continuity of speech features. The terminal device performs a fast Fourier transform (FFT) on each time frame, converting the time-domain signal to the frequency domain, obtaining frequency-domain signals corresponding to multiple time frames, and integrating the frequency-domain signals of all time frames to generate an initial spectrogram of the audio signal. The terminal device uses a set of Mel filters to filter the initial spectrum, converting the linear spectrum to a spectrum on the Mel frequency scale to obtain a Mel spectrum, and performing a logarithmic transformation on each value in the Mel spectrum to compress the dynamic range and enhance the visibility of low-energy signals, obtaining a two-dimensional spectrogram.
[0074] In this embodiment, the Mel filter is used to filter the initial spectrum to obtain the Mel spectrum. Based on the human ear's perception characteristics of frequency, the linear spectrum can be converted to the Mel frequency scale that is more in line with human hearing, so as to distinguish different language types based on the subtle differences in frequency. Further, taking the logarithm of the Mel spectrum can compress the dynamic range of the audio signal, making the characteristics of low-intensity signals more obvious, ensuring that the audio classification model can evenly learn the characteristics of all intensity levels, and improving the ability to accurately classify speech signals.
[0075] In an exemplary embodiment, the enhanced residual network includes multiple feature extraction layers and multiple feature fusion layers. Among them, the multiple feature extraction layers in the enhanced residual network constitute the local feature fusion (LFF) branch. LFF introduces an attention feature fusion (AFF) module in the residual connection, divides the feature map into groups to obtain multiple feature groups, and dynamically weights and fuses the multiple feature groups to enhance local information interaction and fine-grained feature extraction. The multiple feature fusion layers in the enhanced residual network constitute the global feature fusion (GFF) branch. GFF selects multi-scale features at different stages, downsamples, expands the channels and then splices them, and calculates the attention weights through the AFF module for fusion to enhance global feature interaction.
[0076] The multiple feature fusion layers correspond to the non-first feature extraction layers in the multiple feature extraction layers in sequence. Figure 5 It is a processing flow chart of an optional enhanced residual network according to an embodiment of the present application. As Figure 5 shown, in this embodiment, the enhanced residual network includes four feature extraction layers and three feature fusion layers. The four feature extraction layers are respectively represented by Stage1~Stage4, and these four feature extraction layers constitute the local feature fusion (LFF). The three feature fusion layers are respectively denoted as Fusion1~Fusion3, and these three feature fusion layers constitute the global feature fusion (GFF).
[0077] In some embodiments, the enhanced residual network is used to extract features from the two-dimensional spectrogram to obtain the local features of the audio signal to be recognized, including:
[0078] First, the first feature extraction layer in the multiple feature extraction layers extracts features from the two-dimensional spectrogram to obtain the output of the first feature extraction layer.
[0079] Among them, the structure and processing flow of each feature extraction layer can be the same. As Figure 5As shown in the figure, each feature extraction layer includes multiple convolutional layers. Each feature extraction layer among the multiple feature extraction layers is used as the current feature extraction layer for explanation. For the first convolutional layer in the current feature extraction layer, its input is the output of the previous feature extraction layer of the current feature extraction layer. The output of the first convolutional layer in the current feature extraction layer serves as the input of the next convolutional layer in the current feature extraction layer, and so on until the last convolutional layer. Then, according to the Attention Feature Fusion (AFF) module, the first attention weights corresponding to the multiple convolutional layers in the current feature extraction layer are calculated, and the outputs of each convolutional layer in the current feature extraction layer are weighted and summed to obtain the output of the current feature extraction layer. Among them, the first attention weights corresponding to the multiple convolutional layers in the current feature extraction layer are dynamically generated by enhancing the attention mechanism inside the residual network, reflecting the contribution degree of each feature group to the output representation of the current feature extraction layer in the context of the current feature extraction layer.
[0080] Optionally, as Figure 5 shown, the terminal device inputs the two-dimensional spectrogram into the first convolutional layer of the first feature extraction layer among the multiple feature extraction layers. The first convolutional layer extracts features from the two-dimensional spectrogram to obtain the output of the first convolutional layer (i.e., a feature group), and then uses the output of the first convolutional layer as the input of the next convolutional layer, and so on until the last convolutional layer to obtain multiple feature groups. Then, according to the Attention Feature Fusion (AFF) module, the first attention weights corresponding to the multiple convolutional layers in the first feature extraction layer are calculated, and the outputs of each convolutional layer in the first feature extraction layer (i.e., multiple feature groups) are weighted and summed to obtain the output of the first feature extraction layer.
[0081] Second, take the non-first feature extraction layers among the multiple feature extraction layers as the current feature extraction layer in turn and perform the following feature extraction operations: The current feature extraction layer extracts features from the output of the previous feature extraction layer of the current feature extraction layer to obtain multiple feature groups, and according to the first attention weights corresponding to the multiple feature groups, the multiple feature groups are weighted and summed to obtain the output of the current feature extraction layer.
[0082] Among them, except that the input of the non-first feature extraction layer is different from the input of the first feature extraction layer, the remaining structure and processing flow are the same as those of the first feature extraction layer. Therefore, the processing flow of the non-first feature extraction layer will not be repeated here.
[0083] Third, the first feature fusion layer among the multiple feature fusion layers performs feature fusion on the output of the first feature extraction layer and the output of the next feature extraction layer of the first feature extraction layer to obtain the output of the first feature fusion layer.
[0084] Optionally, during the process of the first feature fusion layer fusing the output of the first feature extraction layer and the output of the next feature extraction layer of the first feature extraction layer, the first feature fusion layer selects multi-scale feature maps at different stages. This step is to capture the information of the audio signal at different frequency and time scales. Then, in order to unify the feature maps of different scales to the same size, a downsampling operation is performed to reduce the size of the high-resolution feature map to match the low-resolution feature map. At the same time, by increasing the number of channels of the feature map (i.e., the dimension of the feature), channel expansion is performed to enrich the feature representation, so that subsequent fusion operations can be based on features with more dimensions. Further, the feature maps after downsampling and channel expansion are concatenated together to form a multi-channel feature representation (i.e., the concatenated feature groups), and the attention weights of each concatenated feature group are calculated through the AFF module. The attention weights reflect the importance of each feature group in the global feature representation. Finally, according to the attention weights calculated by the AFF module, weighted summation is performed on the concatenated multi-channel feature representation to obtain the output of the first feature fusion layer.
[0085] IV. Take the non-first feature fusion layers in multiple feature fusion layers as the current feature fusion layer in turn to perform the following feature fusion operations. The feature extraction layer corresponding to the current feature fusion layer is the current feature extraction layer: fuse the output of the current feature extraction layer and the output of the previous feature fusion layer of the current feature fusion layer through the current feature fusion layer to obtain the output of the current feature fusion layer.
[0086] Among them, except that the input of the non-first feature fusion layer is different from the input of the first feature fusion layer, the remaining structure and processing flow are the same as those of the first feature fusion layer. Therefore, the processing flow of the non-first feature fusion layer will not be repeated here.
[0087] V. Perform pooling and fully connected processing on the output of the last feature fusion layer in multiple feature fusion layers to obtain the local features of the audio signal to be recognized.
[0088] Through this embodiment, the output of the non-first feature extraction layer is not simply directly used by subsequent layers, but through interaction with the corresponding feature fusion layer, and feature weighted fusion is performed using the attention mechanism, so that the enhanced residual network can dynamically adjust the importance of different features while extracting features, thereby optimizing the feature representation. Secondly, in the process of feature fusion of the non-first feature fusion layer, the output of the current feature extraction layer and the output of the previous feature fusion layer are considered, which not only enhances the hierarchy of the features, but also realizes the dynamic fusion of the features through the attention mechanism. Compared with traditional feature fusion methods, it can more accurately capture the dependency relationships between far-apart parts in the speech signal and overcome the limitations of traditional CNNs in processing languages with similar features.
[0089] In an exemplary embodiment, the transformer network includes a plurality of encoding blocks. The structures of the plurality of encoding blocks are the same, and the output of each encoding block serves as the input of the next encoding block, enabling information to be deepened and transmitted layer by layer, which helps the model learn more complex feature representations and deeper context dependencies. The encoding blocks are responsible for converting the input sequence into a series of continuous representations. Each encoding block in the plurality of encoding blocks includes a multi-head self-attention network and a feed-forward neural network for learning the context dependencies and feature representations of the input sequence.
[0090] The multi-head self-attention network can simultaneously calculate the interdependencies of elements in the input sequence from multiple different attention heads, thereby capturing the features of the sequence in multiple dimensions and enhancing the expressive power of the model. In this embodiment, the multi-head self-attention network is used to model the global structure of the speech sequence and capture the dependencies between remotely separated parts.
[0091] The feed-forward neural network is used to perform a non-linear transformation on the features after the self-attention operation to further optimize the feature representation, so that the model can better handle the spectral characteristics of the audio signal and improve the classification performance. In this embodiment, the feed-forward neural network includes two linear layers and an activation function, and the output of the multi-head self-attention network is further subjected to feature transformation and learning through the feed-forward neural network.
[0092] In some embodiments, the two-dimensional spectrogram is subjected to feature extraction through the transformer network to obtain the global features of the audio signal to be recognized, including:
[0093] First, through the multi-head self-attention network of the first encoding block among the plurality of encoding blocks, the two-dimensional spectrogram is converted into a one-dimensional vector sequence, and an encoding operation is performed on the one-dimensional vector sequence to obtain the output features of the first encoding block.
[0094] Among them, since the audio signal to be recognized is a one-dimensional signal, it usually becomes a two-dimensional spectrogram after being processed (such as converted into a Mel spectrogram), while the transformer network processes one-dimensional sequence data. Therefore, the two-dimensional spectrogram needs to be unfolded into a one-dimensional vector sequence. The one-dimensional vector sequence refers to the representation form of the audio signal processed by the self-attention mechanism, which converts the complex structure of the two-dimensional spectrogram into a sequence of eigenvalue arranged in time order. The one-dimensional vector sequence can capture the temporal continuity and dependencies, enabling the self-attention mechanism to calculate the dependencies between any two positions in the sequence.
[0095] Optionally, the terminal device unfolds the two-dimensional spectrogram to obtain a one-dimensional vector sequence. The multi-head self-attention network of the first coding block among multiple coding blocks calculates the second attention weights between any two elements in the one-dimensional vector sequence through multiple parallel attention heads. According to the second attention weights between any two elements, the output feature vectors of multiple attention heads are calculated, and the output feature vectors of multiple attention heads are concatenated to obtain the multi-head attention output. The multi-head attention output is linearly transformed through the feed-forward neural network of the first coding block to obtain the output feature of the first coding block.
[0096] Among them, the second attention weight refers to the degree of relationship between any two elements in the one-dimensional vector sequence, which is used to guide the calculation of the multi-head attention output and the weighting of features. The second attention weight reflects the mutual correlation strength of information at different time points in the audio signal to be recognized.
[0097] Second, each non-first coding block among multiple coding blocks is used as the current coding block to perform the following coding operations to obtain the output features of multiple coding blocks. The multi-head self-attention network of the current coding block is the current multi-head self-attention network, and the feed-forward neural network in the current coding block is the current feed-forward neural network: through multiple parallel attention heads, calculate the second attention weights between any two elements in the output features of the previous coding block of the current coding block. According to the second attention weights between any two elements, calculate the output feature vectors of multiple attention heads, and concatenate the output feature vectors of multiple attention heads to obtain the multi-head attention output. The multi-head attention output is linearly transformed through the current feed-forward neural network to obtain the output feature of the current coding block. The output feature of the last coding block is the global feature of the audio signal to be recognized.
[0098] Among them, multiple attention heads refer to in the multi-head self-attention mechanism, the self-attention operation is independently executed in multiple heads (subspaces) in parallel, and each head focuses on different aspects of the input sequence, thereby calculating multiple independent attention weight matrices. In this embodiment, multiple attention heads are used to calculate the second attention weights between any two elements in the output features of the previous coding block, enhancing the multi-angle understanding of the input audio signal by the transformer network.
[0099] The output features of the previous coding block refer to the feature representation generated by the previous coding block before the current coding block processes. It carries the feature information of the audio signal in the previous processing stage, provides input data for the current coding block, and is used for further feature learning and extraction.
[0100] The output feature vectors of multiple attention heads are the feature representations calculated by each attention head, reflecting the features and mutual relationships of elements in the sequence.
[0101] The output of multi-head attention refers to concatenating or fusing the output feature vectors of multiple attention heads to retain the feature information obtained from different subspaces. In this embodiment, the output of multi-head attention is obtained by concatenating the output feature vectors of multiple attention heads, forming a multi-dimensional feature representation, which provides richer and more comprehensive information for subsequent feature learning.
[0102] Optionally, the terminal device uses non-first encoding blocks among multiple encoding blocks as the current encoding block respectively to perform the following encoding operations to obtain the output features of multiple encoding blocks. The multi-head self-attention network of the current encoding block is the current multi-head self-attention network, and the feed-forward neural network in the current encoding block is the current feed-forward neural network: The terminal device processes the output features of the previous encoding block in parallel through multiple attention heads in parallel, calculates the second attention weights between any two elements in the sequence, and based on the calculated second attention weights, each attention head generates its specific output feature vector. The output feature vectors of all attention heads are concatenated to form the output of multi-head attention, and the output of multi-head attention is passed to the current feed-forward neural network for linear transformation and non-linear activation to obtain the output features of the current encoding block. Repeat the above steps until all encoding blocks are processed. The output features of the last encoding block are the global features of the audio signal to be recognized.
[0103] Through this embodiment, multiple encodings are stacked in the transformer network. Each encoding block contains a multi-head self-attention network and a feed-forward neural network. The output of each encoding block will be used as the input of the next encoding block, enabling information to be deepened and transmitted layer by layer, which helps the model learn more complex feature representations and deeper context dependencies; repeating the encoding operations of multi-head self-attention and feed-forward neural network, calculating and fusing the second attention weights between different elements, generating the output of multi-head attention, and then performing linear transformation through the feed-forward network to obtain the output features of deeper encoding blocks. The process of capturing long-range dependencies through the multi-head self-attention network not only improves the richness and depth of feature expression but also solves the limitations of CNNs in processing time-series data, enabling the model to better understand the internal structure of the audio signal to be recognized.
[0104] In an exemplary embodiment, through multiple attention heads in parallel, calculate the second attention weights between any two elements in the output features of the previous encoding block of the current encoding block. According to the second attention weights between any two elements, calculate the output feature vectors of multiple attention heads, including:
[0105] 1. Perform a linear transformation on the output features of the previous encoding block of the current encoding block through the current multi-head self-attention network to obtain query vectors, key vectors, and value vectors. Split the query vectors, key vectors, and value vectors into multiple attention heads respectively, and calculate the second attention weights between any two elements through the multiple attention heads.
[0106] Among them, in the multi-head self-attention mechanism, the query vector represents the query demand of a specific element in the sequence for other elements, and is used to calculate the correlation between this element and other elements. In this embodiment, the query vector is obtained by linear transformation of the output features of the previous encoding block and is used for subsequent attention weight calculation.
[0107] The key vector carries the feature information of each element in the sequence and is used to match with other elements (represented by the query vector) to calculate the similarity between them. In this embodiment, the key vector is also generated by linear transformation of the output features of the previous encoding block and participates in the similarity calculation in the attention mechanism.
[0108] The value vector contains the actual information of each element in the sequence. When performing attention weighted summation, the value vector is weighted and merged according to the matching degree with the query vector to generate the final attention output. In this embodiment, the value vector is transformed from the output features of the previous encoding block and is used for information integration according to the calculated attention weights.
[0109] Optionally, the terminal device performs a linear transformation on the feature vector output by the previous encoding block of the current encoding block through the multi-head self-attention network in the current encoding block to respectively generate query vectors, key vectors, and value vectors; split the generated query vectors, key vectors, and value vectors into multiple sub-vectors, each sub-vector corresponding to an attention head, to achieve parallel processing, and each attention head independently calculates the second attention weight between its corresponding query vector and key vector.
[0110] 2. Calculate the attention scores corresponding to the second attention weights between any two elements through dot product, and perform normalization processing on the attention scores corresponding to the multiple attention heads to obtain the weight matrices of the multiple attention heads.
[0111] Among them, the attention score is the result obtained by calculating the dot product of the query vector and the key vector, which reflects the similarity or correlation between any two elements in the sequence. In this embodiment, the attention score is calculated through dot product to measure the similarity between the query vector and the key vector, and assigns a preliminary weight value to each element to guide the transformer network to perform weighted summation on the value vector.
[0112] The weight matrix is a matrix obtained by normalizing the attention scores. The value of each element represents its relative weight in the sequence and is used for weighted summation calculation. In this embodiment, the weight matrix is composed of the attention scores of multiple attention heads. Through normalization, it is ensured that the importance of each element is taken into account when the feature vectors output by all attention heads are fused, avoiding the omission of some elements and improving the accuracy of feature fusion.
[0113] Optionally, the terminal device performs a dot product operation on the output features of the previous encoding block of the current encoding block, that is, the query vector Q and the key vector K that have been segmented and linearly transformed, to calculate the attention scores between any two elements, obtaining an attention score matrix. The rows of this matrix represent the respective elements of the query vector, the columns represent the respective elements of the key vector, and each value in the matrix is the corresponding attention score; the obtained attention score matrix is normalized by applying an activation function (such as the softmax function) to generate the weight matrices of multiple attention heads.
[0114] Third, according to the weight matrices of multiple attention heads, perform a weighted sum on the value vectors to obtain the output feature vectors of multiple attention heads.
[0115] Optionally, the terminal device uses the weight matrices of multiple attention heads to perform a weighted sum on the corresponding value vectors, generates the output of each attention head, and splices or further linearly transforms the outputs of all attention heads to form the output feature vector of the multi-head attention.
[0116] Through this embodiment, multiple parallel attention heads are introduced, and each head independently processes the query vector, the key vector, and the value vector. This enables the audio classification model to parallelly search for multiple dependencies in different subspaces, reflecting the complex associations of different parts in the speech signal and enhancing the feature capture ability of the audio classification model; using the dot product operation to calculate the attention scores reflects the correlation between elements in the sequence, and the weight matrix is obtained through normalization. Normalization ensures that the output of each attention head is based on relative importance, avoiding bias in feature processing and improving the fairness and accuracy of feature fusion; according to the generated weight matrix, a weighted sum is performed on the value vectors to obtain the output feature vectors of each attention head. This mechanism allows the audio classification model to dynamically adjust its contribution to features according to the relative importance of each element, significantly improving the quality of feature representation and providing a more accurate information basis for subsequent classification and recognition tasks.
[0117] In an exemplary embodiment, the above speech recognition method further includes:
[0118] 1. Obtain multiple training samples. The training samples among the multiple training samples include historical audio signals and label data, and the label data of the multiple training samples is used to indicate the language type to which the historical audio signals in the multiple training samples belong.
[0119] Among them, in order to enable the audio classification model to have the ability to identify different language types, in this embodiment, historical audio signals with label data are used as training samples to train a pre-trained audio classification model, and a trained audio classification model is obtained.
[0120] The historical audio signal refers to the audio data used to train the model. The historical audio signal can be selected from open-source data sets (such as open-source dialect data sets of different language types and open-source standard language data sets of different languages), or can be an audio segment that is manually recorded or collected. Figure 6 It is a training flow chart of an optional audio classification model according to an embodiment of the present application. As Figure 6 shown, in this embodiment, the historical audio signal can be selected from an open-source dialect data set and an open-source Mandarin data set. Training samples are extracted from these two data sets, and a pre-trained audio classification model is trained to obtain a trained audio classification model that can distinguish audio signals of two different language types, namely dialect and Mandarin. It can be understood that if more data sets of different language types are used to train the pre-trained audio classification model, the obtained trained audio classification model can distinguish audio signals of different language types.
[0121] The label data is annotation information that matches the historical audio signal, usually represented in text form, and is used to indicate the language type to which the audio signal belongs.
[0122] 2. Use multiple training samples to train the audio classification model to be trained, and obtain a trained audio classification model. Among them, during the training of the audio classification model, the model parameters of the audio classification model are adjusted according to the difference between the predicted language types corresponding to the multiple training samples output by the audio classification model and the language types indicated by the label data of the multiple training samples.
[0123] Among them, the predicted language type refers to the specific language or dialect category to which the historical audio signal identified by the audio classification model belongs.
[0124] Optionally, as Figure 6As shown in the figure, the terminal device constructs an audio classification model that combines an enhanced residual network, a transformer network, and a feature fusion network, initializes the model parameters, inputs each training sample into the audio classification model, extracts features using the enhanced residual network and the transformer network parts of the audio classification model, and fuses the features through the feature fusion network using the attention mechanism to obtain the predicted language type corresponding to each training sample; compares the predicted language type output by the audio classification model with the language type indicated by the label data of the corresponding training sample, calculates a loss function, such as cross-entropy loss, to quantify the difference between the prediction result and the actual label; based on the calculated loss, uses the gradient descent method for backpropagation to update the parameters of the audio classification model to minimize the difference between the predicted language type and the language type indicated by the label data; repeats the above steps, uses multiple training samples for multiple rounds of iterative training, and continuously adjusts the model parameters until the model loss converges to obtain a trained audio classification model.
[0125] Through this embodiment, in the training stage, labeled historical audio signals are introduced, and the model parameters of the audio classification model are fine-tuned using multiple training samples. During the training process, parameters such as the number of training rounds and the learning rate of the audio classification model are continuously adjusted, enabling the audio classification model to better learn the features and patterns of different language types, continuously optimizing the performance of the audio classification model, providing accurate language type information for the multi-language recognition model, effectively avoiding language confusion during the speech recognition process, and significantly improving the recognition accuracy.
[0126] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.
[0127] In another embodiment of the present application, as Figure 7 shown, a speech recognition device is further provided, including:
[0128] A speech acquisition module 702, configured to acquire continuous audio signals to be recognized.
[0129] A language type recognition module 704, configured to input the audio signals to be recognized into a pre-trained audio classification model to obtain the target language type output by the audio classification model; the audio classification model is used to extract local features and global features of the audio signals to be recognized, fuse the local features and global features, and determine the language type to which the audio signals to be recognized belong based on the fused speech features obtained by the fusion.
[0130] A speech transcription module 706 is configured to input an audio signal to be recognized and a target language type into a pre-trained multi-language recognition model, and obtain the text content output by the multi-language recognition model. The multi-language recognition model is configured to recognize the audio signal to be recognized according to the language recognition rules corresponding to the target language type, and obtain the text content corresponding to the audio signal to be recognized.
[0131] In an exemplary embodiment, the audio classification model includes an enhanced residual network, a transformer network, and a feature fusion network. The language type recognition module 704 is further configured to input the audio signal to be recognized into a pre-trained audio classification model, so that the audio classification model performs the following detection operations on the audio signal to be recognized: converting the audio signal to be recognized into a two-dimensional spectrogram; extracting features of the two-dimensional spectrogram through the enhanced residual network to obtain local features of the audio signal to be recognized; extracting features of the two-dimensional spectrogram through the transformer network to obtain global features of the audio signal to be recognized; fusing the local features and the global features through the feature fusion network to obtain fused speech features; predicting the classification probability of the fused speech features to obtain the probability prediction values of the audio signal to be recognized belonging to different language types; and determining the target language type to which the audio signal to be recognized belongs according to the probability prediction values of the audio signal to be recognized belonging to different language types.
[0132] In an exemplary embodiment, the language type recognition module 704 is further configured to segment continuous audio signals to be recognized into a plurality of time frames. The time frames in the plurality of time frames include the time domain information of the audio signal to be recognized. Performing a Fourier transform on the plurality of time frames to obtain a frequency domain signal corresponding to the plurality of time frames, and generating an initial spectrum based on the frequency domain signal corresponding to the plurality of time frames; filtering the initial spectrum through a set of Mel filters to obtain a Mel spectrum, and performing a logarithm operation on the Mel spectrum to obtain a two-dimensional spectrogram.
[0133] In an exemplary embodiment, the enhanced residual network includes a plurality of feature extraction layers and a plurality of feature fusion layers. The plurality of feature fusion layers correspond to the non-first feature extraction layers in the plurality of feature extraction layers in sequence. The language type recognition module 704 is further configured to extract features from the two-dimensional spectrogram through the first feature extraction layer in the plurality of feature extraction layers to obtain the output of the first feature extraction layer; and use the non-first feature extraction layers in the plurality of feature extraction layers as the current feature extraction layer in sequence to perform the following feature extraction operations: extract features from the output of the previous feature extraction layer of the current feature extraction layer through the current feature extraction layer to obtain a plurality of feature groups, and perform weighted summation on the plurality of feature groups according to the first attention weights corresponding to the plurality of feature groups to obtain the output of the current feature extraction layer; fuse the output of the first feature extraction layer and the output of the next feature extraction layer of the first feature extraction layer through the first feature fusion layer in the plurality of feature fusion layers to obtain the output of the first feature fusion layer; use the non-first feature fusion layers in the plurality of feature fusion layers as the current feature fusion layer in sequence to perform the following feature fusion operations, and the feature extraction layer corresponding to the current feature fusion layer is the current feature extraction layer: fuse the output of the current feature extraction layer and the output of the previous feature fusion layer of the current feature fusion layer through the current feature fusion layer to obtain the output of the current feature fusion layer; perform pooling and fully connected processing on the output of the last feature fusion layer in the plurality of feature fusion layers to obtain the local features of the audio signal to be recognized.
[0134] In an exemplary embodiment, the transformer network includes a plurality of encoding blocks. Each encoding block in the plurality of encoding blocks includes a multi-head self-attention network and a feed-forward neural network. The language type recognition module 704 is further configured to convert the two-dimensional spectrogram into a one-dimensional vector sequence through the multi-head self-attention network of the first encoding block in the plurality of encoding blocks, perform encoding operations on the one-dimensional vector sequence to obtain the output features of the first encoding block; and use the non-first encoding blocks in the plurality of encoding blocks as the current encoding block in sequence to perform the following encoding operations to obtain the output features of the plurality of encoding blocks. The multi-head self-attention network of the current encoding block is the current multi-head self-attention network, and the feed-forward neural network in the current encoding block is the current feed-forward neural network: calculate the second attention weights between any two elements in the output features of the previous encoding block of the current encoding block through a plurality of parallel attention heads, calculate the output feature vectors of the plurality of attention heads according to the second attention weights between any two elements, splice the output feature vectors of the plurality of attention heads to obtain the multi-head attention output; perform a linear transformation on the multi-head attention output through the current feed-forward neural network to obtain the output features of the current encoding block; the output features of the last encoding block are the global features of the audio signal to be recognized.
[0135] In an exemplary embodiment, the language type recognition module 704 is further configured to perform a linear transformation on the output features of the previous encoded block of the current encoded block through the current multi-head self-attention network to obtain a query vector, a key vector, and a value vector, split the query vector, the key vector, and the value vector into multiple attention heads respectively, and calculate the second attention weights between any two elements through the multiple attention heads; calculate the attention scores corresponding to the second attention weights between any two elements through dot product, and perform normalization processing on the attention scores corresponding to the multiple attention heads to obtain the weight matrices of the multiple attention heads; and perform weighted summation on the value vector according to the weight matrices of the multiple attention heads to obtain the output feature vectors of the multiple attention heads.
[0136] In an exemplary embodiment, the language type recognition module 704 is further configured to obtain a plurality of training samples, where the training samples in the plurality of training samples include historical audio signals and label data, and the label data of the plurality of training samples is used to indicate the language type to which the historical audio signals in the plurality of training samples belong; use the plurality of training samples to train the audio classification model to be trained to obtain a trained audio classification model, where, in the process of training the audio classification model, the model parameters of the audio classification model are adjusted according to the difference between the predicted language types corresponding to the plurality of training samples output by the audio classification model and the language types indicated by the label data of the plurality of training samples.
[0137] For the descriptions of the features in the corresponding embodiments of the speech recognition device, reference can be made to the relevant descriptions in the corresponding embodiments of the speech recognition method, which will not be elaborated here one by one.
[0138] An embodiment of the present application further provides an electronic device, including a memory and a processor, where a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned embodiments of the speech recognition method.
[0139] An embodiment of the present application further provides a computer-readable storage medium, where a computer program is stored in the computer-readable storage medium, and the computer program is configured to execute the steps in any of the above-mentioned embodiments of the speech recognition method when running.
[0140] In an exemplary embodiment, the above-mentioned computer-readable storage medium may include, but is not limited to: USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), mobile hard disks, magnetic disks, or optical discs and other various media that can store computer programs.
[0141] Embodiments of the present application also provide a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-mentioned embodiments of the voice recognition method.
[0142] Embodiments of the present application also provide another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-mentioned embodiments of the voice recognition method.
[0143] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0144] The above has introduced in detail a voice recognition method, device, storage medium, and program product provided by the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A speech recognition method, characterized in that, Including: Obtain consecutive audio signals to be recognized; Input the audio signal to be recognized into a pre-trained audio classification model to obtain the target language type output by the audio classification model; the audio classification model is used to extract local features and global features of the audio signal to be recognized, fuse the local features and the global features, and determine the language type to which the audio signal to be recognized belongs based on the fused speech features; Input the audio signal to be recognized and the target language type into a pre-trained multi-language recognition model to obtain the text content output by the multi-language recognition model; the multi-language recognition model is used to recognize the audio signal to be recognized according to the language recognition rules corresponding to the target language type to obtain the text content corresponding to the audio signal to be recognized.
2. The voice recognition method according to claim 1, wherein The audio classification model includes an enhanced residual network, a transformer network, and a feature fusion network; The step of inputting the audio signal to be recognized into a pre-trained audio classification model to obtain the target language type output by the audio classification model includes: Input the audio signal to be recognized into the pre-trained audio classification model so that the audio classification model performs the following detection operations on the audio signal to be recognized: Convert the audio signal to be recognized into a two-dimensional spectrogram; Extract features from the two-dimensional spectrogram through the enhanced residual network to obtain the local features of the audio signal to be recognized; Extract features from the two-dimensional spectrogram through the transformer network to obtain the global features of the audio signal to be recognized; Fuse the local features and the global features through the feature fusion network to obtain the fused speech features; Perform classification probability prediction on the fused speech features to obtain probability prediction values of the audio signal to be recognized belonging to different language types; Determine the target language type to which the audio signal to be recognized belongs according to the probability prediction values of the audio signal to be recognized belonging to different language types.
3. The speech recognition method according to claim 2, wherein The step of converting the audio signal to be recognized into a two-dimensional spectrogram includes: Segment the consecutive audio signals to be recognized into multiple time frames; the time frames in the multiple time frames contain the time domain information of the audio signal to be recognized; Perform Fourier transform on the multiple time frames to obtain frequency domain signals corresponding to the multiple time frames, and generate an initial spectrogram based on the frequency domain signals corresponding to the multiple time frames; Filter the initial spectrogram through a set of Mel filters to obtain a Mel spectrogram, and perform logarithmic processing on the Mel spectrogram to obtain the two-dimensional spectrogram.
4. The voice recognition method according to claim 2, wherein The enhanced residual network includes multiple feature extraction layers and multiple feature fusion layers, and the multiple feature fusion layers correspond to the non-first feature extraction layers in the multiple feature extraction layers in sequence; The step of extracting features from the two-dimensional spectrogram through the enhanced residual network to obtain the local features of the audio signal to be recognized includes: Extract features from the two-dimensional spectrogram through the first feature extraction layer in the multiple feature extraction layers to obtain the output of the first feature extraction layer; Take the non-first feature extraction layers among the multiple feature extraction layers as the current feature extraction layer in sequence and perform the following feature extraction operations: Use the current feature extraction layer to perform feature extraction on the output of the previous feature extraction layer of the current feature extraction layer to obtain multiple feature groups, and perform weighted summation on the multiple feature groups according to the first attention weights corresponding to the multiple feature groups to obtain the output of the current feature extraction layer; Use the first feature fusion layer among the multiple feature fusion layers to perform feature fusion on the output of the first feature extraction layer and the output of the next feature extraction layer of the first feature extraction layer to obtain the output of the first feature fusion layer; Take the non-first feature fusion layers among the multiple feature fusion layers as the current feature fusion layer in sequence and perform the following feature fusion operations. The feature extraction layer corresponding to the current feature fusion layer is the current feature extraction layer: Use the current feature fusion layer to fuse the output of the current feature extraction layer and the output of the previous feature fusion layer of the current feature fusion layer to obtain the output of the current feature fusion layer; Perform pooling and fully connected processing on the output of the last feature fusion layer among the multiple feature fusion layers to obtain the local feature of the audio signal to be recognized.
5. The voice recognition method according to claim 2, wherein The transformer network includes multiple encoding blocks, and the encoding blocks among the multiple encoding blocks include a multi-head self-attention network and a feed-forward neural network; The feature extraction of the two-dimensional spectrogram by the transformer network to obtain the global feature of the audio signal to be recognized includes: Use the multi-head self-attention network of the first encoding block among the multiple encoding blocks to convert the two-dimensional spectrogram into a one-dimensional vector sequence, perform encoding operations on the one-dimensional vector sequence to obtain the output feature of the first encoding block; Take the non-first encoding blocks among the multiple encoding blocks as the current encoding block in sequence and perform the following encoding operations to obtain the output features of the multiple encoding blocks. The multi-head self-attention network of the current encoding block is the current multi-head self-attention network, and the feed-forward neural network in the current encoding block is the current feed-forward neural network: Calculate the second attention weights between any two elements in the output feature of the previous encoding block of the current encoding block through multiple parallel attention heads, calculate the output feature vectors of the multiple attention heads according to the second attention weights between any two elements, splice the output feature vectors of the multiple attention heads to obtain the multi-head attention output; Use the current feed-forward neural network to perform a linear transformation on the multi-head attention output to obtain the output feature of the current encoding block; The output feature of the last encoding block is the global feature of the audio signal to be recognized.
6. The speech recognition method according to claim 5, characterized in that The calculation of the second attention weights between any two elements in the output feature of the previous encoding block of the current encoding block through multiple parallel attention heads and the calculation of the output feature vectors of the multiple attention heads according to the second attention weights between any two elements include: Performing a linear transformation on the output features of the previous encoded block of the current encoded block through the current multi-head self-attention network to obtain a query vector, a key vector, and a value vector, splitting the query vector, the key vector, and the value vector into the multiple attention heads respectively, and calculating the second attention weights between any two elements through the multiple attention heads; Calculating the attention scores corresponding to the second attention weights between any two elements through dot product, and performing normalization processing on the attention scores corresponding to the multiple attention heads to obtain the weight matrices of the multiple attention heads; According to the weight matrices of the multiple attention heads, performing weighted summation on the value vectors to obtain the output feature vectors of the multiple attention heads.
7. The voice recognition method according to any one of claims 1 to 6, characterized in that, The method further includes: Obtaining a plurality of training samples, where the training samples in the plurality of training samples include historical audio signals and label data, and the label data of the plurality of training samples is used to indicate the language type to which the historical audio signals in the plurality of training samples belong; Using the plurality of training samples to train the audio classification model to be trained to obtain the trained audio classification model, where during the training of the audio classification model, the model parameters of the audio classification model are adjusted according to the difference between the predicted language types corresponding to the plurality of training samples output by the audio classification model and the language types indicated by the label data of the plurality of training samples.
8. A voice recognition device, characterized in that, Including: A voice acquisition module for acquiring continuous audio signals to be recognized; A language type recognition module for inputting the audio signal to be recognized into a pre-trained audio classification model to obtain the target language type output by the audio classification model; the audio classification model is used to extract local features and global features of the audio signal to be recognized, fuse the local features and the global features, and determine the language type to which the audio signal to be recognized belongs based on the fused speech features; A voice transcription module for inputting the audio signal to be recognized and the target language type into a pre-trained multilingual recognition model to obtain the text content output by the multilingual recognition model; the multilingual recognition model is used to recognize the audio signal to be recognized according to the language recognition rules corresponding to the target language type to obtain the text content corresponding to the audio signal to be recognized.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, where the computer program, when executed by a processor, implements the steps of the voice recognition method according to any one of claims 1 to 7.
10. A computer program product, comprising a computer program, characterized in that, The computer program, when executed by a processor, implements the steps of the voice recognition method according to any one of claims 1 to 7.