Speech recognition method, speech recognition device, electronic equipment and storage medium

By using a pre-trained speech recognition model to perform phoneme feature extraction and feature fusion for multilingual speech speech recognition, the problem of low accuracy of multilingual speech recognition in the prior art is solved, and higher quality speech-to-text conversion is achieved.

CN120236584APending Publication Date: 2025-07-01SF TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311865603.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-30
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The prior art relies on pronunciation dictionary in multilingual speech recognition, resulting in a long time-consuming construction and maintenance and low recognition accuracy, because coarse-grained dictionaries can only roughly recognize the text content of speech data.

Method used

A speech recognition method is proposed, using the phoneme feature extraction network of the pre-trained speech recognition model to perform type recognition and phoneme feature extraction for multilingual speech fragments, combined with the speech feature extraction network, the speech context network and the linear layer, feature extraction, context extraction and feature fusion are performed, and finally more accurate speech-to-text conversion is achieved.

Benefits of technology

By extracting high-quality speech phoneme features and fusing speech context information, the accuracy of multilingual speech recognition is significantly improved, and the problem of low recognition accuracy in traditional methods is overcome.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236584A_ABST
    Figure CN120236584A_ABST
Patent Text Reader

Abstract

The invention provides a voice recognition method, a voice recognition device, electronic equipment and a storage medium, and belongs to the field of financial science and technology. The method comprises the steps of obtaining target voice data; a phoneme feature extraction network based on a speech recognition model performs type recognition on the speech segment to obtain a target language type of the speech segment, speech phoneme features are extracted based on the target language type and the phoneme feature extraction network, and the speech recognition model comprises a speech feature extraction network, a speech context network and a linear layer; extracting voice characterization features of the voice segments based on the voice feature quantization network; processing the voice representation feature into a voice context feature based on a voice context network, and fusing the voice context feature and the voice phoneme feature into a voice fusion feature; and performing recognition processing on the voice fusion feature based on the linear layer to obtain a text fragment corresponding to the voice fragment, and determining a target text of the target voice data based on the text fragment. According to the invention, the multi-language speech recognition accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of fintech, and in particular, to a speech recognition method, a speech recognition device, an electronic device, and a storage medium. Background Art

[0002] With the rapid development of artificial intelligence technology, intelligent voice customer service is widely used in various business fields. For example, in financial scenarios such as financial business handling and logistics business consultation, multilingual speech recognition technology is often used to recognize the text content involved in various language types of speech.

[0003] However, currently, the recognition of multilingual speech often relies on pronunciation dictionaries. This method often requires building pronunciation dictionaries at the text level corresponding to various language types of speech simultaneously, which has the problem of long time consumption. In addition, for speech data of different language types, the constructed pronunciation dictionaries are often coarse-grained and can only roughly recognize the text content indicated by the speech data, resulting in low speech recognition accuracy. Summary of the Invention

[0004] The main purpose of the embodiments of this application is to propose a speech recognition method, a speech recognition device, an electronic device, and a storage medium, aiming to improve the recognition accuracy of multilingual speech.

[0005] To achieve the above object, a first aspect of the embodiments of this application proposes a speech recognition method, and the method includes:

[0006] Obtain target speech data, where the target speech data includes speech segments of at least two language types;

[0007] Perform type recognition on the speech segments based on the phoneme feature extraction network of a pre-trained speech recognition model to obtain the target language type of the speech segments, and extract the corresponding speech phoneme features of the speech segments based on the target language type and the phoneme feature extraction network, where the speech recognition model includes a speech feature extraction network, a speech context network, and a linear layer;

[0008] Extract features from the speech segments based on the speech feature extraction network to obtain speech representation features;

[0009] Extract context from the speech representation features based on the speech context network to obtain speech context features, and perform feature fusion on the speech context features and the speech phoneme features based on the speech context network to obtain speech fusion features;

[0010] Performing recognition processing on the speech fusion feature based on the linear layer to obtain a text segment corresponding to the speech segment, and determining a target text corresponding to the target speech data based on the text segment.

[0011] In some embodiments, the phoneme feature extraction network includes a language classification sub-network and a phoneme feature extractor corresponding to each candidate language type;

[0012] Performing type recognition on the speech segment by the phoneme feature extraction network based on a pre-trained speech recognition model to obtain a target language type of the speech segment, and extracting a speech phoneme feature corresponding to the speech segment based on the target language type and the phoneme feature extraction network, including:

[0013] Performing language type recognition on the speech segment based on the language classification sub-network and multiple candidate language types to obtain the target language type of the speech segment;

[0014] Determining a phoneme feature extractor corresponding to the target language type from among the phoneme feature extractors corresponding to multiple candidate language types;

[0015] Performing phoneme feature extraction on the speech segment based on the phoneme feature extractor corresponding to the target language type to obtain the speech phoneme feature.

[0016] In some embodiments, the speech context network includes a first feature processing sub-network and a second feature processing sub-network;

[0017] Performing context extraction on the speech representation feature based on the speech context network to obtain a speech context feature, and performing feature fusion on the speech context feature and the speech phoneme feature based on the speech context network to obtain a speech fusion feature, including:

[0018] Performing context feature extraction on the speech representation feature based on the first feature processing sub-network to obtain a speech context feature;

[0019] Performing feature fusion on the speech context feature and the speech phoneme feature to obtain a preliminary fusion feature;

[0020] Performing context feature extraction on the preliminary fusion feature based on the second feature processing sub-network to obtain the speech fusion feature.

[0021] To achieve the above object, a second aspect of the embodiments of the present application proposes a training method for a speech recognition model, the training method including:

[0022] Obtain sample speech data, where the sample speech data includes multiple sample speech segments, the sample texts corresponding to each of the sample speech segments, and the candidate language types;

[0023] Input the sample speech segments, the sample texts, and the candidate language types into a speech recognition model, where the speech recognition model includes a phoneme feature extraction network, a speech feature extraction network, a speech context network, and a linear layer; the phoneme feature extraction network includes a language classification sub-network and phoneme feature extractors corresponding to multiple candidate language types;

[0024] Based on the language classification sub-network, perform type recognition on the sample speech segments to obtain the sample language type of the sample speech segments, and based on the phoneme feature extractor corresponding to the sample language type, perform phoneme feature extraction on the sample speech segments to obtain the sample phoneme features of the sample speech segments;

[0025] Based on the network structure of the phoneme feature extractor corresponding to the sample language type, divide the speech context network into a first sub-network and a second sub-network;

[0026] Based on the speech feature extraction network, perform feature extraction on the sample speech segments to obtain the sample speech representation features of the sample speech segments;

[0027] Based on the first sub-network, perform context feature extraction on the sample speech representation features to obtain first sample context features;

[0028] Perform feature fusion on the first sample context features and the sample phoneme features to obtain sample fusion speech features;

[0029] Based on the second sub-network, perform context feature extraction on the sample fusion speech features to obtain second sample context features;

[0030] Based on the linear layer and the second sample context features, perform speech recognition to obtain the predicted text of the sample speech segments, and based on the comparison between the predicted text and the sample texts, adjust the model parameters of the speech recognition model to train the speech recognition model.

[0031] In some embodiments, the language classification sub-network is trained in the following manner:

[0032] Based on a preset first pre-trained network, perform language type prediction on the sample speech segments to obtain the predicted language type of the sample speech segments;

[0033] Based on the comparison between the predicted language type and the candidate language types, determine a first loss function;

[0034] Based on the first loss function, adjust the parameters of the first pre-trained network to train the first pre-trained network, and obtain the language classification sub-network.

[0035] In some embodiments, the phoneme feature extractors corresponding to the multiple candidate language types are trained as follows:

[0036] For each candidate language type, input the sample speech segment of the candidate language type into a preset second pre-trained network to obtain the phoneme feature sensitivities of each layer of the second pre-trained network for the sample speech segment;

[0037] Based on the phoneme feature sensitivities, adjust the structure of the second pre-trained network to obtain the phoneme feature extractor of the candidate language type.

[0038] In some embodiments, the speech feature extraction network and the speech context network are jointly trained as follows:

[0039] Based on the speech feature extraction network, perform convolutional processing on multiple sample speech segments to obtain the sample speech features of each sample speech segment;

[0040] Select target speech features from multiple sample speech features and mask the target speech features;

[0041] Input the non-target speech features among the multiple sample speech features and the masked target speech features into the speech context network to obtain a mask prediction result;

[0042] Based on the comparison between the target speech features and the mask prediction result, determine a second loss function;

[0043] Based on the second loss function, update the parameters of the speech feature extraction network, and based on the second loss function, update the parameters of the speech context network.

[0044] To achieve the above object, a third aspect of the embodiments of the present application proposes a speech recognition device, and the device includes:

[0045] An acquisition module, configured to acquire target speech data, where the target speech data includes speech segments of at least two language types;

[0046] A phoneme feature extraction module, configured to perform type recognition on the speech segment based on a phoneme feature extraction network of a pre-trained speech recognition model to obtain the target language type of the speech segment, and extract the corresponding speech phoneme features of the speech segment based on the target language type and the phoneme feature extraction network, where the speech recognition model includes a speech feature extraction network, a speech context network, and a linear layer;

[0047] A feature extraction module, configured to perform feature extraction on the speech segment based on the speech feature extraction network to obtain speech representation features;

[0048] A feature fusion module, configured to perform context extraction on the speech extraction features based on the speech context network to obtain speech context features, and perform feature fusion on the speech context features and the speech phoneme features based on the speech context network to obtain speech fusion features;

[0049] An identification module, configured to perform identification processing on the speech fusion features based on the linear layer to obtain a text segment corresponding to the speech segment, and determine a target text corresponding to the target speech data based on the text segment.

[0050] To achieve the above object, a fourth aspect of the embodiments of the present application provides an electronic device, where the electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the method described in the first aspect or the method described in the second aspect is implemented.

[0051] To achieve the above object, a fifth aspect of the embodiments of the present application provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described in the first aspect or the method described in the second aspect is implemented.

[0052] The speech recognition method, training method of speech recognition model, speech recognition device, electronic device and storage medium proposed in this application obtain target speech data, perform type recognition on speech segments based on the phoneme feature extraction network of the speech recognition model to obtain the target language type of the speech segments, and extract the speech phoneme features of the speech segments based on the target language type and the phoneme feature extraction network. It can extract the speech phoneme level information of language segments for different language types, improve the feature quality of the extracted speech phoneme features, and is beneficial to enriching the speech level information of the speech to be recognized. Further, the speech recognition model includes a speech feature extraction network, a speech context network, and a linear layer; feature extraction is performed on speech segments based on the speech feature extraction network to obtain speech representation features, improving the feature processing efficiency. Then, context extraction is performed on the speech extraction features based on the speech context network to obtain speech context features, and feature fusion is performed on the speech context features and speech phoneme features based on the speech context network to obtain speech fusion features, which is beneficial to improving the richness and comprehensiveness of the feature content of the speech fusion features. Finally, recognition processing is performed on the speech fusion features based on the linear layer to obtain the text segment corresponding to the speech segment, and the target text of the target speech data is determined based on the text segment, enabling the target text to more comprehensively reflect the semantic content of the target speech data, thereby improving the recognition accuracy of multilingual speech. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 is a flowchart of the speech recognition method provided by an embodiment of this application;

[0054] Figure 2 is Figure 1 a flowchart of step S102 in

[0055] Figure 3 is Figure 1 a flowchart of step S104 in

[0056] Figure 4 is a schematic diagram of the implementation process of the speech recognition method provided by an embodiment of this application;

[0057] Figure 5 is a flowchart of the training method of the speech recognition model provided by an embodiment of this application;

[0058] Figure 6 is a schematic diagram of the implementation process of the training method of the speech recognition model provided by an embodiment of this application;

[0059] Figure 7 is a flowchart of pre-training the language classification sub-network in the training method of the speech recognition model provided by an embodiment of this application;

[0060] Figure 8It is a flowchart for pre-training a phoneme feature extractor corresponding to a candidate language type in the method for training a speech recognition model provided by an embodiment of the present application;

[0061] Figure 9 It is a schematic diagram of the implementation process for pre-training a phoneme feature extractor corresponding to a candidate language type in the method for training a speech recognition model provided by an embodiment of the present application;

[0062] Figure 10 It is a flowchart for pre-training a speech feature extraction network and a speech context network in the method for training a speech recognition model provided by an embodiment of the present application;

[0063] Figure 11 It is a schematic diagram of the structure of a speech recognition device provided by an embodiment of the present application;

[0064] Figure 12 It is a schematic diagram of the structure of a training device for a speech recognition model provided by an embodiment of the present application;

[0065] Figure 13 It is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0066] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0067] It should be noted that although functional module division is performed in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order from the module division in the device or the flowchart. Terms such as "first" and "second" in the specification, claims and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.

[0068] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0069] First, several nouns involved in the present application are analyzed:

[0070] Artificial Intelligence (AI): It is a new technical science that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence. Artificial intelligence is a branch of computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. The research in this field includes robots, speech recognition, image recognition, natural language processing, and expert systems, etc. Artificial intelligence can simulate the information process of human consciousness and thinking. It also uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results in terms of theories, methods, technologies, and application systems.

[0071] Natural Language Processing (NLP): NLP uses computers to process, understand, and apply human languages (such as Chinese, English, etc.). NLP belongs to a branch of artificial intelligence and is an interdisciplinary field of computer science and linguistics, and is often also called computational linguistics. Natural language processing includes syntactic analysis, semantic analysis, discourse understanding, etc. Natural language processing is commonly used in technical fields such as machine translation, handwritten and printed character recognition, speech recognition and text-to-speech conversion, information intention recognition, information extraction and filtering, text classification and clustering, public opinion analysis, and opinion mining. It involves data mining related to language processing, machine learning, knowledge acquisition, knowledge engineering, artificial intelligence research, and linguistic research related to language computing, etc.

[0072] Information Extraction (NER): It is a text processing technology that extracts factual information such as entities, relationships, events of a specified type from natural language texts and forms structured data for output. Information extraction is a technology for extracting specific information from text data. Text data is composed of some specific units, such as sentences, paragraphs, and texts. Text information is exactly composed of some small specific units, such as characters, words, phrases, sentences, paragraphs, or combinations of these specific units. Extracting noun phrases, personal names, place names, etc. from text data are all text information extraction. Of course, the information extracted by text information extraction technology can be various types of information.

[0073] With the rapid development of artificial intelligence technology, intelligent voice customer service has been widely used in various business fields. For example, in financial scenarios such as financial business handling and logistics business consulting, multilingual speech recognition technology is often used to identify the text content involved in various language types of voices.

[0074] However, currently, the recognition of multi-language speech often relies on pronunciation dictionaries. This method often requires constructing pronunciation dictionaries at the text level corresponding to voices of multiple language types simultaneously, which has the problem of long time consumption. In addition, for voice data of different language types, the constructed pronunciation dictionaries are often coarse-grained and can only roughly recognize the text content indicated by the voice data, resulting in low accuracy of speech recognition.

[0075] Based on this, the embodiments of the present application provide a speech recognition method, a training method for a speech recognition model, a speech recognition device, an electronic device, and a storage medium, aiming to improve the recognition accuracy of multi-language speech.

[0076] The speech recognition method, the training method for the speech recognition model, the speech recognition device, the electronic device, and the storage medium provided by the embodiments of the present application are specifically described through the following embodiments. First, the speech recognition method in the embodiments of the present application is described.

[0077] The embodiments of the present application can obtain and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results in theory, method, technology, and application systems.

[0078] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0079] The speech recognition method provided by the embodiments of the present application relates to the field of fintech. The speech recognition method provided by the embodiments of the present application can be applied to a terminal, or to a server side, or can also be software running on a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server side can be configured as an independent physical server, or can be configured as a server cluster or a distributed system composed of multiple physical servers, or can also be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the speech recognition method, etc., but is not limited to the above forms.

[0080] This application can be used in numerous general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0081] It should be noted that in each specific embodiment of this application, when it comes to performing relevant processing based on data related to the identity or characteristics of an object, such as object information, object behavior data, object historical data, and object location information, the permission or consent of the object will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when an embodiment of this application needs to obtain personal information of an object, it will obtain the separate permission or separate consent of the object through methods such as pop-up windows or jumping to a confirmation page. After clearly obtaining the separate permission or separate consent of the object, it will then obtain the necessary object-related data for the normal operation of the embodiment of this application.

[0082] Figure 1 It is an optional flowchart of the speech recognition method provided by an embodiment of this application. Figure 1 The method in it can include but is not limited to steps S101 to S105.

[0083] Step S101, obtain target speech data;

[0084] Step S102, perform type recognition on the speech segment based on the phoneme feature extraction network of the pre-trained speech recognition model to obtain the target language type of the speech segment, and extract the corresponding speech phoneme features of the speech segment based on the target language type and the phoneme feature extraction network;

[0085] Step S103, perform feature extraction on the speech segment based on the speech feature extraction network to obtain speech extraction features;

[0086] Step S104: Extract context from the speech extraction features based on the speech context network to obtain speech context features, and perform feature fusion on the speech context features and speech phoneme features based on the speech context network to obtain speech fusion features;

[0087] Step S105: Perform recognition processing on the speech fusion features based on the linear layer to obtain the text segment corresponding to the speech segment, and determine the target text corresponding to the target speech data based on the text segment.

[0088] Steps S101 to S105 shown in the embodiments of the present application, by obtaining the target speech data, performing type recognition on the speech segment through the phoneme feature extraction network of the speech recognition model to obtain the target language type of the speech segment, and extracting the speech phoneme features of the speech segment based on the target language type and the phoneme feature extraction network, can extract the speech phoneme level information of the language segment for different language types, improve the feature quality of the extracted speech phoneme features, and is beneficial to enriching the speech level information of the speech to be recognized. Further, the speech recognition model includes a speech feature extraction network, a speech context network, and a linear layer; feature extraction is performed on the speech segment based on the speech feature extraction network to obtain speech representation features, improving the feature processing efficiency. Then, context extraction is performed on the speech extraction features based on the speech context network to obtain speech context features, and feature fusion is performed on the speech context features and speech phoneme features based on the speech context network to obtain speech fusion features, which is beneficial to improving the richness and comprehensiveness of the feature content of the speech fusion features. Finally, recognition processing is performed on the speech fusion features based on the linear layer to obtain the text segment corresponding to the speech segment, and the target text corresponding to the target speech data is determined based on the text segment, so that the target text can more comprehensively reflect the semantic content of the target speech data, thereby improving the recognition accuracy of multilingual speech.

[0089] In step S101 of some embodiments, the target speech data is obtained.

[0090] The target speech data refers to an audio with a certain duration. For example, the target speech data is a conversation with a fixed duration, a speech audio with a fixed duration, or music with a fixed duration.

[0091] The target speech data includes speech segments of at least two language types. Among them, the language types include but are not limited to Chinese, English, Japanese, Korean, French, and so on. The speech segment is a part of the target speech data.

[0092] When specifically implementing this embodiment, the methods for obtaining the target speech data include but are not limited to the following methods:

[0093] (1) Purposefully crawl data from a preset data source through a web crawler to obtain target voice data. Among them, the preset data source includes a preset database or other network platforms that can provide voice materials, etc.

[0094] (2) Obtain target voice data from a public dataset. The public dataset can be the LJSpeech dataset, which contains various types of voice data recorded by speakers.

[0095] For example, in the field of logistics business, the target voice data can be the dialogue audio between the object handling the shipping business and the relevant service personnel.

[0096] After obtaining the target voice data, the target voice data can be divided into multiple voice segments according to the pause interval time between each voice segment in the target voice data. For example, set the pause interval threshold to 2 seconds. When the pause interval time between consecutive voice segments exceeds 2 seconds, the consecutive voice segments are split into two separate voice segments. When the pause interval time between consecutive voice segments does not exceed 2 seconds, the consecutive voice segments are taken as a whole to obtain a voice segment.

[0097] In step S102 of some embodiments, the type of the voice segment is recognized based on the phoneme feature extraction network of the pre-trained speech recognition model to obtain the target language type of the voice segment, and the voice phoneme features corresponding to the voice segment are extracted based on the target language type and the phoneme feature extraction network.

[0098] The speech recognition model is a neural network model with the function of recognizing the semantic content of voice data. The input of the speech recognition model is generally voice data, and the output of the speech recognition model is generally text data, which is used to indicate the semantic content of the input voice data.

[0099] In the embodiments of the present application, the speech recognition model includes a phoneme feature extraction network, a speech feature extraction network, a speech context network, and a linear layer. The speech feature extraction network is used to extract features from the input voice data to capture the speech content information in the voice data. The phoneme feature extraction network is used to extract phoneme features for voice data of a fixed language type to obtain the phoneme-level feature information of this language type in the voice data. The speech context network is used to process the speech features output by the speech feature extraction network, and perform feature fusion and speech context extraction based on the processed speech features and the phoneme-level feature information output by the phoneme feature extraction network to obtain more comprehensive speech features. The linear layer is mainly used to recognize the semantic content of the voice data according to the more comprehensive speech features output by the speech context network, so as to output a text data that can represent the main content of the voice data.

[0100] The target language type of a speech segment is used to indicate the language type to which the speech segment belongs.

[0101] Speech phoneme features are used to characterize the phoneme-level feature information unique to a speech segment under the target language type.

[0102] To save space, the training process of the speech recognition model in the embodiments of this application will be described in detail below and will not be elaborated here.

[0103] Please refer to Figure 2 , in some embodiments, the phoneme feature extraction network includes a language classification sub-network and a phoneme feature extractor corresponding to each candidate language type; step S102 may include but is not limited to steps S201 to S203:

[0104] Step S201, based on the language classification sub-network and multiple candidate language types, perform language type recognition on the speech segment to obtain the target language type of the speech segment;

[0105] Step S202, among the phoneme feature extractors corresponding to multiple candidate language types, determine the phoneme feature extractor corresponding to the target language type;

[0106] Step S203, based on the phoneme feature extractor corresponding to the target language type, perform phoneme feature extraction on the speech segment to obtain speech phoneme features.

[0107] The following will describe steps S201 to S203 in detail.

[0108] In step S201 of some embodiments, first input the speech segment into the language classification sub-network, where the language classification sub-network is a neural network structure with language classification capabilities. Then, the language classification sub-network extracts features from the speech segment to obtain the speech feature information of the speech segment; further, the softmax function in the language classification sub-network performs multi-classification processing on the speech segment according to the speech feature information to determine the probability distribution of the speech segment on each candidate language type, and obtains the probability vector of the speech segment on each candidate language type. Finally, according to the magnitude of the probability vector, the candidate language type with the largest probability vector is determined as the target language type of the speech segment.

[0109] In step S202 of some embodiments, search for the candidate language type that is the same as the target language type among multiple candidate language types, and determine the phoneme feature extractor corresponding to the candidate language type that is the same as the target language type as the phoneme feature extractor corresponding to the target language type.

[0110] In step S203 of some embodiments, first, the speech segment is input into the phoneme feature extractor corresponding to the target language type. Then, the speech segment is subjected to feature extraction through each hierarchical structure of the phoneme feature extractor. The output of the previous hierarchical structure is used as the input of the next hierarchical structure, and the speech phoneme information in the speech segment is extracted layer by layer. The output of the last hierarchical structure of the phoneme feature extractor is used as the speech phoneme feature of the speech segment.

[0111] The advantage of this embodiment is that a corresponding phoneme feature extractor is set for each different candidate language type. First, the language classification sub-network is used to determine the target language type of the input speech segment. According to the target language type, the phoneme feature extractor corresponding to the target language type is selected to perform phoneme feature extraction on the speech segment, which can extract the speech phoneme hierarchical information of the language segment for different language types, improve the feature quality of the extracted speech phoneme features, facilitate enriching the speech hierarchical information of the speech to be recognized, and improve the speech recognition effect.

[0112] In step S103 of some embodiments, the speech segment is subjected to feature extraction based on the speech feature extraction network to obtain the speech representation feature.

[0113] Feature extraction refers to the process of capturing the speech content information of the speech segment.

[0114] In the specific implementation of this embodiment, the speech feature extraction network includes a plurality of sequentially connected convolutional layers. Specifically, first, the speech segment is input into the speech feature extraction network. Then, the continuous speech features in the speech segment are subjected to feature extraction through a plurality of convolutional layers in sequence, and the speech feature output by the last convolutional layer is used as the speech representation feature.

[0115] In step S104 of some embodiments, context extraction is performed on the speech representation feature based on the speech context network to obtain the speech context feature, and feature fusion is performed on the speech context feature and the speech phoneme feature based on the speech context network to obtain the speech fusion feature.

[0116] The speech fusion feature is used to indicate the overall speech feature information of the speech segment and the phoneme hierarchical feature information of the speech segment.

[0117] In the embodiment of the present application, the speech context network is a neural network based on a transformer. The speech context network includes a first feature processing sub-network and a second feature processing sub-network. The first feature processing sub-network and the second feature processing sub-network of the speech context network are formed by dividing multiple hierarchical structures of the speech context network according to the hierarchical structure of the phoneme feature extractor. Among them, the hierarchical structure of the first feature processing sub-network is basically the same as the hierarchical structure of the phoneme feature extractor.

[0118] Please refer to Figure 3 , in some embodiments, step S104 may include but is not limited to steps S301 to S303:

[0119] Step S301, extracting context features from the speech representation features based on the first feature processing sub-network to obtain speech context features;

[0120] Step S302, fusing the speech context features and the speech phoneme features to obtain preliminary fusion features;

[0121] Step S303, extracting context features from the preliminary fusion features based on the second feature processing sub-network to obtain speech fusion features.

[0122] The following is a detailed description of steps S301 to S303.

[0123] In step S301 of some embodiments, when extracting context features from the speech representation features based on the first feature processing sub-network, an attention mechanism is introduced, and the context-related speech features of the speech representation features are obtained through attention calculation, and the obtained context-related speech feature information is fused into the speech representation features to form speech context features, so that the speech context features contain richer speech feature information.

[0124] In step S302 of some embodiments, the speech context features and the speech phoneme features are added or concatenated to achieve the feature fusion of the speech context features and the speech phoneme features, and preliminary fusion features are obtained. Among them, the preliminary fusion features contain both the phoneme-level feature information of the speech segment and the speech feature information of the speech segment.

[0125] In step S303 of some embodiments, the specific implementation process of extracting context features from the preliminary fusion features based on the second feature processing sub-network is basically the same as the specific implementation process of extracting context features from the speech quantization features based on the first feature processing sub-network in step S301 above. The difference is that the features targeted by the context feature extraction of the two are different. For the sake of brevity, it will not be elaborated.

[0126] The advantage of this embodiment is that in the process of speech recognition, the speech representation features and the speech phoneme features after context extraction are fused, so that the speech recognition can simultaneously utilize the phoneme-level information and the discrete speech feature information of the speech, improve the recognition accuracy of speech for various language types, and thus improve the accuracy of speech recognition.

[0127] In step S105 of some embodiments, the speech fusion features are recognized based on a linear layer to obtain a text segment corresponding to the speech segment, and the target text corresponding to the target speech data is determined based on the text segment.

[0128] The text segment is used to indicate the semantic content of the speech segment.

[0129] The target text is used to indicate the overall semantic content of the target speech data.

[0130] In the specific implementation of this embodiment, first, the speech fusion features are input into the linear layer. Then, the linear layer performs feature decoding on the speech fusion features to generate a text sequence corresponding to the speech fusion features, and the text sequence is used as the text segment corresponding to the speech segment. Further, for multiple speech segments in the target speech data, in chronological order, the text segments corresponding to each speech segment are spliced to form a complete text, thereby obtaining the target text corresponding to the target speech data.

[0131] Please refer to Figure 4 , which is a detailed implementation diagram of the speech recognition method according to the embodiments of the present application. Specifically, first, the target speech data is input into the speech recognition model. On the one hand, the speech feature extraction network extracts features from the speech segments in the target speech data to obtain the speech representation features of each speech segment, and its specific implementation process is similar to the above step S103. On the other hand, for each speech segment in the target speech data, the target language type of each speech segment is determined based on the language classification sub-network in the phoneme feature extraction network, and according to the target language type, the speech segment is input into the phoneme feature extractor corresponding to the target language type for phoneme feature extraction to obtain the speech phoneme features corresponding to each speech segment, and its specific implementation process is similar to the above step S102. Further, for each speech segment, the speech representation features of the speech segment are input into the first feature processing sub-network of the speech context network. Then, the output of the first feature processing sub-network is fused with the speech phoneme features of the speech segment, and the fused features are input into the second feature processing sub-network for context feature extraction to obtain speech fusion features, and its specific implementation process is similar to the above steps S301 - S303. Finally, the linear layer performs feature decoding on the speech fusion features of each speech segment to obtain the text segments of each speech segment, and the text segments are spliced according to the order of the speech segments to form the target text, and its specific implementation process is similar to the above step S105. For the sake of brevity, it will not be elaborated here.

[0132] Next, the training method of the speech recognition model according to the embodiments of the present application will be described in detail.

[0133] Figure 5It is an alternative flowchart of the training method of the speech recognition model provided by the embodiments of the present application. Figure 5 The training method in

[0134] Step S501: Obtain sample speech data.

[0135] Step S502: Input the sample speech segment, the sample text, and the candidate language type into the speech recognition model.

[0136] Step S503: Based on the language classification sub-network, perform type recognition on the sample speech segment to obtain the sample language type of the sample speech segment, and based on the phoneme feature extractor corresponding to the sample language type, perform phoneme feature extraction on the sample speech segment to obtain the sample phoneme feature of the sample speech segment.

[0137] Step S504: Based on the network structure of the phoneme feature extractor corresponding to the sample language type, divide the speech context network into a first sub-network and a second sub-network.

[0138] Step S505: Based on the speech feature extraction network, perform feature extraction on the sample speech segment to obtain the sample speech representation feature of the sample speech segment.

[0139] Step S506: Based on the first sub-network, perform context feature extraction on the sample speech representation feature to obtain the first sample context feature.

[0140] Step S507: Perform feature fusion on the first sample context feature and the sample phoneme feature to obtain the sample fused speech feature.

[0141] Step S508: Based on the second sub-network, perform context feature extraction on the sample fused speech feature to obtain the second sample context feature.

[0142] Step S509: Based on the linear layer and the second sample context feature, perform speech recognition to obtain the predicted text of the sample speech segment, and based on the comparison between the predicted text and the sample text, adjust the model parameters of the speech recognition model to train the speech recognition model.

[0143] The following will describe steps S501 to S509 in detail.

[0144] In step S501 of some embodiments, the sample speech data includes multiple sample speech segments, the sample text corresponding to each sample speech segment, and the candidate language type. Among them, the sample text is used to indicate the semantic content of the sample speech segment, and the candidate language type is used to indicate the language type to which the audio in the sample speech segment belongs.

[0145] In the embodiments of the present application, the process of obtaining the sample speech data is basically the same as the process of obtaining the target speech data in step S101 above. For the sake of brevity, it will not be elaborated here. Among them, the sample text and candidate language type of the sample speech segment can be determined in advance by manual annotation or artificial intelligence technology, without limitation.

[0146] In step S502 of some embodiments, the sample speech segment, the sample text, and the candidate language type are input into the speech recognition model. The speech recognition model includes a phoneme feature extraction network, a speech feature extraction network, a speech context network, and a linear layer; the phoneme feature extraction network includes a language classification sub-network and phoneme feature extractors corresponding to multiple candidate language types.

[0147] In the embodiments of the present application, the language classification sub-network in the phoneme feature extraction network and the phoneme feature extractors corresponding to multiple candidate language types are all pre-trained; the speech feature extraction network and the speech context network are also jointly pre-trained. Among them, the training of the language classification sub-network is supervised training; the training of the phoneme feature extractors corresponding to multiple candidate language types and the joint training of the speech feature extraction network and the speech context network are unsupervised training.

[0148] For the sake of brevity, in the embodiments of the present application, the training process of the language classification sub-network, the training of the phoneme feature extractors, and the joint training of the speech feature extraction network and the speech context network will be described in detail below, and will not be elaborated here.

[0149] In step S503 of some embodiments, the specific implementation process of type recognition of the sample speech segment based on the language classification sub-network is similar to step S201 above. The specific implementation process of phoneme feature extraction of the sample speech segment based on the phoneme feature extractor corresponding to the sample language type is similar to steps S202 - S203 above. For the sake of brevity, it will not be elaborated here.

[0150] In step S504 of some embodiments, at least a part of the network structures of the speech context network and the phoneme feature extractor are the same. Generally, the network structure of the speech context network has one more level than that of the phoneme feature extractor. Based on this, first determine the levels in the network structure of the phoneme feature extractor corresponding to the sample language type; then, according to the levels in the network structure of the phoneme feature extractor, starting from the first level of the speech context network, the hierarchical structure part with the same level as the network structure of the phoneme feature extractor is divided into the first sub-network, and the part of the hierarchical structure that is more than the network structure of the phoneme feature extractor in the speech context network is used as the second sub-network.

[0151] For example, the speech context network contains 10 layers; the network structure of the phoneme feature extractor contains 7 levels. Based on this, the first 7 levels of the speech context network are used as the first sub-network, and the remaining 3 levels are used as the second sub-network.

[0152] In step S505 of some embodiments, the specific implementation process of step S505 is similar to the specific implementation process of the above step S103. The difference is that step S505 is feature quantization during model training, while step S103 is feature quantization during model application, and the implementation stages of the two steps are different. To save space, it will not be elaborated here.

[0153] In steps S506 - S508 of some embodiments, the specific implementation process of steps S506 - S508 is similar to the specific implementation process of the above steps S301 to S303. The difference is that steps S506 - S508 are feature fusion during model training, while steps S301 to S303 are feature fusion during model application, and the implementation stages of the two steps are different. To save space, it will not be elaborated here.

[0154] In step S509 of some embodiments, the specific implementation process of performing speech recognition based on the linear layer and the second sample context feature to obtain the predicted text of the sample speech segment is similar to the specific implementation process of the above step S105. To save space, it will not be elaborated here.

[0155] Furthermore, based on the comparison between the predicted text and the sample text, the model parameters of the speech recognition model are adjusted to train the speech recognition model.

[0156] In the specific implementation of this embodiment, the difference between the predicted text and the sample text can characterize the speech recognition accuracy of the speech recognition model. When the predicted text and the sample text are relatively close, it indicates that the speech recognition accuracy of the speech recognition model is better. Specifically, since the sample text is predetermined, supervised training can be adopted for the comparison between the predicted text and the sample text, and a loss function such as the cross-entropy loss function is used to calculate the difference between the predicted text and the sample text to obtain a loss function value. Further, according to the magnitude relationship between the loss function value and the preset loss threshold, the model parameters of the speech recognition model are adjusted, and the speech recognition model is iteratively trained according to the foregoing steps S501 - S509 until at a certain iteration round, the loss function value is less than or equal to the loss threshold, the update of the model parameters is stopped, and the model parameters of the current iteration round are used as the final model parameters to obtain the trained speech recognition model.

[0157] Such as Figure 6As shown, it is a detailed implementation diagram for training a speech recognition model. Specifically, the sample speech data is input into the speech recognition model. On the one hand, the sample speech segments in the sample speech data are feature-extracted through the speech feature extraction network to obtain the sample speech representation features of each sample speech segment. The specific implementation process is similar to the above-mentioned step S103. On the other hand, for each sample speech segment in the sample speech data, the sample language type of each sample speech segment is determined based on the language classification sub-network in the phoneme feature extraction network, and according to the sample language type, the sample speech segment is input into the phoneme feature extractor corresponding to the sample language type for phoneme feature extraction to obtain the sample speech phoneme features corresponding to each sample speech segment. The specific implementation process is similar to the above-mentioned step S102. Further, for each sample speech segment, the sample speech representation feature of the sample speech segment is input into the first sub-network of the speech context network. Then, the output of the first sub-network is feature-fused with the sample speech phoneme feature of the sample speech segment, and the fused feature is input into the second sub-network for context feature extraction to obtain the second sample context feature. The specific implementation process is similar to the above-mentioned steps S301 - S303. Finally, the second sample context features of each sample speech segment are feature-decoded through a linear layer to obtain the sample text segments of each sample speech segment, and the sample text segments are concatenated according to the order of the sample speech segments to form a predicted text. The specific implementation process is similar to the above-mentioned step S105. Further, according to the difference between the sample text and the predicted text, the loss value loss1 of the speech recognition model is determined, and according to the difference degree between the loss value loss1 and the preset loss threshold, the model parameters of the speech recognition model are continuously adjusted to train the speech recognition model. The specific implementation process is similar to the above-mentioned step S509. To save space, it will not be elaborated further.

[0158] The training method of the speech recognition model in the embodiments of the present application forms a speech recognition model through the combination of a multi-language recognition network (i.e., the speech feature extraction network and the speech context network), a language classification sub-network, and phoneme feature extractors of various language types, and adopts a supervised learning method to train the speech recognition model, which can realize the fine-tuning processing of the model parameters of the speech recognition model, so that the speech recognition accuracy of the speech recognition model reaches the preset requirements, thereby improving the speech recognition accuracy of the speech recognition model for multi-language speech.

[0159] Please refer to Figure 7 , in some embodiments, the pre-training process of the language classification sub-network includes but is not limited to steps S701 to S703:

[0160] Step S701, based on a preset first pre-training network, predict the language type of the sample speech segment to obtain the predicted language type of the sample speech segment;

[0161] Step S702: Determine a first loss function based on the comparison between the predicted language type and the candidate language type.

[0162] Step S703: Based on the first loss function, adjust the parameters of the first pre-trained network to train the first pre-trained network and obtain a language classification sub-network.

[0163] The following provides a detailed description of steps S701 to S703.

[0164] In step S701 of some embodiments, the first pre-trained network can be a common multi-language pre-trained model. The input of the first pre-trained network is sample speech segments of various language types and the candidate language types of each sample speech segment. Specifically, the specific implementation process of the first pre-trained network for predicting the language type of the sample speech segment is similar to the specific implementation process of step S201 above. For the sake of brevity, it will not be elaborated here.

[0165] In step S702 of some embodiments, the difference between the predicted language type and the candidate language type can characterize the language classification accuracy of the first pre-trained network. When the predicted language type of a certain sample speech segment is the same as the candidate language type, it indicates that the classification of the first pre-trained network for this sample speech segment is accurate. Specifically, since the candidate language type is predetermined, supervised training can be adopted for the comparison between the predicted language type and the candidate language type, and loss functions such as the cross-entropy loss function are used to calculate the difference between the predicted language type and the candidate language type to obtain the first loss function.

[0166] In step S703 of some embodiments, according to the first loss function, adjust the network parameters of the first pre-trained network with the goal of minimizing the first loss function. Iteratively train the first pre-trained network according to the aforementioned steps S701 - S703 until, in a certain iteration round, the output value of the first loss function is the smallest, stop updating the network parameters, use the network parameters of the current iteration round as the final network parameters, and use the first pre-trained network of the current iteration round as the language classification sub-network.

[0167] The advantage of this embodiment is that the first pre-trained network is trained in a supervised training manner, enabling the classification ability of the first pre-trained network for speech segments of different language types to meet the predetermined requirements, obtaining a language classification sub-network, and enabling the use of the trained language classification sub-network to determine the language type of each speech segment during the speech recognition process, thereby improving the selection accuracy and reliability of the phoneme feature extractor.

[0168] Please refer to Figure 8, in some embodiments, the process of training the phoneme feature extractor corresponding to the candidate language type may include but is not limited to steps S801 to S802:

[0169] Step S801, for each candidate language type, input the sample speech segment of the candidate language type into the preset second pre-trained network to obtain the phoneme feature sensitivity of each layer of the second pre-trained network to the sample speech segment;

[0170] Step S802, based on the phoneme feature sensitivity, adjust the structure of the second pre-trained network to obtain the phoneme feature extractor of the candidate language type.

[0171] The following is a detailed description of steps S801 to S802.

[0172] In step S801 of some embodiments, for each candidate language type, input the sample speech segment of the candidate language type into the preset second pre-trained network, and sequentially extract features from the sample speech segment based on each layer of the second pre-trained network, so that each layer generates a sample phoneme feature. Further, for each layer, detect the feature integrity of the sample phoneme feature, and use the detection result as the phoneme feature sensitivity of the layer to the sample speech segment.

[0173] It should be noted that the feature integrity detection can be implemented by using a method of scoring based on an activation function, and the output of the activation function is used as the phoneme feature sensitivity. Among them, the activation function can be a softmax function, a sigmiod function, etc.

[0174] In step S802 of some embodiments, first, compare the phoneme feature sensitivities of each layer and determine which layer has the maximum phoneme feature sensitivity. Then, determine the layer with the maximum phoneme feature sensitivity as the last layer of the phoneme feature extractor, and discard all layers after the layer with the maximum phoneme feature sensitivity in the second pre-trained network, so as to realize the structural adjustment of the second pre-trained network and obtain the phoneme feature extractor of the candidate language type.

[0175] Such as Figure 9As shown, it is a detailed implementation diagram for optimizing the structure of the second pre-training network. Specifically, the second pre-training network includes n layers, which are successively connected as the first layer, the second layer, …, the kth layer, the (k + 1)th layer, …, the nth layer, where n and k are both positive integers and k is less than n. For each candidate language type, the sample speech segments of this candidate language type in the sample speech data are input into the second pre-training network, and feature extraction is successively performed on the sample speech segments based on each layer of the second pre-training network, so that each layer generates a sample phoneme feature. Further, for each layer, the integrity of the sample phoneme feature is detected, and the detection result is used as the phoneme feature sensitivity of the layer to the sample speech segment. Among them, the sensitivity of the first layer is 0.3, the sensitivity of the second layer is 0.24; the sensitivity of the third layer is 0.15; the sensitivity of the fourth layer is 0.2, …, the sensitivity of the kth layer is 0.56, the sensitivity of the (k + 1)th layer is 0.88, …, the sensitivity of the nth layer is 0.66. Since the sensitivity of the (k + 1)th layer is the largest, the (k + 1)th layer is determined as the last layer of the phoneme feature extractor. The first layer to the (k + 1)th layer of the second pre-training network are retained, the (k + 2)th layer to the nth layer are discarded, and the neural network formed by the retained first layer to the (k + 1)th layer is used as the phoneme feature extractor for the trained candidate language type.

[0176] The advantage of this embodiment is that it utilizes the characteristic that different layers in the second pre-training network have different sensitivities to the phoneme information of speech segments; when training the second pre-training network, the phoneme feature sensitivity is calculated for each layer, and according to the magnitudes of the phoneme feature sensitivities of each layer, the network structure of the second pre-training network is optimized. The layer with the largest phoneme feature sensitivity is used as the last layer, and the layers after the layer with the largest phoneme feature sensitivity are discarded, so that the phoneme layer feature information output by the trained phoneme feature extractor is more accurate, achieving the effect of improving the feature quality of the extracted speech phoneme features, thereby making the phoneme layer information incorporated in the speech recognition process more rich and accurate to improve the speech recognition accuracy.

[0177] Please refer to Figure 10 , in some embodiments, the process of jointly training the speech feature extraction network and the speech context network may include but is not limited to steps S1010 to S1050:

[0178] Step S1010, perform convolution processing on multiple sample speech segments based on the speech feature extraction network to obtain the sample speech features of each sample speech segment;

[0179] Step S1020, select the target speech feature from multiple sample speech features and mask the target speech feature;

[0180] Step S1030: Input the non-target speech features among multiple sample speech features and the masked target speech features into the speech context network to obtain a mask prediction result;

[0181] Step S1040: Determine the second loss function based on the comparison between the target speech features and the mask prediction result;

[0182] Step S1050: Update the parameters of the speech feature extraction network based on the second loss function, and update the parameters of the speech context network based on the second loss function.

[0183] The following will describe steps S1010 to S1050 in detail.

[0184] In the embodiments of the present application, in order to improve the effect of feature fusion, the speech context network and the second pre-training network adopt the same neural network structure. For example, both the speech context network and the second pre-training network adopt the wav2vec2.0 model, and both the speech context network and the second pre-training network are composed of multiple layers of transformer structures. The speech feature extraction network can adopt a multi-layer CNN model.

[0185] The above Figure 9 For example, when the phoneme feature extractor is from the first layer to the k + 1 layer, the speech context network is from the first layer to the n layer. The neural network formed by the first layer to the k + 1 layer is used as the first sub-network of the speech context network, and the neural network formed by the k + 2 layer to the n layer is used as the second sub-network of the speech context network.

[0186] In step S1010 of some embodiments, multiple sample speech segments are sequentially input into the speech feature extraction network. Based on the speech feature extraction network, convolutional processing is performed on the multiple sample speech segments to extract the speech feature information of each sample speech segment, and the sample speech feature of each sample speech segment is obtained.

[0187] In step S1020 of some embodiments, a predetermined number of sample speech features are randomly selected from the multiple sample speech features as the target speech features, and the selected target speech features are replaced with predetermined values or set to null values to implement mask processing of the target speech features. Among them, the predetermined number can be determined according to the training accuracy. When the training accuracy is high, the predetermined number is set to a larger value, and when the training accuracy is low, the predetermined number is set to a smaller value.

[0188] For example, for a sample speech data, the sample speech features include sample speech feature 1, sample speech feature 2, sample speech feature 3, sample speech feature 4, sample speech feature 5, and sample speech feature 6. Two of them are randomly selected as target speech features for masking, and the resulting masked feature sequence is "sample speech feature 1, mask, sample speech feature 3, sample speech feature 4, mask, sample speech feature 6".

[0189] In step S1030 of some embodiments, first, the non-target speech features and the masked target speech features among the multiple sample speech features are input into the speech context network in the chronological order of the speech segments. The speech context network is used to learn the context correlation information between the non-target speech features and the masked target speech features, and predict the true feature information of the masked target speech features to generate a mask prediction result.

[0190] For example, for the masked feature sequence "sample speech feature 1, mask, sample speech feature 3, sample speech feature 4, mask, sample speech feature 6", the speech context network predicts the feature information at the two mask positions according to the feature sequence, obtains prediction feature 1 and prediction feature 2, and uses prediction feature 1 and prediction feature 2 as the mask prediction results to compare prediction feature 1 with sample speech feature 2 and compare prediction feature 2 with sample speech feature 5.

[0191] In step S1040 of some embodiments, the specific implementation process of determining the second loss function based on the comparison between the target speech features and the mask prediction results is similar to the specific implementation process of step S702 above. For the sake of brevity, it will not be elaborated here.

[0192] In step S1050 of some embodiments, when updating the network parameters of the speech feature extraction network and the speech context network based on the second loss function, the network parameters of the speech feature extraction network and the speech context network are jointly adjusted, and minimizing the output value of the second loss function is used as the training objective. Iterative training is performed according to steps S1010 to S1050 until a certain iteration round when the output value of the second loss function is minimized. The network parameters of the speech feature extraction network and the speech context network at this iteration round are used as the final parameters to complete the joint training of the speech feature extraction network and the speech context network.

[0193] The advantage of this embodiment is that in an unsupervised training manner, first, a voice extraction network is used to determine the sample voice features of each sample voice segment, and a certain part of the sample voice features is randomly masked. Then, based on the context information of the voice segment, a voice context network predicts the masked sample voice features to obtain a masked prediction result. By comparing the masked prediction result with the actual sample voice features for consistency, it is determined whether the voice processing capabilities of the voice feature extraction network and the voice context network meet the predetermined requirements. Thus, through continuous iterative training, a voice feature extraction network and a voice context network that meet the predetermined requirements are obtained, which can effectively improve the learning of the voice context network and the voice extraction network for voice feature information and improve the speech recognition accuracy of the trained voice context network and voice extraction network.

[0194] Please refer to Figure 11 , the embodiment of the present application further provides a speech recognition device that can implement the above speech recognition method. The device includes:

[0195] An acquisition module 1110, configured to acquire target voice data, where the target voice data includes voice segments of at least two language types;

[0196] A phoneme feature extraction module 1120, configured to perform type recognition on the voice segment based on the phoneme feature extraction network of a pre-trained speech recognition model to obtain the target language type of the voice segment, and extract the corresponding speech phoneme features of the voice segment based on the target language type and the phoneme feature extraction network, where the speech recognition model includes a voice feature extraction network, a voice context network, and a linear layer;

[0197] A feature extraction module 1130, configured to perform feature extraction on the voice segment based on the voice feature extraction network to obtain a voice representation feature;

[0198] A feature fusion module 1140, configured to perform context extraction on the voice representation feature based on the voice context network to obtain a voice context feature, and perform feature fusion on the voice context feature and the speech phoneme feature based on the voice context network to obtain a voice fusion feature;

[0199] An identification module 1150, configured to perform identification processing on the voice fusion feature based on the linear layer to obtain a text segment corresponding to the voice segment, and determine a target text corresponding to the target voice data based on the text segment.

[0200] The specific implementation manner of this speech recognition device is basically the same as the specific embodiment of the above speech recognition method, and will not be elaborated here.

[0201] Please refer to Figure 12, an embodiment of the present application further provides a training device for a speech recognition model, which can implement the above-mentioned training method of the speech recognition model. The training device includes:

[0202] A sample acquisition module 1210, configured to acquire sample speech data, where the sample speech data includes multiple sample speech segments, the sample texts corresponding to each sample speech segment, and the candidate language types;

[0203] An input module 1220, configured to input the sample speech segment, the sample text, and the candidate language type into the speech recognition model, where the speech recognition model includes a phoneme feature extraction network, a speech feature extraction network, a speech context network, and a linear layer; the phoneme feature extraction network includes a language classification sub-network and phoneme feature extractors corresponding to multiple candidate language types;

[0204] A type recognition module 1230, configured to perform type recognition on the sample speech segment based on the language classification sub-network to obtain the sample language type of the sample speech segment, and perform phoneme feature extraction on the sample speech segment based on the phoneme feature extractor corresponding to the sample language type to obtain the sample phoneme feature of the sample speech segment;

[0205] A partitioning module 1240, configured to partition the speech context network into a first sub-network and a second sub-network based on the network structure of the phoneme feature extractor corresponding to the sample language type;

[0206] A first extraction module 1250, configured to perform feature extraction on the sample speech segment based on the speech feature extraction network to obtain the sample speech representation feature of the sample speech segment;

[0207] A second extraction module 1260, configured to perform context feature extraction on the sample speech representation feature based on the first sub-network to obtain the first sample context feature;

[0208] A fusion module 1270, configured to perform feature fusion on the first sample context feature and the sample phoneme feature to obtain the sample fusion speech feature;

[0209] A third extraction module 1280, configured to perform context feature extraction on the sample fusion speech feature based on the second sub-network to obtain the second sample context feature;

[0210] An adjustment module 1290, configured to perform speech recognition based on the linear layer and the second sample context feature to obtain the predicted text of the sample speech segment, and adjust the model parameters of the speech recognition model based on the comparison between the predicted text and the sample text to train the speech recognition model.

[0211] The specific implementation manner of the training device for the speech recognition model is basically the same as the specific embodiment of the above-mentioned training method of the speech recognition model, and will not be elaborated here.

[0212] The embodiment of the present application also provides an electronic device, which includes: a memory, a processor, a program stored on the memory and executable on the processor, and a data bus for realizing the connection and communication between the processor and the memory. When the program is executed by the processor, it realizes the above-mentioned speech recognition method or the training method of the speech recognition model. The electronic device can be any intelligent terminal including a tablet computer, a vehicle-mounted computer, etc.

[0213] Please refer to Figure 13 , Figure 13 which schematically shows the hardware structure of the electronic device in another embodiment. The electronic device includes:

[0214] A processor 1301, which can be implemented in ways such as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present application;

[0215] A memory 1302, which can be implemented in forms such as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1302 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1302 and are called by the processor 1301 to execute the speech recognition method or the training method of the speech recognition model in the embodiments of the present application;

[0216] An input / output interface 1303, which is used to realize information input and output;

[0217] A communication interface 1304, which is used to realize the communication interaction between this device and other devices, and can communicate through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.);

[0218] A bus 1305, which transmits information between the various components of the device (such as the processor 1301, the memory 1302, the input / output interface 1303, and the communication interface 1304);

[0219] Among them, the processor 1301, the memory 1302, the input / output interface 1303, and the communication interface 1304 realize the communication connection among themselves inside the device through the bus 1305.

[0220] An embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the above-mentioned speech recognition method or the training method of the speech recognition model.

[0221] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely provided with respect to the processor, and these remote memories can be connected to the processor through a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0222] The speech recognition method, the training method of the speech recognition model, the speech recognition device, the electronic device, and the computer-readable storage medium provided by the embodiments of the present application obtain target speech data, perform type recognition on speech segments based on the phoneme feature extraction network of the speech recognition model to obtain the target language type of the speech segment, and extract the speech phoneme features of the speech segment based on the target language type and the phoneme feature extraction network. It can extract the speech phoneme level information of the language segment for different language types, improve the feature quality of the extracted speech phoneme features, and is beneficial to enriching the speech level information of the speech to be recognized. Further, the speech recognition model includes a speech feature extraction network, a speech context network, and a linear layer; the speech representation features are obtained by extracting features from the speech segment based on the speech feature quantization network, which can improve the feature processing efficiency. Then, the speech context features are extracted from the speech representation features based on the speech context network, and the speech context features and the speech phoneme features are feature-fused based on the speech context network to obtain speech fusion features, which is beneficial to improving the richness and comprehensiveness of the feature content of the speech fusion features. Finally, the speech fusion features are recognized and processed based on the linear layer to obtain the text segment corresponding to the speech segment, and the target text of the target speech data is determined based on the text segment, so that the target text can more comprehensively reflect the semantic content of the target speech data, thereby improving the recognition accuracy of multi-language speech.

[0223] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0224] Those skilled in the art can understand that Figure 1-7 the technical solutions shown in

[0225] do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown in the figures, or combine certain steps, or different steps.

[0226] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and appropriate combinations thereof.

[0227] The terms "first", "second", "third", "fourth", etc. (if any) in the description of the present application and the above drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products, or devices.

[0228] It should be understood that in the present application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or a similar expression means any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0229] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. The displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.

[0230] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place, or they may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0231] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0232] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. And the aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store programs.

[0233] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings. However, this does not limit the scope of the rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of the rights of the embodiments of the present application.

Claims

1. A speech recognition method, characterized in that, The method includes: Obtaining target speech data, where the target speech data includes speech segments of at least two language types; Performing type recognition on the speech segments based on the phoneme feature extraction network of a pre-trained speech recognition model to obtain the target language type of the speech segments, and extracting the corresponding speech phoneme features of the speech segments based on the target language type and the phoneme feature extraction network, where the speech recognition model includes a speech feature extraction network, a speech context network, and a linear layer; Performing feature extraction on the speech segments based on the speech feature extraction network to obtain speech representation features; Performing context extraction on the speech representation features based on the speech context network to obtain speech context features, and performing feature fusion on the speech context features and the speech phoneme features based on the speech context network to obtain speech fusion features; Performing recognition processing on the speech fusion features based on the linear layer to obtain the text segment corresponding to the speech segment, and determining the target text corresponding to the target speech data based on the text segment.

2. The voice recognition method according to claim 1, wherein The phoneme feature extraction network includes a language classification sub-network and a phoneme feature extractor corresponding to each candidate language type; The performing type recognition on the speech segments based on the phoneme feature extraction network of a pre-trained speech recognition model to obtain the target language type of the speech segments, and extracting the corresponding speech phoneme features of the speech segments based on the target language type and the phoneme feature extraction network includes: Performing language type recognition on the speech segments based on the language classification sub-network and multiple candidate language types to obtain the target language type of the speech segments; Determining the phoneme feature extractor corresponding to the target language type among the phoneme feature extractors corresponding to multiple candidate language types; Performing phoneme feature extraction on the speech segments based on the phoneme feature extractor corresponding to the target language type to obtain the speech phoneme features.

3. The voice recognition method according to claim 1, wherein The speech context network includes a first feature processing sub-network and a second feature processing sub-network; The performing context extraction on the speech representation features based on the speech context network to obtain speech context features, and performing feature fusion on the speech context features and the speech phoneme features based on the speech context network to obtain speech fusion features includes: Performing context feature extraction on the speech representation features based on the first feature processing sub-network to obtain speech context features; Performing feature fusion on the speech context features and the speech phoneme features to obtain a preliminary fusion feature; Performing context feature extraction on the preliminary fusion feature based on the second feature processing sub-network to obtain the speech fusion features.

4. A training method for a speech recognition model, characterized in that, The training method includes: Obtaining sample speech data, where the sample speech data includes multiple sample speech segments, the sample text corresponding to each sample speech segment, and candidate language types; Input the sample speech segment, the sample text, and the candidate language type into a speech recognition model, where the speech recognition model includes a phoneme feature extraction network, a speech feature extraction network, a speech context network, and a linear layer; the phoneme feature extraction network includes a language classification sub-network and phoneme feature extractors corresponding to multiple candidate language types; Based on the language classification sub-network, perform type recognition on the sample speech segment to obtain the sample language type of the sample speech segment, and based on the phoneme feature extractor corresponding to the sample language type, perform phoneme feature extraction on the sample speech segment to obtain the sample phoneme features of the sample speech segment; Based on the network structure of the phoneme feature extractor corresponding to the sample language type, divide the speech context network into a first sub-network and a second sub-network; Based on the speech feature extraction network, perform feature extraction on the sample speech segment to obtain the sample speech representation features of the sample speech segment; Based on the first sub-network, perform context feature extraction on the sample speech representation features to obtain first sample context features; Perform feature fusion on the first sample context features and the sample phoneme features to obtain sample fused speech features; Based on the second sub-network, perform context feature extraction on the sample fused speech features to obtain second sample context features; Based on the linear layer and the second sample context features, perform speech recognition to obtain the predicted text of the sample speech segment, and based on the comparison between the predicted text and the sample text, adjust the model parameters of the speech recognition model to train the speech recognition model.

5. The training method according to claim 4, wherein The language classification sub-network is trained through the following method: Based on a preset first pre-training network, perform language type prediction on the sample speech segment to obtain the predicted language type of the sample speech segment; Based on the comparison between the predicted language type and the candidate language type, determine a first loss function; Based on the first loss function, adjust the parameters of the first pre-training network to train the first pre-training network to obtain the language classification sub-network.

6. The training method according to claim 4, characterized in that, The phoneme feature extractors corresponding to the multiple candidate language types are trained through the following method: For each candidate language type, input the sample speech segment of the candidate language type into a preset second pre-training network to obtain the phoneme feature sensitivities of each layer of the second pre-training network for the sample speech segment; Based on the phoneme feature sensitivities, adjust the structure of the second pre-training network to obtain the phoneme feature extractor of the candidate language type.

7. The training method according to any one of claims 4 to 6, characterized in that The speech feature extraction network and the speech context network are jointly trained through the following method: Based on the speech feature extraction network, perform convolutional processing on multiple sample speech segments to obtain the sample speech features of each sample speech segment; Select target speech features from the multiple sample speech features and mask the target speech features; Input the non-target speech features among the multiple sample speech features and the masked target speech features into the speech context network to obtain a mask prediction result; Determine a second loss function based on the comparison between the target speech features and the mask prediction result; Update the parameters of the speech feature extraction network based on the second loss function, and update the parameters of the speech context network based on the second loss function.

8. A voice recognition device, characterized in that, The device includes: An acquisition module, configured to acquire target speech data, where the target speech data includes speech segments of at least two language types; A phoneme feature extraction module, configured to perform type recognition on the speech segment based on the phoneme feature extraction network of a pre-trained speech recognition model to obtain the target language type of the speech segment, and extract the corresponding speech phoneme feature of the speech segment based on the target language type and the phoneme feature extraction network, where the speech recognition model includes a speech feature extraction network, a speech context network, and a linear layer; A feature extraction module, configured to extract features from the speech segment based on the speech feature extraction network to obtain speech representation features; A feature fusion module, configured to perform context extraction on the speech representation features based on the speech context network to obtain speech context features, and perform feature fusion on the speech context features and the speech phoneme features based on the speech context network to obtain speech fusion features; An identification module, configured to perform identification processing on the speech fusion features based on the linear layer to obtain a text segment corresponding to the speech segment, and determine a target text corresponding to the target speech data based on the text segment.

9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it implements the speech recognition method according to any one of claims 1 to 3, or the training method of the speech recognition model according to any one of claims 4 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the speech recognition method according to any one of claims 1 to 3, or the training method of the speech recognition model according to any one of claims 4 to 7.