Speech recognition method and device, equipment, storage medium and program product

By coding and fusion decoding the voice of hearing-impaired people twice, the problem of poor performance of speech recognition for hearing-impaired people in the prior art is solved, and higher accuracy and applicability of speech recognition are achieved.

CN119943035APending Publication Date: 2025-05-06ANHUI IFLYTEK UNIVERSAL LANGUAGE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510064642.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing speech recognition methods have poor voice recognition effects on hearing-impaired people. How to improve the accuracy of speech recognition for hearing-impaired people has become a technical problem that needs to be solved urgently.

Method used

By first encoding and second encoding of the recognized speech, encoding features representing the speech intelligibility level and speech content are obtained, and then fused and decoded to improve the accuracy of speech recognition.

Benefits of technology

By considering the intelligibility level of speech and integrating coding features for decoding, the accuracy of speech recognition for hearing-impaired people is significantly improved and the scope of application of speech recognition is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943035A_ABST
    Figure CN119943035A_ABST
Patent Text Reader

Abstract

The invention discloses a voice recognition method and device, equipment, a storage medium and a program product, and relates to the technical field of artificial intelligence, and the method comprises the steps: carrying out the first coding of a to-be-recognized voice, and obtaining a first coding feature; the first coding feature is an initial coding feature representing the intelligibility level of the speech to be recognized; performing second coding on the to-be-recognized voice to obtain a second coding feature; the second coding feature is an initial coding feature representing the voice content of the to-be-recognized voice; fusing the first coding feature and the second coding feature to obtain a fused coding feature; and decoding the fused coding features to obtain a speech recognition result. According to the voice recognition method provided by the embodiment of the invention, when the coding features of the to-be-recognized voice are decoded, the intelligibility level of the to-be-recognized voice is considered, so that the accuracy of voice recognition is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a speech recognition method, apparatus, device, storage medium and program product. Background Art

[0002] Speech recognition is a technology that automatically converts speech into text. It has a wide range of applications, such as voice assistants, smart homes, and car voice interactions.

[0003] The current speech recognition method has a good effect on speech recognition for people with normal hearing, but a poor effect on speech recognition for people with hearing impairments. Therefore, how to improve the accuracy of speech recognition for people with hearing impairments has become a technical problem that needs to be solved urgently. Summary of the invention

[0004] In view of the above problems, the present application provides a speech recognition method, apparatus, device, storage medium and program product to improve the accuracy of speech recognition. The specific scheme is as follows:

[0005] The first aspect of the present application provides a speech recognition method, comprising:

[0006] Performing a first encoding on the speech to be recognized to obtain a first encoding feature; the first encoding feature represents the intelligibility level of the speech to be recognized;

[0007] Performing a second encoding on the speech to be recognized to obtain a second encoding feature; the second encoding feature represents the speech content of the speech to be recognized;

[0008] Fusing the first coding feature and the second coding feature to obtain a fused coding feature;

[0009] The fused coding features are decoded to obtain a speech recognition result.

[0010] In a possible implementation, the first encoding of the speech to be recognized includes:

[0011] Extracting acoustic features of the speech to be recognized;

[0012] Encoding the acoustic features through a first network to obtain a first hidden layer feature;

[0013] The first hidden layer feature is encoded through a second network to obtain a second hidden layer feature as the first encoded feature; the complexity of the second network is greater than the complexity of the first network.

[0014] In a possible implementation, performing a second encoding on the speech to be recognized includes:

[0015] Extracting acoustic features of the speech to be recognized;

[0016] Encoding the acoustic features through a third network to obtain a third hidden layer feature;

[0017] The third hidden layer feature is encoded based on the attention mechanism through the third network to obtain the fourth hidden layer feature as the second encoded feature.

[0018] In a possible implementation, the process of performing the first encoding and the second encoding on the speech to be recognized and decoding the fused coding feature is implemented by a speech recognition system, and the speech recognition system includes:

[0019] A first encoding module, used for performing a first encoding on the speech to be recognized to obtain a first encoding feature;

[0020] A second encoding module, used for performing a second encoding on the speech to be recognized to obtain a second encoding feature;

[0021] A decoding module, used for decoding the fused coding features to obtain a speech recognition result;

[0022] The speech recognition system is obtained by jointly training a pre-trained intelligibility level recognition model and a pre-trained speech recognition model; the pre-trained intelligibility level recognition model includes the first encoding module; the pre-trained speech recognition model includes the second encoding module and the decoding module.

[0023] In a possible implementation, the pre-trained intelligibility level recognition model further includes a classification module; and the process of pre-training the intelligibility level recognition model includes:

[0024] Inputting speech data in the first training set into the intelligibility level recognition model, encoding the speech data by the first encoding module to obtain a first encoding feature, and classifying the first encoding feature of the speech data by the classification module to obtain a classification result; the classification result represents the intelligibility level of the speech data; the first training set includes speech data of hearing-impaired people and speech data of normal hearing people;

[0025] Based on the loss between the classification result and the class label of the speech data, the parameters of the intelligibility level recognition model are updated.

[0026] In a possible implementation, the process of pre-training the speech recognition model includes:

[0027] Inputting speech data in a second training set into the speech recognition model, encoding the speech data by the second encoding module to obtain a second encoding feature, and decoding the second encoding feature of the speech data by the decoding module to obtain a speech recognition result; the second training set includes speech data of people with normal hearing;

[0028] The parameters of the speech recognition model are updated based on the loss between the speech recognition result of the speech data and the text label of the speech data, and the loss between the second encoded feature and the text label of the speech data.

[0029] In a possible implementation, the pre-trained intelligibility level recognition model further includes a classification module; and the process of jointly training the pre-trained intelligibility level recognition model and the pre-trained speech recognition model includes:

[0030] Input the speech data in the third training set into the first encoding module and the second encoding module respectively, obtain the first encoding feature obtained by encoding the speech data by the first encoding module, and obtain the second encoding feature obtained by encoding the speech data by the second encoding module; fuse the first encoding feature and the second encoding feature of the speech data to obtain a fused encoding feature; decode the fused encoding feature by the decoding module to obtain a decoded feature, normalize the decoded feature to obtain a speech recognition result; classify the decoded feature by the classification module to obtain a classification result; the third training set includes speech data of hearing-impaired people and speech data of normal-hearing people;

[0031] Based on the loss between the classification result and the class label of the speech data, and the loss between the speech recognition result of the speech data and the text label of the speech data, the parameters of the decoding module and the classification module are updated.

[0032] A second aspect of the present application provides a speech recognition device, comprising:

[0033] A first encoding unit, configured to perform a first encoding on the speech to be recognized to obtain a first encoding feature; the first encoding feature represents an intelligibility level of the speech to be recognized;

[0034] A second encoding unit is used to perform a second encoding on the speech to be recognized to obtain a second encoding feature; the second encoding feature represents the speech content of the speech to be recognized;

[0035] a fusion unit, configured to fuse the first coding feature and the second coding feature to obtain a fused coding feature;

[0036] A decoding unit is used to decode the fused coding feature to obtain a speech recognition result.

[0037] The third aspect of the present application provides a computer program product, including computer-readable instructions. When the computer-readable instructions are executed on an electronic device, the electronic device implements the speech recognition method of the first aspect or any implementation of the first aspect.

[0038] A fourth aspect of the present application provides an electronic device, comprising at least one processor and a memory connected to the processor, wherein:

[0039] The memory is used to store computer programs;

[0040] The processor is used to execute the computer program so that the electronic device can implement the speech recognition method of the above-mentioned first aspect or any implementation manner of the first aspect.

[0041] A fifth aspect of the present application provides a computer storage medium, which carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement the speech recognition method of the above-mentioned first aspect or any implementation of the first aspect.

[0042] By means of the above technical scheme, the speech recognition method, apparatus, device, storage medium and program product provided by the present application first encode the speech to be recognized to obtain a first encoding feature representing the intelligibility level of the speech to be recognized; second encode the speech to be recognized to obtain a second encoding feature representing the speech content of the speech to be recognized; fuse the first encoding feature and the second encoding feature to obtain a fused encoding feature; decode the fused encoding feature to obtain a speech recognition result. The speech recognition method provided by the embodiment of the present application introduces the intelligibility level of the speech to be recognized, obtains a fused encoding feature that fuses the intelligibility level feature and the speech content feature of the speech to be recognized, and the process of decoding the fused encoding feature takes into account the intelligibility level of the speech to be recognized, thereby improving the accuracy of speech recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the originals and elements are not necessarily drawn to scale.

[0044] Figure 1 A flow chart of an implementation of the speech recognition method provided in this application;

[0045] Figure 2A flowchart for implementing the first encoding of the speech to be recognized provided by the present application;

[0046] Figure 3 A flowchart for implementing the second encoding of the speech to be recognized provided by the present application;

[0047] Figure 4 A schematic diagram of the structure of the speech recognition system provided in this application;

[0048] Figure 5 A flowchart for implementing pre-training of an intelligibility level recognition model provided in this application;

[0049] Figure 6 A flowchart for implementing pre-training of a speech recognition model provided in this application;

[0050] Figure 7 An architecture diagram for jointly training a pre-trained intelligibility level recognition model and a pre-trained speech recognition model provided in the present application;

[0051] Figure 8 A schematic diagram of the structure of the speech recognition device provided by this application;

[0052] Fig. 9 A schematic diagram of the structure of an electronic device provided in this application. DETAILED DESCRIPTION

[0053] The following describes the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. The terms used in the implementation method section of the present application are only used to explain the specific embodiments of the present application, and are not intended to limit the present application.

[0054] The embodiments of the present application are described below in conjunction with the accompanying drawings. Those skilled in the art will appreciate that, with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0055] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and need not be used to describe a specific order or sequential order. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances, which is only to describe the distinction mode adopted by the objects of the same attributes when describing in the embodiments of the present application. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, so that the process, method, system, product or equipment comprising a series of units need not be limited to those units, but may include other units that are not clearly listed or inherent to these processes, methods, products or equipment.

[0056] Due to many factors such as hearing loss, hearing-impaired people have poor speech clarity when expressing themselves. The more severe the hearing loss, the worse the speech clarity. Poor speech clarity makes it difficult for the listener to understand the speech content. The current speech recognition solution has a good speech recognition effect for people with normal hearing, but a poor speech recognition effect for people with hearing impairment.

[0057] This application is proposed to improve the accuracy of speech recognition for hearing-impaired people.

[0058] In order to better understand the difference between the present application and the existing solutions, the existing speech recognition method is first described:

[0059] The existing speech recognition method is: encode the speech to be recognized to obtain the encoding features, and decode the encoding features to obtain the speech recognition results.

[0060] The solution of this application is explained below.

[0061] like Figure 1 As shown, a flow chart of an implementation of the speech recognition method provided in an embodiment of the present application may include:

[0062] Step S101: performing a first encoding on the speech to be recognized to obtain a first encoding feature.

[0063] The first coding feature represents the intelligibility level of the speech to be recognized. In other words, the first coding feature can be used to determine the intelligibility level of the speech to be recognized.

[0064] Optionally, the acoustic features of the speech to be recognized may be extracted, and the extracted acoustic features may be first encoded to obtain first encoded features. As an example, the first encoding may be encoding based on an attention mechanism or may not be encoding based on an attention mechanism.

[0065] Optionally, the acoustic features of the speech to be recognized can be extracted, and the extracted acoustic features can be subjected to a first sub-encoding to obtain a first hidden feature; the first hidden feature can be subjected to a second sub-encoding to obtain a second hidden feature, i.e., a first encoding feature. The second sub-encoding is different from the first sub-encoding, i.e., the way of encoding the acoustic features is different from the way of encoding the first hidden feature. As an example, the second sub-encoding can be an encoding based on an attention mechanism or can be an encoding not based on an attention mechanism.

[0066] In the embodiment of the present application, the acoustic features may include but are not limited to any of the following: perceptual linear prediction features (PLP), Mel-Frequency Cepstral Coefficients (MFCC), FilterBank, etc.

[0067] This application divides the speech intelligibility level into 5 levels, which are optional. Different levels correspond to different intelligibility ranges:

[0068] The intelligibility range corresponding to the first level intelligibility level is: [0%, 30%);

[0069] The intelligibility range corresponding to the second-level intelligibility level is: [30%, 50%);

[0070] The intelligibility range corresponding to the third level of intelligibility is: [50%, 70%);

[0071] The intelligibility range corresponding to the four-level intelligibility level is: [70%, 95%);

[0072] The intelligibility range corresponding to the five levels of intelligibility is: [95%, 100%].

[0073] The intelligibility of speech data refers to the proportion of the understood content of the speech data to all the content corresponding to the speech data. Different intelligibility levels correspond to different levels of hearing loss of the speaker. The greater the hearing loss, the lower the clarity of the speaker's speech, and the lower the corresponding intelligibility level. Among the above five levels, the first level of intelligibility corresponds to the lowest speech clarity, and the fifth level of intelligibility corresponds to the highest speech clarity.

[0074] At present, in the medical field, hearing impairment is classified into five levels according to the degree of hearing loss, namely:

[0075] Normal hearing: The hearing level is between -10 decibels and 25 decibels, which is the normal hearing range. People can hear and distinguish various sounds clearly.

[0076] Mild hearing loss: Hearing loss is between 26-40 decibels. At this time, patients may have difficulty hearing when hearing small sounds or whispers, and usually need to slightly increase the volume to hear. Although this degree of hearing loss has a relatively small impact on daily life, in certain specific situations, such as in noisy environments, patients may find it difficult to communicate.

[0077] Moderate hearing loss: Hearing loss is 41-60 decibels. Patients may have difficulty in daily verbal communication and need the speaker to repeat the content in order to hear clearly. Patients with moderate hearing loss may need hearing aids to help improve their hearing.

[0078] Severe hearing loss: Hearing loss is between 61-90 decibels, which usually has a greater impact on daily life. Patients may not be able to communicate with others normally and need the help of louder sounds or hearing aids to assist hearing.

[0079] Extreme hearing loss: Hearing loss exceeds 90 decibels. Patients may find it difficult to sense the existence of sounds, and even with hearing aids, they may find it difficult to hear external sounds clearly.

[0080] Optionally, the present application may define the intelligibility level of speech data of a person with profound hearing loss as a first-level intelligibility level;

[0081] The intelligibility level of speech data of people with severe hearing loss is defined as the secondary intelligibility level;

[0082] The intelligibility level of speech data of people with moderate hearing loss is defined as the third level of intelligibility;

[0083] The intelligibility level of speech data for people with mild hearing loss is defined as four levels of intelligibility;

[0084] The intelligibility level of speech data of normal hearing persons is defined as five levels of intelligibility levels.

[0085] Step S102: Perform a second encoding on the speech to be recognized to obtain a second encoding feature. The second encoding feature represents the speech content of the speech to be recognized. In other words, the second encoding feature can be used to determine the speech content of the speech to be recognized.

[0086] Optionally, the acoustic features of the speech to be recognized may be extracted, and the extracted acoustic features may be second-encoded to obtain second coded features. As an example, the second encoding may be an encoding based on an attention mechanism or may not be an encoding based on an attention mechanism. The second encoding is different from the first encoding.

[0087] It should be noted that step S101 and step S102 may be executed successively or simultaneously, and the present application does not specifically limit the execution order of these two steps.

[0088] Step S103: Fusing the first coding feature and the second coding feature to obtain a fused coding feature.

[0089] The fused coding features carry both intelligibility level features and speech content features.

[0090] Optionally, the first coding feature and the second coding feature may be spliced ​​(usually spliced ​​in the depth direction) to obtain a fused coding feature.

[0091] Alternatively, the first coding feature and the second coding feature may be spliced ​​together (usually in the depth direction), and the spliced ​​coding feature may be dimensionally transformed to obtain a fused coding feature.

[0092] Alternatively, the first coding feature and the second coding feature may be summed to obtain a fused coding feature.

[0093] Alternatively, the first coding feature and the second coding feature may be averaged to obtain a fused coding feature.

[0094] Step S104: Decode the fused coding features to obtain speech recognition results.

[0095] Since the fused coding features carry the intelligibility level information of the speech to be recognized, the decoding of the fused coding features adopts decoding that is adapted to the intelligibility level of the speech to be recognized, and the obtained speech recognition result is a recognition result that is adapted to the intelligibility level of the speech to be recognized.

[0096] The speech recognition method provided in the embodiment of the present application introduces the intelligibility level of the speech to be recognized, obtains a fused coding feature that combines the intelligibility level characteristics of the speech to be recognized and the speech content characteristics, and the process of decoding the fused coding feature takes into account the intelligibility level of the speech to be recognized, thereby improving the accuracy of speech recognition.

[0097] In an optional embodiment, a flowchart of implementing the first encoding of the speech to be recognized is as follows: Figure 2 As shown, it may include:

[0098] Step S201: extracting acoustic features of the speech to be recognized.

[0099] Usually, the acoustic features of each speech frame of the speech to be recognized are extracted, that is, the speech to be recognized is divided into frames to obtain multiple speech frames, and then the acoustic features of each speech frame are extracted.

[0100] Step S202: Encode the acoustic features through the first network to obtain the first hidden layer features of the speech to be recognized.

[0101] The first network may be a simple neural network, for example, it may be composed of two or more linear layers, or may be composed of two or more convolutional layers, or may adopt other network structures. The linear layer is used to perform linear transformation on input information.

[0102] The first network encodes the acoustic features of the speech to be recognized (i.e., the acoustic features of each speech frame) to obtain the first hidden features of each speech frame. That is, the first hidden features of the speech to be recognized are composed of the first hidden features of each speech frame of the speech to be recognized.

[0103] Step S203: Encode the first hidden layer feature of the speech to be recognized through the second network to obtain the second hidden layer feature of the speech to be recognized, that is, the first encoded feature. The complexity of the second network is greater than the complexity of the first network.

[0104] The second network may be a complex neural network, for example, it may be a Transform network (which may be composed of multiple Transform modules), or it may be a Long Short-term Memory Network (LSTM), or it may be a Bidirectional Long Short-term Memory Network (Bi-LSTM), or other network structures may be used.

[0105] The second network encodes the first hidden layer features of the speech to be recognized (i.e., the first hidden layer features of each speech frame) to obtain the second hidden layer features of each speech frame. That is, the second hidden layer features of the speech to be recognized are composed of the second hidden layer features of each speech frame of the speech to be recognized.

[0106] In an optional embodiment, a flowchart of implementing the second encoding of the speech to be recognized is as follows: Figure 3 As shown, it may include:

[0107] Step S301: extracting acoustic features of the speech to be recognized.

[0108] Usually, the acoustic features of each speech frame of the speech to be recognized are extracted.

[0109] Step S302: Encode the acoustic features of the speech to be recognized through a third network to obtain a third hidden layer feature of the speech to be recognized.

[0110] The third network may be a simple neural network, for example, it may be composed of two or more linear layers, or may be composed of two or more convolutional layers, or may adopt other network structures.

[0111] The third network encodes the acoustic features of the speech to be recognized (i.e., the acoustic features of each speech frame) to obtain the third hidden features of each speech frame. That is, the third hidden features of the speech to be recognized are composed of the third hidden features of each speech frame of the speech to be recognized.

[0112] Step S303: The third latent feature of the speech to be recognized is encoded based on the attention mechanism through the fourth network to obtain the fourth latent feature of the speech to be recognized, that is, the second encoded feature.

[0113] The fourth network encodes the third hidden layer feature of the speech to be recognized (i.e., the third hidden layer feature of each speech frame) to obtain the fourth hidden layer feature of each speech frame, i.e., the second encoding feature. That is, the fourth hidden layer feature of the speech to be recognized is composed of the fourth hidden layer features of each speech frame of the speech to be recognized.

[0114] As an example, the fourth network can be an encoder based on an attention mechanism, such as a Transform network, or other neural networks based on an attention mechanism.

[0115] In an optional embodiment, the above-mentioned process of performing the first encoding and the second encoding on the speech to be recognized and decoding the fused coding features can be implemented by a speech recognition system. A structural schematic diagram of the speech recognition system provided in the embodiment of the present application is as follows: Figure 4 As shown, it may include:

[0116] The first encoding module 401 is used to perform first encoding on the speech to be recognized to obtain a first encoding feature. The structure of the first encoding module 401 can be composed of the first network and the second network mentioned above.

[0117] The second encoding module 402 is used to perform a second encoding on the speech to be recognized to obtain a second encoding feature. The second encoding module 402 can be an encoder based on the attention mechanism, and can be composed of the third network and the fourth network mentioned above.

[0118] In the case where the first coding feature is the first coding feature of each speech frame of the speech to be recognized, and the second coding feature is the second coding feature of each speech frame of the speech to be recognized, assuming that the number of speech frames of the speech to be recognized is K, feature fusion can be performed by any of the following fusion methods:

[0119] The first coding feature and the second coding feature of the same speech frame can be concatenated to obtain the fused coding features of K speech frames. For example, if the first coding feature and the second coding feature of a speech frame are both 512-dimensional vectors, then the concatenation of the first coding feature and the second coding feature will obtain a 1024-dimensional vector.

[0120] Alternatively, the first encoded feature and the second encoded feature of the same speech frame can be concatenated, and the concatenated encoded feature is subjected to dimensionality transformation to obtain the fused encoded feature of K speech frames. For example, if the first encoded feature and the second encoded feature of a certain speech frame are both vectors of 512 dimensions, then the concatenation of the first encoded feature and the second encoded feature results in a vector of 1024 dimensions. Then, the 1024-dimensional vector is subjected to dimensionality transformation to obtain a 512-dimensional vector. The 1024-dimensional vector can be subjected to dimensionality transformation through a preset network.

[0121] Alternatively, the first encoded feature and the second encoded feature of the same speech frame can be summed to obtain the fused encoded feature of K speech frames.

[0122] Alternatively, the first encoded feature and the second encoded feature of the same speech frame can be averaged to obtain the fused encoded feature of K speech frames.

[0123] Optionally, in order to reduce the computational amount, the first encoded feature of K speech frames can be downsampled to obtain the downsampled first encoded feature of J (J < K) speech frames (for example, the first encoded features of two adjacent speech frames are fused into the encoded feature of one speech frame, then K = 2J); the second encoded feature of K speech frames is downsampled to obtain the downsampled second encoded feature of J speech frames; the downsampled first encoded feature of J speech frames and the downsampled second encoded feature of J speech frames are fused based on any of the foregoing fusion methods to obtain the fused encoded feature of J speech frames.

[0124] The decoding module 403 is configured to decode the fused encoded feature to obtain a speech recognition result. The decoding module 403 can be a decoder based on an attention mechanism.

[0125] Optionally, the decoding module 403 can decode the fused encoded feature to obtain a decoded feature, and perform normalization processing (such as softmax) on the decoded feature to obtain a speech recognition result.

[0126] Assume that the input to the decoding module 403 is the fused encoded feature of L (L is J or K) speech frames. Then, the decoding module 403 can decode the L fused encoded features to obtain L decoded features, and each decoded feature is a vector of Q dimensions, where Q is the total number of words in the dictionary used for speech recognition modeling.

[0127] Among them, the speech recognition system is obtained by jointly training a pre-trained intelligibility level recognition model and a pre-trained speech recognition model; the pre-trained intelligibility level recognition model includes the first encoding module 401; the pre-trained speech recognition model includes the second encoding module 402 and the decoding module 403.

[0128] The pre-trained intelligibility level recognition model is used to classify the input speech data according to the first encoding feature obtained by encoding the input speech data by the first encoding module 401, that is, to determine to which intelligibility level the input speech data belongs.

[0129] In an optional embodiment, the pre-trained intelligibility level recognition model further includes a classification module; based on this, an implementation flowchart of pre-training the intelligibility level recognition model provided in the embodiment of the present application is as follows: Figure 5 As shown, it may include:

[0130] Step S501: Input the speech data in the first training set into the intelligibility level recognition model to obtain the classification result output by the intelligibility level recognition model. Specifically, the first encoding module encodes the input speech data to obtain a first encoding feature, and the classification module classifies the first encoding feature of the speech data to obtain a classification result; the classification result represents the intelligibility level of the speech data.

[0131] The input to the intelligibility level recognition model may be acoustic features of the speech data.

[0132] The first training set includes speech data of hearing-impaired persons and speech data of hearing-normal persons. Optionally, in the first training set, the proportion of speech data of hearing-normal persons is less than the proportion of speech data of hearing-impaired persons. As an example, the proportion of speech data of hearing-normal persons in the first training set may be in the range of 1 / 10 to 1 / 5.

[0133] Optionally, the classification module may be composed of M (M is an integer greater than 0) linear layers and a normalization layer (eg, a softmax layer). The output of the classification module may be the probability that the speech data belongs to each intelligibility level.

[0134] Step S502: Based on the loss between the classification result and the level label of the speech data, the parameters of the intelligibility level recognition model are updated.

[0135] The level label of speech data can be a one-dimensional vector consisting of N elements, where N is the number of intelligibility levels. For example, N is 5. Different positions in the one-dimensional vector correspond to different levels. If the speech data belongs to the i-th (i∈{1, 2, 3, ..., N}) intelligibility level, then in the level label, the element at the position corresponding to the i-th intelligibility level is 1, and the elements at other positions are 0.

[0136] The parameters of the intelligibility level recognition model may be updated based on the cross entropy loss between the classification result and the level label of the speech data.

[0137] Optionally, the loss between the classification result and the class label of the speech data may also adopt other losses, such as logarithmic loss, mean square error loss, etc.

[0138] The intelligibility level recognition model obtained when the training end condition is met is the pre-trained intelligibility level recognition model.

[0139] In an optional embodiment, a flowchart of implementing pre-training of a speech recognition model provided in an embodiment of the present application is as follows: Figure 6 As shown, it may include:

[0140] Step S601: Input the speech data in the second training set into the speech recognition model, obtain the second coding feature obtained by the speech recognition model encoding the speech data, and the speech recognition result output by the speech recognition model. Specifically, the second encoding module encodes the speech data to obtain the second coding feature, and the decoding module decodes the second coding feature of the speech data to obtain the speech recognition result. The second training set includes speech data of people with normal hearing.

[0141] The second training set contains only speech data from people with normal hearing, and does not include speech data from people with hearing impairment.

[0142] The input to the speech recognition model may be the acoustic features of the speech data.

[0143] Step S602: Based on the loss between the speech recognition result of the speech data and the text label of the speech data, and the loss between the second encoding feature and the text label of the speech data, the parameters of the speech recognition model are updated.

[0144] The text label is composed of the individual words corresponding to the speech data.

[0145] The parameters of the speech recognition model can be updated based on the cross entropy loss between the speech recognition result of the speech data and the text label of the speech data, and the connectionist temporal classification loss (CTC) between the second encoding feature and the text label of the speech data.

[0146] Optionally, the parameters of the speech recognition model may be updated based only on the loss between the speech recognition result of the speech data and the text label of the speech data.

[0147] Optionally, the loss between the speech recognition result of the speech data and the text label of the speech data may also adopt other losses, such as logarithmic loss, mean square error loss, etc.

[0148] like Figure 7As shown, it is an architecture diagram for jointly training a pre-trained intelligibility level recognition model and a pre-trained speech recognition model provided in an embodiment of the present application. Based on the architecture diagram, the method for jointly training a pre-trained intelligibility level recognition model and a pre-trained speech recognition model provided in an embodiment of the present application can be:

[0149] The speech data in the third training set are respectively input into the first encoding module and the second encoding module to obtain the first encoding feature obtained by encoding the speech data by the first encoding module, and the second encoding feature obtained by encoding the speech data by the second encoding module; the first encoding feature and the second encoding feature of the speech data are fused to obtain the fused encoding feature; the fused encoding feature is decoded by the decoding module to obtain the decoding feature, and the decoding feature is normalized to obtain the speech recognition result; the decoding feature is classified by the classification module to obtain the classification result.

[0150] The acoustic features of the speech data may be input into the first encoding module and the second encoding module. The acoustic features of the same speech data may be input into the first encoding module and the second encoding module each time.

[0151] The third training set includes speech data of hearing-impaired persons and speech data of hearing-normal persons. Optionally, in the first training set, the proportion of speech data of hearing-normal persons is less than the proportion of speech data of hearing-impaired persons. As an example, the proportion of speech data of hearing-normal persons in the third training set may be in the range of 1 / 10 to 1 / 5.

[0152] The parameters of the decoding module and the classification module are updated based on the loss between the classification result and the class label of the speech data, and the loss between the speech recognition result of the speech data and the text label of the speech data.

[0153] When the pre-trained intelligibility level recognition model and the pre-trained speech recognition model are jointly trained, the parameters of the first encoding and the second encoding are frozen, and only the parameters of the decoding module and the classification module are updated.

[0154] The parameters of the decoding module and the classification module can be updated based on the cross entropy loss between the classification result and the class label of the speech data, and the cross entropy loss between the speech recognition result of the speech data and the text label of the speech data.

[0155] In an optional embodiment, the pre-trained speech recognition model can be further trained using the speech data in the fourth training set to obtain a trained speech recognition model.

[0156] For example, the speech data in the fourth training set can be input into a pre-trained speech recognition model to obtain a second coding feature obtained by encoding the speech data by the pre-trained speech recognition model, and a speech recognition result output by the pre-trained speech recognition model. Specifically, the second encoding module encodes the speech data to obtain a second coding feature, and the decoding module decodes the second coding feature of the speech data to obtain a speech recognition result. The fourth training set includes speech data of hearing-impaired people.

[0157] The fourth training set only contains speech data of hearing-impaired people, and does not include speech data of people with normal hearing.

[0158] The input to the pre-trained speech recognition model can be the acoustic features of the speech data.

[0159] Based on the loss between the speech recognition result of the speech data and the text label of the speech data, and the loss between the second encoded feature and the text label of the speech data, the parameters of the pre-trained speech recognition model are updated.

[0160] The parameters of the speech recognition model can be updated based on the cross entropy loss between the speech recognition result of the speech data and the text label of the speech data, and the connectionist temporal classification loss (CTC) between the second encoding feature and the text label of the speech data.

[0161] Optionally, the parameters of the pre-trained speech recognition model may be updated based only on the loss between the speech recognition result of the speech data and the text label of the speech data.

[0162] Optionally, the loss between the speech recognition result of the speech data and the text label of the speech data may also adopt other losses, such as logarithmic loss, mean square error loss, etc.

[0163] The speech recognition model trained based on the above method has achieved certain improvements in the speech recognition effect for the hearing-impaired. However, the pronunciation habits of people with different levels of hearing impairment are too different. If the fourth training set includes a large number of speech data of people with different levels of hearing impairment, recognition confusion will occur during the training of the speech recognition model, and it will be impossible to learn the pronunciation habits corresponding to each level of hearing impairment. The speech recognition model trained based on the above method will only have a good speech recognition effect for people with a certain level of hearing impairment, but the accuracy of speech recognition for people of other levels is very low.

[0164] The speech recognition method combined with the intelligibility level based on the present application can not only effectively recognize the speech of people with normal hearing, but also effectively recognize the speech of people with different levels of hearing impairment, thereby improving the accuracy and scope of application of speech recognition.

[0165] Corresponding to the method embodiment, the present application also provides a speech recognition device, such as Figure 8 As shown, a structural diagram of a speech recognition device provided in an embodiment of the present application may include:

[0166] A first encoding unit 801, a second encoding unit 802, a fusion unit 803 and a decoding unit 804;

[0167] The first encoding unit 801 is used to perform a first encoding on the speech to be recognized to obtain a first encoding feature; the first encoding feature represents the intelligibility level of the speech to be recognized;

[0168] The second encoding unit 802 is used to perform a second encoding on the speech to be recognized to obtain a second encoding feature; the second encoding feature represents the speech content of the speech to be recognized;

[0169] The fusion unit 803 is used to fuse the first coding feature and the second coding feature to obtain a fused coding feature;

[0170] The decoding unit 804 is used to decode the fused coding features to obtain a speech recognition result.

[0171] The speech recognition device provided in the embodiment of the present application takes into account the intelligibility level of the speech to be recognized when decoding the encoded features of the speech to be recognized, that is, a decoding method that is compatible with the intelligibility level of the speech to be recognized is used for decoding, thereby improving the accuracy of speech recognition.

[0172] In an optional embodiment, when the first encoding unit 801 performs the first encoding on the speech to be recognized, it is used to:

[0173] Extracting acoustic features of the speech to be recognized;

[0174] Encoding the acoustic features through a first network to obtain a first hidden layer feature;

[0175] The first hidden layer feature is encoded through a second network to obtain a second hidden layer feature as the first encoded feature; the complexity of the second network is greater than the complexity of the first network.

[0176] In an optional embodiment, when the second encoding unit 802 performs the second encoding on the speech to be recognized, it is used to:

[0177] Extracting acoustic features of the speech to be recognized;

[0178] Encoding the acoustic features through a third network to obtain a third hidden layer feature;

[0179] The third hidden layer feature is encoded based on the attention mechanism through the third network to obtain the fourth hidden layer feature as the second encoded feature.

[0180] In an optional embodiment, the first encoding unit 801 performs a first encoding on the speech to be recognized, the second encoding unit 802 performs a second encoding on the speech to be recognized, and the decoding unit 804 decodes the fused coding feature through a speech recognition system, and the speech recognition system includes:

[0181] A first encoding module, used for performing a first encoding on the speech to be recognized to obtain a first encoding feature;

[0182] A second encoding module, used for performing a second encoding on the speech to be recognized to obtain a second encoding feature;

[0183] A decoding module, used for decoding the fused coding features to obtain a speech recognition result;

[0184] The speech recognition system is obtained by jointly training a pre-trained intelligibility level recognition model and a pre-trained speech recognition model; the pre-trained intelligibility level recognition model includes the first encoding module; the pre-trained speech recognition model includes the second encoding module and the decoding module.

[0185] In an optional embodiment, the pre-trained intelligibility level recognition model further includes a classification module; the pre-trained intelligibility level recognition model is trained in the following manner:

[0186] Inputting speech data in the first training set into the intelligibility level recognition model, encoding the speech data by the first encoding module to obtain a first encoding feature, and classifying the first encoding feature of the speech data by the classification module to obtain a classification result; the classification result represents the intelligibility level of the speech data; the first training set includes speech data of hearing-impaired people and speech data of normal hearing people;

[0187] Based on the loss between the classification result and the class label of the speech data, the parameters of the intelligibility level recognition model are updated.

[0188] In an optional embodiment, the pre-trained speech recognition model is trained in the following manner:

[0189] Inputting speech data in a second training set into the speech recognition model, encoding the speech data by the second encoding module to obtain a second encoding feature, and decoding the second encoding feature of the speech data by the decoding module to obtain a speech recognition result; the second training set includes speech data of people with normal hearing;

[0190] The parameters of the speech recognition model are updated based on the loss between the speech recognition result of the speech data and the text label of the speech data, and the loss between the second encoded feature and the text label of the speech data.

[0191] In an optional embodiment, the pre-trained intelligibility level recognition model further includes a classification module; the speech recognition system is obtained by performing the following joint training on the pre-trained intelligibility level recognition model and the pre-trained speech recognition model:

[0192] Input the speech data in the third training set into the first encoding module and the second encoding module respectively, obtain the first encoding feature obtained by encoding the speech data by the first encoding module, and obtain the second encoding feature obtained by encoding the speech data by the second encoding module; fuse the first encoding feature and the second encoding feature of the speech data to obtain a fused encoding feature; decode the fused encoding feature by the decoding module to obtain a decoded feature, normalize the decoded feature to obtain a speech recognition result; classify the decoded feature by the classification module to obtain a classification result; the third training set includes speech data of hearing-impaired people and speech data of normal-hearing people;

[0193] Based on the loss between the classification result and the class label of the speech data, and the loss between the speech recognition result of the speech data and the text label of the speech data, the parameters of the decoding module and the classification module are updated.

[0194] The present application also provides an electronic device in an embodiment. Fig. 9 As shown, it shows a structural schematic diagram of an electronic device suitable for implementing the speech recognition method in the embodiment of the present application. The electronic device in the embodiment of the present application can be a terminal device (for example, a car machine, a large-screen device, a smart home, a mobile phone, a tablet computer, a laptop computer, a desktop computer, etc.), or a server (can be a single server, can be a server cluster, or can be a cloud server, etc.). Fig. 9 The electronic device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0195] like Fig. 9As shown, the electronic device may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage device 908 into a random access memory (RAM) 903. When the electronic device is powered on, various programs and data required for the operation of the electronic device are also stored in the RAM 903. The processing device 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0196] Typically, the following devices may be connected to the I / O interface 905: an input device 906 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 907 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 908 including, for example, a memory card, a hard disk, etc.; and a communication device 909. The communication device 909 may allow the electronic device to communicate with other devices wirelessly or by wire to exchange data. Although Fig. 9 An electronic device having various devices is shown, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.

[0197] An embodiment of the present application also provides a computer program product including computer-readable instructions. When the computer-readable instructions are executed on an electronic device, the electronic device implements any of the speech recognition methods provided in the embodiments of the present application.

[0198] A computer-readable storage medium is also provided in an embodiment of the present application. The storage medium carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement any speech recognition method provided in the embodiment of the present application.

[0199] It should be noted that the device embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed over multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, in the drawings of the device embodiments provided by the present application, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines.

[0200] Through the description of the above implementation mode, the technicians in the field can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course, it can also be implemented by special hardware including special integrated circuits, special CPUs, special memories, special components, etc. In general, all functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be various, such as analog circuits, digital circuits or special circuits. However, for the present application, software program implementation is a better implementation mode in more cases. Based on such an understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer floppy disk, a U disk, a mobile hard disk, a ROM, a RAM, a disk or an optical disk, etc., including a number of instructions to enable a computer device (which can be a personal computer, a training device, or a network device, etc.) to execute the methods described in each embodiment of the present application.

[0201] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product. Professionals and technicians may use different methods to implement the described functions for each specific solution, but such implementation should not be considered beyond the scope of this application.

[0202] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from a website site, a computer, a training device, or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, training device, or data center. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device, a data center, etc. that includes one or more available media integrations. The available medium may be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)), etc.

[0203] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0204] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A speech recognition method, characterized in that: include: Performing a first encoding on the speech to be recognized to obtain a first encoding feature; The first coding feature represents the intelligibility level of the speech to be recognized; Performing a second encoding on the speech to be recognized to obtain a second encoding feature; the second encoding feature represents the speech content of the speech to be recognized; Fusing the first coding feature and the second coding feature to obtain a fused coding feature; The fused coding features are decoded to obtain a speech recognition result.

2. The method according to claim 1, characterized in that The first encoding of the speech to be recognized comprises: Extracting acoustic features of the speech to be recognized; Encoding the acoustic features through a first network to obtain a first hidden layer feature; The first hidden layer feature is encoded through a second network to obtain a second hidden layer feature as the first encoded feature; the complexity of the second network is greater than the complexity of the first network.

3. The method according to claim 1, characterized in that The second encoding of the to-be-recognized speech comprises: Extracting acoustic features of the speech to be recognized; Encoding the acoustic features through a third network to obtain a third hidden layer feature; The third hidden layer feature is encoded based on the attention mechanism through the third network to obtain the fourth hidden layer feature as the second encoded feature.

4. The method according to claim 1, characterized in that: The process of performing first encoding and second encoding on the speech to be recognized and decoding the fused coding feature is implemented by a speech recognition system, and the speech recognition system includes: A first encoding module, used for performing a first encoding on the speech to be recognized to obtain a first encoding feature; A second encoding module, used for performing a second encoding on the speech to be recognized to obtain a second encoding feature; A decoding module, used for decoding the fused coding features to obtain a speech recognition result; The speech recognition system is obtained by jointly training a pre-trained intelligibility level recognition model and a pre-trained speech recognition model; the pre-trained intelligibility level recognition model includes the first encoding module; the pre-trained speech recognition model includes the second encoding module and the decoding module.

5. The method according to claim 4, characterized in that The pre-trained intelligibility level recognition model further includes a classification module; the process of pre-training the intelligibility level recognition model includes: Inputting speech data in the first training set into the intelligibility level recognition model, encoding the speech data by the first encoding module to obtain a first encoding feature, and classifying the first encoding feature of the speech data by the classification module to obtain a classification result; the classification result represents the intelligibility level of the speech data; the first training set includes speech data of hearing-impaired people and speech data of normal hearing people; Based on the loss between the classification result and the class label of the speech data, the parameters of the intelligibility level recognition model are updated.

6. The method according to claim 4, characterized in that The process of pre-training the speech recognition model includes: Inputting speech data in a second training set into the speech recognition model, encoding the speech data by the second encoding module to obtain a second encoding feature, and decoding the second encoding feature of the speech data by the decoding module to obtain a speech recognition result; the second training set includes speech data of people with normal hearing; The parameters of the speech recognition model are updated based on the loss between the speech recognition result of the speech data and the text label of the speech data, and the loss between the second encoded feature and the text label of the speech data.

7. The method according to claim 4, characterized in that The pre-trained intelligibility level recognition model further includes a classification module; the process of jointly training the pre-trained intelligibility level recognition model and the pre-trained speech recognition model includes: Input the speech data in the third training set into the first encoding module and the second encoding module respectively, obtain the first encoding feature obtained by encoding the speech data by the first encoding module, and obtain the second encoding feature obtained by encoding the speech data by the second encoding module; fuse the first encoding feature and the second encoding feature of the speech data to obtain a fused encoding feature; decode the fused encoding feature by the decoding module to obtain a decoded feature, normalize the decoded feature to obtain a speech recognition result; classify the decoded feature by the classification module to obtain a classification result; the third training set includes speech data of hearing-impaired people and speech data of normal-hearing people; Based on the loss between the classification result and the class label of the speech data, and the loss between the speech recognition result of the speech data and the text label of the speech data, the parameters of the decoding module and the classification module are updated.

8. A speech recognition device, characterized in that: include: A first encoding unit, used for performing a first encoding on the speech to be recognized to obtain a first encoding feature; The first coding feature represents the intelligibility level of the speech to be recognized; A second encoding unit is used to perform a second encoding on the speech to be recognized to obtain a second encoding feature; the second encoding feature represents the speech content of the speech to be recognized; a fusion unit, configured to fuse the first coding feature and the second coding feature to obtain a fused coding feature; A decoding unit is used to decode the fused coding feature to obtain a speech recognition result.

9. A computer program product, characterized in that It comprises computer-readable instructions, and when the computer-readable instructions are executed on an electronic device, the electronic device implements the speech recognition method as claimed in any one of claims 1 to 7.

10. An electronic device, characterized in that: The electronic device comprises at least one processor and a memory connected to the processor, wherein: The memory is used to store computer programs; The processor is used to execute the computer program so that the electronic device can implement the speech recognition method as described in any one of claims 1 to 7.

11. A computer storage medium, characterized in that: The storage medium carries one or more computer programs, and when the one or more computer programs are executed by an electronic device, the electronic device can implement the speech recognition method as described in any one of claims 1 to 7.