Method for identifying speech, method for training a model, and device
By acquiring and analyzing the harmonic characteristics of speech, using a combined model of convolutional neural network and deep neural network, the problem of poor speech recognition accuracy in vehicle-mounted speech wake-up is solved, and more efficient speech recognition and wake-up operations are achieved.
Patent Information
- Application Number
- CN202111505060.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-10
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2041-12-10
AI Technical Summary
The problem of poor speech recognition accuracy in the prior art is that it is difficult to accurately recognize wake-up speech in vehicle voice wake-up scenarios.
By obtaining the harmonic characteristics of the speech to be recognized, and using a speech recognition model composed of a convolutional neural network, a gated cyclic unit and a deep neural network, speech recognition is performed based on harmonic characteristics, including converting time domain speech into frequency domain speech, extracting harmonic characteristics using a Mel filter group, and judging wake-up speech with a probability threshold.
The accuracy of speech recognition is improved, especially in vehicle-mounted voice wake-up scenarios, which can more accurately recognize wake-up voices, reduce the amount of model calculations and improve the accuracy of wake-up.
Smart Images

Figure CN114141246B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, specifically to the fields of artificial intelligence and deep learning technologies. Background Art
[0002] Currently, with the increasingly widespread application of speech recognition, the learning of speech signals becomes particularly important.
[0003] For example, in the scenario of in-vehicle voice wake-up, by learning the features of speech signals, it is possible to identify whether the current speech is the speech that triggers the wake-up condition, thereby achieving accurate in-vehicle voice wake-up. In practice, it is found that the current method of identifying speech has the problem of poor speech recognition accuracy. Summary of the Invention
[0004] The present disclosure provides a method for identifying speech, a method for training a model, and an apparatus.
[0005] According to one aspect of the present disclosure, there is provided a method for identifying speech, including: obtaining the speech to be identified; determining the harmonic features of the speech to be identified; determining the speech recognition result based on the harmonic features and a pre-trained speech recognition model; and outputting the speech recognition result.
[0006] According to another aspect of the present disclosure, there is provided a method for training a model, including: obtaining sample speech and sample annotation data; determining the harmonic features of the sample speech; determining the sample recognition result based on the harmonic features and the model to be trained; and training the model to be trained based on the sample recognition result and the sample annotation data to obtain a speech recognition model.
[0007] According to another aspect of the present disclosure, there is provided an apparatus for identifying speech, including: a speech acquisition unit configured to obtain the speech to be identified; a feature determination unit configured to determine the harmonic features of the speech to be identified; a result determination unit configured to determine the speech recognition result based on the harmonic features and a pre-trained speech recognition model; and a result output unit configured to output the speech recognition result.
[0008] According to another aspect of the present disclosure, there is provided an apparatus for training a model, including: a sample acquisition unit configured to obtain sample speech and sample annotation data; a sample feature determination unit configured to determine the harmonic features of the sample speech; a sample result determination unit configured to determine the sample recognition result based on the harmonic features and the model to be trained; and a model training unit configured to train the model to be trained based on the sample recognition result and the sample annotation data to obtain a speech recognition model.
[0009] According to another aspect of the present disclosure, there is provided an electronic device, including: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement any one of the above methods for recognizing speech or the method for training a model.
[0010] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute any one of the above methods for recognizing speech or the method for training a model.
[0011] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program, which when executed by a processor implements any one of the above methods for recognizing speech or the method for training a model.
[0012] According to the technology of the present disclosure, there is provided a method for recognizing speech, which can improve the accuracy of speech recognition.
[0013] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0015] Figure 1 is an exemplary system architecture diagram to which an embodiment of the present disclosure can be applied;
[0016] Figure 2 is a flowchart of an embodiment of the method for recognizing speech according to the present disclosure;
[0017] Figure 3 is a schematic diagram of an application scenario of the method for recognizing speech according to the present disclosure;
[0018] Figure 4 is a flowchart of another embodiment of the method for recognizing speech according to the present disclosure;
[0019] Figure 5 is a flowchart of an embodiment of the method for training a model according to the present disclosure;
[0020] Figure 6 is a schematic structural diagram of an embodiment of the device for recognizing speech according to the present disclosure;
[0021] Figure 7It is a schematic structural diagram of an embodiment of an apparatus for training a model according to the present disclosure;
[0022] Figure 8 It is a block diagram of an electronic device for implementing the method for recognizing speech or the method for training a model according to an embodiment of the present disclosure. Detailed implementation manners
[0023] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted below for clarity and conciseness.
[0024] It should be noted that, without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other. The present disclosure will be described in detail below with reference to the drawings and in conjunction with the embodiments.
[0025] As Figure 1 shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0026] The terminal devices 101, 102, 103 interact with the server 105 through the network 104 to receive or send messages, etc. Among them, the terminal devices 101, 102, 103 may have a voice wake-up function. The terminal devices 101, 102, 103 may recognize the received voice. If the recognized voice is a wake-up voice, the device is started. Optionally, an application software that can be woken up by voice may be running in the terminal devices 101, 102, 103. If the recognized voice is a wake-up voice, the application software is started to perform corresponding wake-up operations. And after the terminal devices 101, 102, 103 obtain the voice to be recognized, the voice to be recognized may be uploaded to the server 105 through the network 104, so that the server 105 recognizes the voice to be recognized and returns the voice recognition result.
[0027] The terminal devices 101, 102, and 103 can be either hardware or software. When the terminal devices 101, 102, and 103 are hardware, they can be various electronic devices, including but not limited to mobile phones, computers, tablets, in-vehicle devices, and so on. When the terminal devices 101, 102, and 103 are software, they can be installed in the above-listed electronic devices. They can be implemented as multiple software or software modules (for example, to provide distributed services), or they can be implemented as a single software or software module. Specific limitations are not made here.
[0028] The server 105 can be a server that provides various services. For example, the server 105 can obtain the voice to be recognized sent by the terminal devices 101, 102, and 103, determine the harmonic characteristics of the voice to be recognized, input the harmonic characteristics into a pre-trained speech recognition model, obtain the corresponding speech recognition result, and return the speech recognition result to the terminal devices 101, 102, and 103 through the network 104. In addition, during the model training phase, the server 105 can also obtain the sample voice and sample annotation data sent by the terminal devices 101, 102, and 103, determine the harmonic characteristics of the sample voice, determine the sample recognition result based on the harmonic characteristics and the model to be trained, and train the model to be trained based on the sample recognition result and the sample annotation data to obtain the speech recognition model.
[0029] It should be noted that the server 105 can be either hardware or software. When the server 105 is hardware, it can be implemented as a distributed server cluster composed of multiple servers, or it can be implemented as a single server. When the server 105 is software, it can be implemented as multiple software or software modules (for example, to provide distributed services), or it can be implemented as a single software or software module. Specific limitations are not made here.
[0030] It should be noted that the method for recognizing speech and the method for training the model provided by the embodiments of the present disclosure can be executed by the terminal devices 101, 102, and 103, or can be executed by the server 105. The device for generating the recognized speech and the device for training the model can be set in the terminal devices 101, 102, and 103, or can be set in the server 105.
[0031] It should be understood that Figure 1 the numbers of the terminal devices, network, and server in
[0032] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, network, and server. Figure 2 Continuing to refer to
[0033] Step 201: Obtain the speech to be recognized.
[0034] In this embodiment, the execution entity (such as Figure 1 the terminal devices 101, 102, 103 or the server 105 in
[0035] Step 202: Determine the harmonic characteristics of the speech to be recognized.
[0036] In this embodiment, after obtaining the speech to be recognized, the execution entity can analyze the characteristics of the speech to obtain the harmonic characteristics of the speech to be recognized. Among them, the harmonic characteristics are waveform characteristics that can reflect the time dimension and frequency dimension of the speech to be recognized.
[0037] Specifically, for the time-domain speech to be recognized, the execution entity can first convert the time-domain speech to be recognized into a frequency-domain speech to be recognized, and then determine the harmonic characteristics, so as to more fully extract the harmonic characteristics. In addition, the execution entity can also group the frequency-domain speech to be recognized according to the frequency range to obtain several sub-speech groups, and then extract the harmonic characteristics for each sub-speech group, so as to adopt a more fine-grained feature extraction method, which can improve the feature extraction effect. Moreover, the execution entity can also use the Mel filter bank to extract the harmonic characteristics for each sub-speech group. Specifically, the triangular filters in the Mel filter bank are used to convert the frequency range corresponding to each sub-speech group into a frequency range on the Mel scale, and then the harmonic characteristics corresponding to the sub-speech group are extracted in the frequency range on the Mel scale.
[0038] Step 203: Determine the speech recognition result based on the harmonic characteristics and the pre-trained speech recognition model.
[0039] In this embodiment, the execution entity may use the above harmonic features as input data for a pre-trained speech recognition model to obtain a speech recognition result output by the pre-trained speech recognition model. The speech recognition result here may correspond to different speech recognition functions and output corresponding results. For example, when the speech recognition function is a voice wake-up function, the speech recognition result here may be used to indicate whether the speech to be recognized is a wake-up speech. When the speech recognition function is a semantic analysis function, the speech recognition result here may be used to indicate the semantic text content corresponding to the speech to be recognized. When the speech recognition function is a voice verification function, the speech recognition result here may be used to indicate whether the speech to be recognized matches a preset verification speech.
[0040] Preferably, the pre-trained speech recognition model can be trained using a network structure composed of a convolutional neural network, a gated recurrent unit, and a deep neural network. Among them, the convolutional neural network is used to downsample the harmonic features, and the gated recurrent unit is used to further reduce the number of model parameters and the amount of computation, normalize the features, and perform mapping processing on the output to obtain the features to be recognized mapped to a low-dimensional space. The deep neural network is used to recognize the features to be recognized mapped to the low-dimensional space to obtain the final speech recognition result.
[0041] In some alternative implementation manners of this embodiment, determining the speech recognition result based on the harmonic features and the pre-trained speech recognition model may include: determining the probability that the speech to be recognized is a wake-up speech based on the harmonic features and the pre-trained speech recognition model.
[0042] In this implementation manner, if the speech recognition function is a voice wake-up function, the speech recognition result may be probability information used to indicate that the speech to be recognized is a wake-up speech. Among them, the execution entity may pre-store the speech features of the wake-up speech, compare the speech features of the wake-up speech with the harmonic features of the speech to be recognized, and determine the probability information that the speech to be recognized is a wake-up speech based on the similarity between the two. Optionally, the execution entity may also preset a probability threshold. If the speech recognition result indicates that the probability that the speech to be recognized is a wake-up speech is greater than the preset probability threshold, it is determined that the speech to be recognized is a wake-up speech, and the relevant device is further awakened to perform corresponding operations.
[0043] Step 204, output the speech recognition result.
[0044] In this embodiment, after the execution entity determines the speech recognition result, the speech recognition result may be output. If it is in the case of a voice wake-up device, the execution entity may transmit the speech recognition result to the device to be awakened to wake up the device.
[0045] Continue to refer to Figure 3, which shows a schematic diagram of an application scenario of the method for identifying speech according to the present disclosure. In Figure 3 In the application scenario, the execution subject can obtain the speech to be recognized 301, convert the speech to be recognized 301 into a speech signal in the frequency domain, and then group the speech signal in the frequency domain to obtain a plurality of speech group sets 302. Each speech group in the speech group set 302 can correspond to a frequency range. For example, if the frequency range corresponding to the speech signal in the frequency domain is 0 to 8000 Hz, the speech signal in the frequency domain can be divided into 8 groups according to this frequency range, and the frequency range span of each group is 1000 Hz. Then, for each speech group, the harmonic Mel filter bank (Filter Banks, FBank) features corresponding to the speech group can be extracted to obtain a harmonic feature set 303. Specifically, for each speech group, the Mel filter bank can be used to convert the frequency range corresponding to the speech group into a frequency range on the Mel scale, and the harmonic Mel filter bank features can be extracted in the frequency range on the Mel scale. Then, the execution subject can input the harmonic feature set 303 into the speech recognition model 304 to obtain the probability 305 that the speech to be recognized output by the speech recognition model 304 is a preset wake-up speech. If the probability 305 that the speech to be recognized is a preset wake-up speech is greater than a preset probability threshold, the corresponding device can be awakened, such as awakening a vehicle-mounted device.
[0046] The method for identifying speech provided in the above embodiments of the present disclosure can determine a speech recognition result based on the harmonic features of the speech to be recognized and a pre-trained speech recognition model. Since the harmonic features can reflect the speech features of the speech to be recognized in the time dimension and the frequency dimension, determining the final speech recognition result based on the harmonic features can improve the accuracy of speech recognition.
[0047] Continue to refer to Figure 4 , which shows a flow chart 400 of another embodiment of the method for identifying speech according to the present disclosure. As Figure 4 shown, the method for identifying speech in this embodiment may include the following steps:
[0048] Step 401, obtain the speech to be recognized.
[0049] In this embodiment, for a detailed description of step 401, please refer to the detailed description of step 201, which will not be repeated here.
[0050] Step 402, in response to determining that the speech to be recognized is a time-domain speech, determine the frequency-domain speech corresponding to the speech to be recognized.
[0051] In this embodiment, the execution entity can determine the speech signal category of the speech to be recognized. Here, the speech signal category can include time-domain speech and frequency-domain speech. If it is determined that the speech to be recognized is frequency-domain speech, speech conversion processing may not be performed, and the subsequent speech harmonic feature extraction can be directly performed using the frequency-domain speech. If it is determined that the speech to be recognized is time-domain speech, the time-domain speech can be converted into frequency-domain speech, and then the subsequent speech harmonic feature extraction can be performed using the converted frequency-domain speech. By adopting this method, speech harmonic feature extraction can be performed based on the frequency-domain speech, and the feature extraction effect is better.
[0052] Specifically, in response to determining that the speech to be recognized is time-domain speech, the execution entity can perform a fast Fourier transform on the time-domain speech to convert the time-domain speech into frequency-domain speech. Among them, the fast Fourier transform is a commonly used time-domain to frequency-domain transformation method now, and the specific implementation principle will not be elaborated here. And, in this embodiment, other existing various time-domain to frequency-domain transformation methods can also be used to implement the conversion of time-domain speech to frequency-domain speech.
[0053] Step 403: Determine the speech group set corresponding to the frequency-domain speech based on the frequency range of the frequency-domain speech.
[0054] In this embodiment, the execution entity can divide the frequency-domain speech into several speech groups according to the frequency range of the frequency-domain speech to obtain a speech group set. Among them, the frequency range of the frequency-domain speech can be preset, or the corresponding frequency range can be obtained based on the analysis of the frequency of the frequency-domain speech. This embodiment does not limit this. Optionally, the execution entity can preset the number of speech groups to be divided. Based on the frequency range of the frequency-domain speech and the number of speech groups, the corresponding frequency range of each speech group can be determined, and then the frequency-domain speech can be divided into each speech group according to this frequency range to obtain the speech group set corresponding to the frequency-domain speech.
[0055] In some optional implementation manners of this embodiment, determining the speech group set corresponding to the frequency-domain speech based on the frequency range of the frequency-domain speech includes: determining a set of frequency sub-ranges based on the frequency range of the frequency-domain speech; for at least one frequency sub-range in the set of frequency sub-ranges, determining the target frequency sub-range on the Mel scale corresponding to this frequency sub-range; determining the speech group set corresponding to the frequency-domain speech based on at least one target frequency sub-range.
[0056] In this implementation manner, the execution entity can determine each frequency sub-range corresponding to the number of speech groups based on the frequency range of the frequency-domain speech and the preset number of speech groups, that is, determine the set of frequency sub-ranges. Among them, the range spans of the respective frequency sub-ranges in the set of frequency sub-ranges can be the same or different. This embodiment does not limit this.
[0057] For example, if the frequency range of the frequency-domain speech is from 0 to 8000 Hz and the preset number of speech groups is 8 groups, 8 frequency sub-ranges can be determined. If the range spans of each frequency sub-range are the same, the range span of these 8 frequency sub-ranges can be 1000 Hz each.
[0058] After that, for at least one frequency sub-range, the execution entity can use a Mel filter bank to determine the target frequency sub-range on the Mel scale corresponding to this frequency sub-range. Among them, the Mel filter bank can include multiple triangular filters. For the conversion of the Mel scale, it can be calculated based on the conversion formula using the triangular filters. After converting to obtain each target frequency sub-range on the Mel scale, each sub-speech in the frequency-domain speech can be divided into each target frequency sub-range according to the frequency of each sub-speech, and a speech group corresponding to each target frequency sub-range is obtained. The speech group includes multiple sub-speeches.
[0059] Step 404, for at least one speech group in the speech group set, determine the harmonic feature corresponding to this speech group.
[0060] In this embodiment, the execution entity can extract the corresponding FBank feature for each speech group in the speech group set. Among them, since each speech group has been grouped according to the frequency dimension, the FBank feature extracted at this time contains the corresponding harmonic feature.
[0061] For the determination of the harmonic feature, please refer to the detailed description of step 202 for details and will not be elaborated here.
[0062] Step 405, based on the harmonic features corresponding to at least one speech group, determine the harmonic feature of the speech to be recognized.
[0063] In this embodiment, the harmonic feature of the speech to be recognized is the harmonic feature corresponding to each of the above speech groups.
[0064] Step 406, perform downsampling on the harmonic feature based on a convolutional neural network to obtain a downsampled feature.
[0065] In this embodiment, the speech recognition model includes a convolutional neural network, a gated recurrent unit, and a deep neural network. Among them, the convolutional neural network is used to perform downsampling on the feature, the gated recurrent unit is used to map the feature to a low-dimensional space, and the deep neural network is used to recognize the feature in the low-dimensional space to obtain a speech recognition result.
[0066] Specifically, after the execution entity determines the harmonic characteristics, the harmonic characteristics can be input into a convolutional neural network first, so that the convolutional neural network downsamples the harmonic characteristics to reduce the computational amount of the characteristics. After that, the execution entity can obtain the downsampled characteristics after the downsampling process and input the downsampled characteristics into a gated recurrent unit for further processing.
[0067] Step 407: Based on the gated recurrent unit, perform normalization mapping on the downsampled characteristics to obtain mapped characteristics.
[0068] In this embodiment, the gated recurrent unit can perform normalization processing on the downsampled unit and map the normalized characteristics to a low-dimensional space for output to obtain mapped characteristics. Using the gated recurrent unit for normalization mapping can further reduce the model's computational amount. Moreover, the gated recurrent unit is a lightweight structure. Using the gated recurrent unit for mapping processing can effectively reduce the number of model parameters and computational counts.
[0069] Step 408: Based on the deep neural network and the mapped characteristics, determine the speech recognition result.
[0070] In this embodiment, the execution entity can input the mapped characteristics into the deep neural network to obtain the speech recognition result output by the neural network.
[0071] Among them, for the determination of the speech recognition result, please refer to the detailed description in step 203, which will not be elaborated here.
[0072] Step 409: Output the speech recognition result.
[0073] In this embodiment, for the detailed description of step 409, please refer to the detailed description of step 204, which will not be elaborated here.
[0074] Step 410: In response to determining that the probability that the speech recognition result indicates that the speech to be recognized is a wake-up speech is greater than a preset probability threshold, wake up the target device.
[0075] In this embodiment, if the speech recognition function is a speech wake-up function, the speech recognition result can be the probability that the speech to be recognized is a wake-up speech. If this probability is greater than the preset probability threshold, it means that the speech to be recognized is a wake-up speech. At this time, the target device can be woken up. The target device here can be the execution entity or other electronic devices that have been pre-connected to the execution entity. This embodiment does not make any limitations in this regard. For example, if this embodiment is applied to the application scenario of in-vehicle speech wake-up, the target device here can be an in-vehicle device.
[0076] The method for identifying speech provided by the above embodiments of the present disclosure can also convert the time-domain speech into frequency-domain speech, divide the frequency-domain speech into a set of speech groups according to the frequency range of the frequency-domain speech, and use Mel triangular filters to extract FBank features including harmonic features, which can achieve the full extraction of harmonic features. Moreover, a speech recognition model composed of a convolutional neural network, a gated recurrent unit, and a deep neural network can perform speech recognition in a lightweight manner, reducing the model calculation amount. Also, by comparing the probability that the speech recognition result indicates that the speech to be recognized is a wake-up speech with a preset probability threshold, accurate speech wake-up can be achieved, improving the accuracy of speech wake-up.
[0077] Continuing to refer Figure 5 , a flowchart 500 of an embodiment of a method for training a model according to the present disclosure is shown. The method for training a model in this embodiment includes the following steps:
[0078] Step 501, obtain sample speech and sample annotation data.
[0079] In this embodiment, an execution subject (such as Figure 1 the terminal devices 101, 102, 103 or the server 105) can obtain sample speech for training a speech recognition model and sample annotation data for the sample speech. Among them, the sample annotation data can be used to describe that the sample speech is a wake-up speech or the sample speech is not a wake-up speech.
[0080] Step 502, determine the harmonic features of the sample speech.
[0081] In this embodiment, the execution subject can determine the harmonic features of the sample speech and obtain a sample recognition result based on an analysis of the harmonic features of the sample speech.
[0082] In some alternative implementation manners of this embodiment, determining the harmonic features of the sample speech includes: in response to determining that the sample speech is time-domain speech, determining the sample frequency-domain speech corresponding to the sample speech; based on the frequency range of the sample frequency-domain speech, determining the set of sample speech groups corresponding to the sample frequency-domain speech; for at least one sample speech group in the set of sample speech groups, determining the harmonic features corresponding to the sample speech group; based on the harmonic features corresponding to at least one sample speech group, determining the harmonic features of the sample speech.
[0083] In this implementation manner, for the determination method of the harmonic features of the sample speech, please also refer to the determination method of the harmonic features of the speech to be recognized, which will not be elaborated here.
[0084] In some other alternative implementation manners of this embodiment, based on the frequency range of the sample frequency-domain speech, determining a set of sample speech groups corresponding to the sample frequency-domain speech includes: based on the frequency range of the sample frequency-domain speech, determining a set of sample frequency sub-ranges; for at least one sample frequency sub-range in the set of sample frequency sub-ranges, determining a target sample frequency sub-range on the Mel scale corresponding to the sample frequency sub-range; based on at least one target sample frequency sub-range, determining a set of sample speech groups corresponding to the sample frequency-domain speech.
[0085] In this implementation manner, for the determination method of the set of sample speech groups, please refer to the determination method of the set of speech groups together, and details are not described herein again.
[0086] Step 503, determining a sample recognition result based on the harmonic feature and the model to be trained.
[0087] In this embodiment, inputting the harmonic feature into the model to be trained can obtain the sample recognition result output by the model to be trained.
[0088] In some alternative implementation manners of this embodiment, the model to be trained includes a convolutional neural network, a gated recurrent unit, and a deep neural network; and determining a sample recognition result based on the harmonic feature and the model to be trained includes: performing downsampling on the harmonic feature based on the convolutional neural network to obtain a downsampled feature; performing normalized mapping on the downsampled feature based on the gated recurrent unit to obtain a mapped feature; determining a sample recognition result based on the deep neural network and the mapped feature.
[0089] In this implementation manner, for the determination method of the sample recognition result, please refer to the determination method of the speech recognition result together, and details are not described herein again.
[0090] In some other alternative implementation manners of this embodiment, the sample annotation data includes wake-up speech and non-wake-up speech; and determining a sample recognition result based on the harmonic feature and the model to be trained includes: determining the probability that the sample speech is wake-up speech based on the harmonic feature and the model to be trained.
[0091] In this implementation manner, for the speech recognition function being the speech wake-up function, the sample annotation data may include wake-up speech and non-wake-up speech, and the sample recognition result indicates the probability that the sample speech is wake-up speech. For the speech recognition function being the semantic analysis function, the sample annotation data may be the sample text content. For the speech recognition function being the speech verification function, the sample annotation data may include verification speech and non-verification speech.
[0092] Step 504, training the model to be trained based on the sample recognition result and the sample annotation data to obtain a speech recognition model.
[0093] In this embodiment, the execution entity may continuously adjust the parameters of the model to be trained based on the difference between the sample recognition result and the sample annotation data until the model converges, thereby obtaining a speech recognition model.
[0094] The method for training a model provided in the above embodiments of the present disclosure can improve the accuracy of speech recognition.
[0095] Further referring to Figure 6 , as an implementation of the methods shown in the above respective figures, the present disclosure provides an embodiment of a device for recognizing speech. This device embodiment corresponds to Figure 2 the method embodiment shown, and this device can be specifically applied to a terminal device or a server.
[0096] As Figure 6 shown, the speech recognition device 600 in this embodiment includes: a speech acquisition unit 601, a feature determination unit 602, a result determination unit 603, and a result output unit 604.
[0097] The speech acquisition unit 601 is configured to acquire the speech to be recognized.
[0098] The feature determination unit 602 is configured to determine the harmonic features of the speech to be recognized.
[0099] The result determination unit 603 is configured to determine the speech recognition result based on the harmonic features and a pre-trained speech recognition model.
[0100] The result output unit 604 is configured to output the speech recognition result.
[0101] In some optional implementation manners of this embodiment, the feature determination unit 602 is further configured to: in response to determining that the speech to be recognized is a time-domain speech, determine the corresponding frequency-domain speech of the speech to be recognized; based on the frequency range of the frequency-domain speech, determine the set of speech groups corresponding to the frequency-domain speech; for at least one speech group in the set of speech groups, determine the harmonic features corresponding to the speech group; based on the harmonic features corresponding to at least one speech group, determine the harmonic features of the speech to be recognized.
[0102] In some optional implementation manners of this embodiment, the feature determination unit 602 is further configured to: based on the frequency range of the frequency-domain speech, determine the set of frequency sub-ranges; for at least one frequency sub-range in the set of frequency sub-ranges, determine the target frequency sub-range on the Mel scale corresponding to the frequency sub-range; based on at least one target frequency sub-range, determine the set of speech groups corresponding to the frequency-domain speech.
[0103] In some alternative implementation manners of this embodiment, the pre-trained speech recognition model includes a convolutional neural network, a gated recurrent unit, and a deep neural network; and, the result determination unit 603 is further configured to: perform downsampling on the harmonic features based on the convolutional neural network to obtain downsampled features; perform normalization mapping on the downsampled features based on the gated recurrent unit to obtain mapped features; and determine the speech recognition result based on the deep neural network and the mapped features.
[0104] In some alternative implementation manners of this embodiment, the result determination unit 603 is further configured to: determine the probability that the speech to be recognized is a wake-up speech based on the harmonic features and the pre-trained speech recognition model.
[0105] In some alternative implementation manners of this embodiment, it further includes: a device wake-up unit, configured to wake up the target device in response to determining that the probability that the speech to be recognized is a wake-up speech is greater than a preset probability threshold.
[0106] It should be understood that the units 601 to 604 described in the apparatus 600 for recognizing speech respectively correspond to the respective steps in the method described in the reference Figure 2 Therefore, the operations and features described above for the method for recognizing speech also apply to the apparatus 600 and the units included therein, and will not be elaborated herein.
[0107] Further referring to Figure 7 , as an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of an apparatus for training a model. This apparatus embodiment corresponds to the method embodiment shown in Figure 5 and this apparatus can be specifically applied to a terminal device or a server.
[0108] As Figure 7 shown, the apparatus 700 for training a model in this embodiment includes: a sample acquisition unit 701, a sample feature determination unit 702, a sample result determination unit 703, and a model training unit 704.
[0109] The sample acquisition unit 701 is configured to acquire sample speech and sample annotation data.
[0110] The sample feature determination unit 702 is configured to determine the harmonic features of the sample speech.
[0111] The sample result determination unit 703 is configured to determine a sample recognition result based on the harmonic features and the model to be trained.
[0112] The model training unit 704 is configured to train the model to be trained based on the sample recognition result and the sample annotation data to obtain a speech recognition model.
[0113] In some alternative implementation manners of this embodiment, the sample feature determination unit 702 is further configured to: in response to determining that the sample voice is a time-domain voice, determine the sample frequency-domain voice corresponding to the sample voice; based on the frequency range of the sample frequency-domain voice, determine the set of sample voice groups corresponding to the sample frequency-domain voice; for at least one sample voice group in the set of sample voice groups, determine the harmonic feature corresponding to the sample voice group; and based on the harmonic features corresponding to the at least one sample voice group, determine the harmonic feature of the sample voice.
[0114] In some alternative implementation manners of this embodiment, the sample feature determination unit 702 is further configured to: based on the frequency range of the sample frequency-domain voice, determine the set of sample frequency sub-ranges; for at least one sample frequency sub-range in the set of sample frequency sub-ranges, determine the target sample frequency sub-range on the Mel scale corresponding to the sample frequency sub-range; and based on the at least one target sample frequency sub-range, determine the set of sample voice groups corresponding to the sample frequency-domain voice.
[0115] In some alternative implementation manners of this embodiment, the model to be trained includes a convolutional neural network, a gated recurrent unit, and a deep neural network; and the sample result determination unit 703 is further configured to: perform downsampling on the harmonic feature based on the convolutional neural network to obtain a downsampled feature; perform normalized mapping on the downsampled feature based on the gated recurrent unit to obtain a mapped feature; and based on the deep neural network and the mapped feature, determine the sample recognition result.
[0116] In some alternative implementation manners of this embodiment, the sample annotation data includes wake-up voices and non-wake-up voices; and the sample result determination unit 703 is further configured to: based on the harmonic feature and the model to be trained, determine the probability that the sample voice is a wake-up voice.
[0117] It should be understood that the units 701 to 704 described in the apparatus 700 for training the model respectively correspond to the respective steps in the method described in the reference Figure 5 Therefore, the operations and features described above for the method for training the model also apply to the apparatus 700 and the units included therein, and will not be elaborated herein.
[0118] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0119] Figure 8FIG. shows a schematic block diagram of an exemplary electronic device 800 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0120] As Figure 8 shown, the device 800 includes a computing unit 801 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0121] A plurality of components in the device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0122] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above, such as the method for recognizing speech or the method for training a model. For example, in some embodiments, the method for recognizing speech or the method for training a model can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the method for recognizing speech or the method for training a model described above can be executed. Alternatively, in other embodiments, the computing unit 801 can be configured to execute the method for recognizing speech or the method for training a model in any other suitable way (e.g., by means of firmware).
[0123] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0124] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0125] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0126] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).
[0127] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of a communication network include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0128] A computer system may include a client and a server. The client and the server are generally far away from each other and usually interact via a communication network. The relationship between the client and the server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server incorporating a blockchain.
[0129] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0130] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A method for speech recognition, comprising: Obtaining the speech to be recognized; In response to determining that the speech to be recognized is a time-domain speech, determining the frequency-domain speech corresponding to the speech to be recognized; Based on the frequency range of the frequency-domain speech, determining a set of frequency sub-ranges; for at least one frequency sub-range in the set of frequency sub-ranges, determining the target frequency sub-range on the Mel scale corresponding to this frequency sub-range; based on at least one target frequency sub-range, determining a set of speech groups corresponding to the frequency-domain speech; for at least one speech group in the set of speech groups, determining the harmonic features corresponding to this speech group; based on the harmonic features corresponding to at least one speech group, determining the harmonic features of the speech to be recognized; Based on the harmonic features and a pre-trained speech recognition model, determining a speech recognition result; Outputting the speech recognition result.
2. The method according to claim 1, wherein The pre-trained speech recognition model includes a convolutional neural network, a gated recurrent unit, and a deep neural network; And The determining a speech recognition result based on the harmonic features and a pre-trained speech recognition model includes: Performing downsampling on the harmonic features based on the convolutional neural network to obtain downsampled features; Performing normalized mapping on the downsampled features based on the gated recurrent unit to obtain mapped features; Based on the deep neural network and the mapped features, determining the speech recognition result.
3. The method according to claim 1, wherein, The determining a speech recognition result based on the harmonic features and a pre-trained speech recognition model includes: Based on the harmonic features and the pre-trained speech recognition model, determining the probability that the speech to be recognized is a wake-up speech.
4. The method according to claim 3, wherein Further comprising: In response to determining that the probability that the speech to be recognized is a wake-up speech is greater than a preset probability threshold, waking up the target device.
5. A method for training a model, comprising: Obtaining sample speech and sample annotation data; In response to determining that the sample speech is a time-domain speech, determining the sample frequency-domain speech corresponding to the sample speech; Based on the frequency range of the sample frequency-domain speech, determining a set of sample frequency sub-ranges; for at least one sample frequency sub-range in the set of sample frequency sub-ranges, determining the target sample frequency sub-range on the Mel scale corresponding to this sample frequency sub-range; based on at least one target sample frequency sub-range, determining a set of sample speech groups corresponding to the sample frequency-domain speech; for at least one sample speech group in the set of sample speech groups, determining the harmonic features corresponding to this sample speech group; Based on the harmonic features corresponding to at least one sample speech group, determining the harmonic features of the sample speech; Based on the harmonic features and a model to be trained, determining a sample recognition result; Based on the sample recognition result and the sample annotation data, training the model to be trained to obtain a speech recognition model.
6. The method according to claim 5, wherein, The model to be trained includes a convolutional neural network, a gated recurrent unit, and a deep neural network; And The determining a sample recognition result based on the harmonic features and a model to be trained includes: Performing downsampling on the harmonic features based on the convolutional neural network to obtain downsampled features; Normalize and map the downsampled features based on the gated recurrent unit to obtain mapped features; Determine the sample recognition result based on the deep neural network and the mapped features.
7. The method according to claim 5, wherein The sample annotation data includes wake-up speech and non-wake-up speech; and Determining the sample recognition result based on the harmonic features and the model to be trained includes: Determine the probability that the sample speech is wake-up speech based on the harmonic features and the model to be trained.
8. A device for speech recognition, comprising: A speech acquisition unit configured to acquire speech to be recognized; A feature determination unit configured to: in response to determining that the speech to be recognized is time-domain speech, determine the frequency-domain speech corresponding to the speech to be recognized; based on the frequency range of the frequency-domain speech, determine a set of frequency sub-ranges; for at least one frequency sub-range in the set of frequency sub-ranges, determine the target frequency sub-range on the Mel scale corresponding to this frequency sub-range; based on at least one target frequency sub-range, determine a set of speech groups corresponding to the frequency-domain speech; for at least one speech group in the set of speech groups, determine the harmonic features corresponding to this speech group; Determine the harmonic features of the speech to be recognized based on the harmonic features corresponding to at least one speech group; A result determination unit configured to determine a speech recognition result based on the harmonic features and a pre-trained speech recognition model; A result output unit configured to output the speech recognition result.
9. The apparatus according to claim 8, wherein, The pre-trained speech recognition model includes a convolutional neural network, a gated recurrent unit, and a deep neural network; and The result determination unit is further configured to: Downsample the harmonic features based on the convolutional neural network to obtain downsampled features; Normalize and map the downsampled features based on the gated recurrent unit to obtain mapped features; Determine the speech recognition result based on the deep neural network and the mapped features.
10. The apparatus according to claim 8, wherein, The result determination unit is further configured to: Determine the probability that the speech to be recognized is wake-up speech based on the harmonic features and the pre-trained speech recognition model.
11. The apparatus according to claim 10, wherein Further comprising: A device wake-up unit configured to wake up the target device in response to determining that the probability that the speech to be recognized is wake-up speech is greater than a preset probability threshold.
12. A device for training a model, comprising: A sample acquisition unit configured to acquire sample speech and sample annotation data; A sample feature determination unit configured to: in response to determining that the sample speech is time-domain speech, determine the sample frequency-domain speech corresponding to the sample speech; based on the frequency range of the sample frequency-domain speech, determine a set of sample frequency sub-ranges; for at least one sample frequency sub-range in the set of sample frequency sub-ranges, determine the target sample frequency sub-range on the Mel scale corresponding to this sample frequency sub-range; based on at least one target sample frequency sub-range, determine a set of sample speech groups corresponding to the sample frequency-domain speech; for at least one sample speech group in the set of sample speech groups, determine the harmonic features corresponding to this sample speech group; Determine the harmonic feature of the sample speech based on the harmonic features corresponding to at least one sample speech group; A sample result determination unit, configured to determine a sample recognition result based on the harmonic feature and the model to be trained; A model training unit, configured to train the model to be trained based on the sample recognition result and the sample annotation data to obtain a speech recognition model.
13. The device according to claim 12, wherein, The model to be trained includes a convolutional neural network, a gated recurrent unit, and a deep neural network; and The sample result determination unit is further configured to: Perform downsampling on the harmonic feature based on the convolutional neural network to obtain a downsampled feature; Perform normalized mapping on the downsampled feature based on the gated recurrent unit to obtain a mapped feature; Determine the sample recognition result based on the deep neural network and the mapped feature.
14. The apparatus according to claim 13, wherein The sample annotation data includes wake-up speech and non-wake-up speech; and The sample result determination unit is further configured to: Determine the probability that the sample speech is wake-up speech based on the harmonic feature and the model to be trained.
15. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-7.
16. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-7.
17. A computer program product, comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-7.
Citation Information
Patent Citations
Acoustic interval detection method and device
US20060053003A1
Pitch Dependent Speech Recognition Engine
US20080167862A1