Speech recognition method and device with tone language, computer equipment and medium

By extracting and splicing the clear and turbidity measurement features and acoustic frequency domain features in speech data, and inputting the end-to-end speech detection model, the problem of low accuracy of speech recognition in tone languages ​​is solved, and higher accuracy of speech recognition is achieved.

CN120220651APending Publication Date: 2025-06-27MOBILITY ASIA SMART TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311702402.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-12
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art has low recognition accuracy when performing speech recognition in tone languages, especially in the processing of languages ​​such as Chinese.

Method used

By extracting the clear and turbidity measurement features and acoustic frequency domain features from the target speech data and splicing the two, the splicing features are obtained, and the end-to-end speech detection model is input for speech recognition.

Benefits of technology

The accuracy of speech recognition is improved. By combining clear and turbidity measurement features and acoustic frequency domain features, the features are richer and the characteristics of the pronunciation language itself are integrated, thereby improving the effect of speech recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220651A_ABST
    Figure CN120220651A_ABST
Patent Text Reader

Abstract

The invention relates to a speech recognition method and device with tone languages, computer equipment and a storage medium. The method comprises the following steps: acquiring target voice data, wherein the target voice data is voice data of a language with tones; extracting voiceless and turbid measurement features from the target voice data, wherein the voiceless and turbid measurement features are features used for describing pronunciation types of consonants in the target voice data; extracting acoustic frequency domain features from the target voice data; splicing the clear and turbid measurement features with the acoustic frequency domain features to obtain spliced features; inputting the splicing features into an end-to-end voice detection model to obtain a voice recognition result output by the end-to-end voice detection model; wherein the end-to-end voice detection model carries out model training according to sample splicing characteristics obtained by splicing sample voiced and turbid measurement characteristics of sample voice data and a sample acoustic frequency domain in advance. By adopting the method, the accuracy of speech recognition can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of speech recognition, and particularly to a method, apparatus, computer device, and storage medium for speech recognition of tonal languages. Background Art

[0002] With the development of speech recognition technology, speech recognition is widely applied in the field of artificial intelligence, such as intelligent vehicles, smart phones, and smart speakers. When performing speech recognition processing on audio data, feature extraction is often required first. Regardless of the speech recognition model for which language, the quality of feature extraction is always associated with the quality of model training.

[0003] In traditional technologies, speech recognition feature extraction technologies include MFCC (Mel-scale Frequency Cepstral Coefficients), PLP (Perceptual Linear Predictive), FBANK (Filter Bank), and other technologies. Among them, MFCC feature extraction and PLP feature extraction are usually used in speech recognition systems with a GMM-HMM (Gaussian mixture model-Hidden Markov Model) architecture; while the FBANK feature extraction technology is usually used in speech recognition systems with a DNN-HMM (Deep-Learning Neural Network-Hidden Markov Model) architecture or an end-to-end architecture. Although these feature extraction technologies have achieved certain speech recognition performance, for tonal languages (such as Chinese), there are still problems of low recognition accuracy. Summary of the Invention

[0004] Based on this, it is necessary to provide a method, apparatus, computer device, and storage medium for speech recognition of tonal languages that can improve the accuracy of speech recognition for the above technical problems.

[0005] In a first aspect, a method for speech recognition of tonal languages is provided. The method includes:

[0006] Obtain target speech data, where the target speech data is speech data of a tonal language;

[0007] Extract a voiceless / voiced metric feature from the target speech data, where the voiceless / voiced metric feature is a feature used to describe the pronunciation type of consonants in the target speech data;

[0008] Extract acoustic frequency domain features from the target speech data;

[0009] Concatenate the voicelessness measure features and the acoustic frequency domain features to obtain concatenated features;

[0010] Input the concatenated features into an end-to-end speech detection model to obtain the speech recognition result output by the end-to-end speech detection model; wherein, the end-to-end speech detection model is pre-trained according to the sample concatenated features obtained by concatenating the sample voicelessness measure features and the sample acoustic frequency domain of the sample speech data.

[0011] In some embodiments, extracting the voicelessness measure features from the target speech data includes:

[0012] Preprocess the target speech data to obtain the amplitude spectra corresponding to each speech frame of the target speech data;

[0013] Generate a harmonic spectrum correlation function according to each amplitude spectrum;

[0014] Calculate the voicelessness measure features according to the harmonic spectrum correlation function.

[0015] In some embodiments, preprocessing the target speech data to obtain the amplitude spectra corresponding to each speech frame of the target speech data includes:

[0016] Perform non-linear processing on the target speech data to obtain a non-linear speech signal;

[0017] Perform frame segmentation and windowing on the non-linear speech signal to obtain a plurality of speech frame signals;

[0018] Perform Fourier transform on the speech frame signals to obtain the amplitude spectra corresponding to each speech frame signal.

[0019] In some embodiments, calculating the voicelessness measure features according to the harmonic spectrum correlation function includes:

[0020] Predict the target frequency point when the harmonic spectrum correlation function takes the maximum value according to the argmax function;

[0021] Generate intermediate parameters according to the sampling frequency, the total number of frequency points, and the minimum value of the fundamental frequency, and determine the value range of the frequency points according to the intermediate parameters;

[0022] Generate the voicelessness measure features according to the harmonic spectrum correlation function, the target frequency point, the intermediate parameters, and the value range of the frequency points.

[0023] In some embodiments, generating the harmonic spectrum correlation function according to each amplitude spectrum includes:

[0024] Perform cumulative processing and summation processing on each amplitude spectrum according to the number of harmonics and the length of the frequency window to obtain the harmonic spectrum correlation function.

[0025] In some embodiments, the end-to-end speech detection model includes a Conformer encoder, a CTC decoder, and an attention decoder. The spliced features are input into the end-to-end speech detection model to obtain the speech recognition result output by the end-to-end speech detection model, including:

[0026] Perform causal convolution processing on the input spliced features through the Conformer encoder to obtain encoded spliced features;

[0027] Use the CTC decoder to perform decoding processing on the encoded spliced features based on the prefix beam search method to obtain at least one first recognition result;

[0028] Use the attention decoder to re-score each first recognition result to obtain a second recognition result, and use the second recognition result as the speech recognition result.

[0029] In some embodiments, the method further includes:

[0030] Construct a fusion loss function, which includes a CTC loss function for speech frame-level decoding and an encoder-decoder loss function based on the attention mechanism for label-level decoding;

[0031] Optimize and train the end-to-end speech detection model according to the fusion loss function and the adaptive moment estimation algorithm.

[0032] In a second aspect, a speech recognition device for a tonal language is provided. The device includes:

[0033] A speech data acquisition module for acquiring target speech data, where the target speech data is speech data of a tonal language;

[0034] A voiceless / voiced metric feature extraction module for extracting voiceless / voiced metric features from the target speech data, where the voiceless / voiced metric features are features used to describe the pronunciation type of consonants in the target speech data;

[0035] An acoustic frequency domain feature extraction module for extracting acoustic frequency domain features from the target speech data;

[0036] A feature splicing module for splicing the voiceless / voiced metric features and the acoustic frequency domain features to obtain spliced features;

[0037] A speech recognition module for inputting the spliced features into the end-to-end speech detection model to obtain the speech recognition result output by the end-to-end speech detection model; wherein, the end-to-end speech detection model is pre-trained according to the sample spliced features obtained by splicing the sample voiceless / voiced metric features and the sample acoustic frequency domain of the sample speech data.

[0038] A computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the speech recognition method for tonal languages in any one or more embodiments of the first aspect described above.

[0039] A computer-readable storage medium stores a computer program thereon. When the computer program is executed by a processor, it implements the steps of the speech recognition method for tonal languages in any one or more embodiments of the first aspect described above.

[0040] For the above speech recognition method, device, computer device, and storage medium for tonal languages, by separately extracting the voiceless / voiced metric features and the acoustic frequency domain features from the speech data of the tonal language, and splicing the two features to obtain the spliced features, and applying the spliced features to the end-to-end speech detection model for speech recognition tasks. With this solution, by integrating the voiceless / voiced metric features on the basis of the acoustic frequency domain features, the features are made richer and the characteristics of the spoken language itself are fused, thereby improving the speech recognition accuracy of the end-to-end speech detection model for tonal languages. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 It is an application environment diagram of the speech recognition method for tonal languages in some embodiments;

[0042] Figure 2 It is a flowchart of the speech recognition method for tonal languages in some embodiments;

[0043] Figure 3 It is a flowchart of the steps for extracting the voiceless / voiced metric features from the target speech data in some embodiments;

[0044] Figure 4 It is a flowchart of the Chinese speech recognition method in some application examples;

[0045] Figure 5 It is a structural block diagram of the speech recognition device for tonal languages in some embodiments;

[0046] Figure 6 It is an internal structure diagram of a computer device in some embodiments. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0048] The speech recognition method for tonal languages provided by the present application can be applied, for example, as Figure 1in the application environment shown. Among them, Figure 1 the vehicle 100 shown may include an in-vehicle terminal 110. The in-vehicle terminal 110 may include at least one memory and at least one processor. A computer program is stored in the at least one memory. When the computer program is executed by the at least one processor, a vehicle pose calculation method according to an exemplary embodiment of the present application is executed. Here, the in-vehicle terminal 110 does not have to be a single electronic device, but may also be any aggregate of devices or circuits that can execute the above computer program alone or jointly.

[0049] In the in-vehicle terminal 110, the processor may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. By way of example and not limitation, the processor may also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, etc.

[0050] In the in-vehicle terminal 110, the processor may run the computer program stored in the memory. The computer program may be divided into one or more modules / units (such as computer program 1, computer program 2,...). One or more modules / units are stored in the memory and executed by the processor to complete the method of the embodiment of the present application. One or more modules / units may be a series of computer program instructions capable of completing a specific function, and the instructions may describe the execution process of the computer program in the in-vehicle terminal device. For example, the detail compensation model in the embodiment of the present application may be one of the modules / units.

[0051] The memory may be integrated with the processor. For example, RAM or flash memory may be arranged within an integrated circuit microprocessor, etc. In addition, the memory may include a separate device, such as an external disk drive, a storage array, or any other storage device that can be used by a database system. The memory and the processor may be operatively coupled or may communicate with each other, for example, through an I / O port, a network connection, etc., so that the processor can read the files stored in the memory.

[0052] In addition, the in-vehicle terminal 110 may further include a display device (such as a liquid crystal display) and a user interaction interface (such as a keyboard, a mouse, a touch input device, etc.). All components of the in-vehicle terminal 110 may be connected to each other via a bus and / or a network.

[0053] Specifically, the vehicle-mounted terminal 110 obtains target voice data, where the target voice data is the voice data of a language with tones; extracts the voiceless / voiced metric feature from the target voice data, where the voiceless / voiced metric feature is a feature used to characterize the type of consonants in the target voice data; extracts the acoustic frequency domain feature from the target voice data; splices the voiceless / voiced metric feature and the acoustic frequency domain feature to obtain a spliced feature; inputs the spliced feature into an end-to-end voice detection model to obtain a voice recognition result output by the end-to-end voice detection model; where the end-to-end voice detection model is pre-trained according to the sample spliced feature obtained by splicing the sample voiceless / voiced metric feature and the sample acoustic frequency domain of the sample voice data. Further, other devices associated with the vehicle 100 can be controlled according to the voice recognition result, etc.

[0054] In some other embodiments, the voice recognition of the tonal language provided in this application can also be used in other application scenarios. For example, it can be applied to a computer device. It should be noted that the execution subject can be a configuration device for virtual network card resources, and this device can be implemented as part or all of a computer device through software, hardware, or a combination of software and hardware. Among them, the computer device can be a terminal, a client, or a server. The server can be a single server or a server cluster composed of multiple servers. The terminal can be other intelligent hardware devices such as a vehicle-mounted terminal, a smart speaker device, a smart phone, a personal computer, a tablet computer, a wearable device, and a smart robot, etc.

[0055] In some embodiments, as Figure 2 shown, a voice recognition method for a tonal language is provided. Taking the case where this method is applied to the Figure 1 vehicle-mounted terminal as an example, it includes the following steps:

[0056] Step S202: Obtain target voice data, where the target voice data is the voice data of a language with tones.

[0057] Among them, the target voice data is the voice data to be recognized, which can be voice data such as voice commands issued by the collected user. The language type of the target voice data is a language containing tones. For example, it can be Chinese. For Chinese, tones are of great significance in Chinese pronunciation. Syllables composed of the same initials and finals have completely different meanings with different tones. In Chinese speech, consonants can be divided into two types: voiceless and voiced. Voiceless consonants refer to consonants where the vocal cords do not vibrate during pronunciation. For example, / p / and / t / sounds. Voiced consonants are consonants where the vocal cords vibrate during pronunciation. For example, / b / and / d / sounds. Therefore, in a Chinese speech, the difference in the measurement of voiceless and voiced consonants will affect the tone and also the voice recognition result.

[0058] Specifically, an audio data of a target language with tones emitted by a user can be received through a voice acquisition device as target voice data to be recognized.

[0059] Step S204: Extract a voiceless / voiced metric feature from the target voice data, where the voiceless / voiced metric feature is a feature used to characterize the pronunciation type of consonants in the target voice data.

[0060] Among them, the voiceless / voiced metric feature is used to describe the sound characteristics of consonants, and thus can be used to characterize the pronunciation type of consonants, etc.

[0061] Specifically, the target voice data can be processed to extract sound characteristics that can describe the pronunciation type of consonants from the target voice data, thereby generating a voiceless / voiced metric feature.

[0062] Step S206: Extract an acoustic frequency domain feature from the target voice data.

[0063] Among them, the acoustic frequency domain feature refers to a feature extracted after performing an operation of mapping the audio signal in the voice data from the time domain to the frequency domain. For example, it may include FBANK features, or may also include other acoustic features applicable to an end-to-end speech recognition system.

[0064] Specifically, the vehicle-mounted terminal performs a transformation process on the target voice data from the time domain to the frequency domain, and extracts feature information that can describe the characteristics of the target voice data in the frequency domain as the acoustic frequency domain feature of the target voice data.

[0065] Step S208: Concatenate the voiceless / voiced metric feature and the acoustic frequency domain feature to obtain a concatenated feature.

[0066] Specifically, the voiceless / voiced metric feature extracted in step S204, which is used to describe the pronunciation type of consonants in the target voice data, is concatenated with the acoustic frequency domain feature extracted in step S206, which is used to describe the characteristics of the frequency domain of the target voice data, so that the two are fused to generate a concatenated feature that can more comprehensively describe the acoustic characteristics of the target voice data.

[0067] Step S210: Input the concatenated feature into an end-to-end speech detection model to obtain a speech recognition result output by the end-to-end speech detection model; among them, the end-to-end speech detection model is pre-trained according to a sample concatenated feature obtained by concatenating a sample voiceless / voiced metric feature and a sample acoustic frequency domain of sample voice data.

[0068] Among them, the end-to-end speech detection model can be pre-trained. Historical voice data can be collected to construct sample voice data for training, and a sample concatenated feature for training can be obtained in a similar manner to generating the above-mentioned concatenated feature, and the sample concatenated feature is used as an input for model training.

[0069] Specifically, the concatenated features after feature fusion can be used as input and input into an end-to-end speech detection model. The end-to-end speech detection model processes the input concatenated features, thereby outputting a speech recognition result.

[0070] In the above speech recognition method for tonal languages, the voiceless / voiced metric features and the acoustic frequency domain features are respectively extracted from the target speech data with tones, and the two types of features are concatenated to obtain concatenated features. The concatenated features are applied to an end-to-end speech detection model to perform a speech recognition task. For a specific language, the richer the features used and the more in line with the language characteristics, the stronger the generalization ability of the trained speech recognition model and the higher the accuracy. By adopting this solution, by aggregating the voiceless / voiced metric features on the basis of the acoustic frequency domain features, it is possible to improve the influence of consonants of different pronunciation types on speech data in the end-to-end speech detection model, and fuse the characteristics of the pronunciation language itself, thereby improving the speech recognition accuracy of the end speech detection model for tonal languages.

[0071] In some embodiments, referring to Figure 3 as shown, Figure 3 FIG. shows a schematic flowchart of steps for extracting voiceless / voiced metric features from target speech data in some embodiments. Specifically, it may include the following steps:

[0072] Step S301: Perform preprocessing on the target speech data to obtain the magnitude spectra corresponding to each speech frame of the target speech data;

[0073] Step S302: Generate a harmonic spectrum correlation function according to each magnitude spectrum;

[0074] Step S303: Calculate the voiceless / voiced metric features according to the harmonic spectrum correlation function.

[0075] Since the frequency characteristics of vocal cord vibration are different during the voiceless and voiced sound production processes, it is possible to process the magnitude spectra corresponding to each speech frame of the target speech data to generate corresponding harmonic spectrum correlation functions, and use the harmonic spectrum correlation functions to reflect the frequency characteristics of vocal cord vibration during the sound production process corresponding to different speech frames, thereby improving the recognition accuracy.

[0076] In some embodiments, referring to Figure 3 as shown, step S301 may further include the following steps:

[0077] Step S3011: Perform non-linear processing on the target speech data to obtain a non-linear speech signal.

[0078] Specifically, the square value function can be used to perform non-linear processing on the speech signal xu of the input Chinese speech data to obtain non-linear speech information y i, the formula can be referred to as follows:

[0079] y i = x i 2

[0080] Step S3012: Perform frame segmentation and windowing on the non-linear speech signal to obtain multiple speech frame signals.

[0081] Specifically, the non-linear signal y i can be subjected to speech frame segmentation and a Hamming window function can be added to achieve windowing, thereby obtaining multiple speech frame signals that are frame-segmented and windowed.

[0082] Step S3013: Perform Fourier transform on the speech frame signals to obtain the amplitude spectra corresponding to the respective speech frame signals.

[0083] In some embodiments, calculating the voicing measure feature according to the harmonic spectrum correlation function includes: predicting the target frequency point when the harmonic spectrum correlation function takes the maximum value according to the argmax function; generating an intermediate parameter according to the sampling frequency, the total number of frequency points, and the minimum value of the fundamental frequency, and determining the value range of the frequency point according to the intermediate parameter; generating the voicing measure feature according to the harmonic spectrum correlation function, the target frequency point, the intermediate parameter, and the value range of the frequency point.

[0084] In some embodiments, generating the harmonic spectrum correlation function according to the respective amplitude spectra includes: performing cumulative processing and summing processing on the respective amplitude spectra according to the number of harmonics and the length of the frequency window to obtain the harmonic spectrum correlation function.

[0085] In the above embodiments, exemplarily, the amplitude spectra corresponding to the respective speech frame signals can be input into the following harmonic spectrum correlation function to calculate the voicing measure feature:

[0086]

[0087] where t represents the t-th frame, f represents the frequency point, Y(t, f) represents the amplitude spectrum at the f-th frequency point in the t-th frame, n represents the length of the frequency window, and h represents the number of harmonics.

[0088] Exemplarily, the voicing measure feature extracted based on the harmonic spectrum correlation function can be expressed as follows:

[0089]

[0090]

[0091] where the range of the fundamental frequency is F0_min ≤ f ≤ F0_max, F 0_min = 50HZ, F 0_max = 400HZ. fmax The frequency point position at which the harmonic spectrum correlation function obtains the maximum value. The intermediate parameter here f s represents the sampling frequency, and N represents the total number of frequency points. The value of f ranges from to and excludes f max .

[0092] In some embodiments, the end-to-end speech detection model includes a Conformer encoder, a CTC decoder, and an attention decoder. The spliced features are input into the end-to-end speech detection model, and the speech recognition result output by the end-to-end speech detection model is obtained, including:

[0093] Performing causal convolution processing on the input spliced features through the Conformer encoder to obtain encoded spliced features;

[0094] Using a CTC (Connectionist Temporal Classification) decoder to perform decoding processing on the encoded spliced features based on the prefix beam search method to obtain at least one first recognition result;

[0095] Using the attention decoder to re-score each first recognition result to obtain a second recognition result, and taking the second recognition result as the speech recognition result.

[0096] In the above embodiments, by combining the spliced features with the Conformer encoder, the CTC decoder, and the attention decoder, an end-to-end speech detection model based on the voiceless / voiced metric features can be obtained, thereby improving the accuracy of speech recognition of the target language (especially Chinese).

[0097] In some embodiments, the method further includes: constructing a fusion loss function, where the fusion loss function includes a CTC loss function for speech frame-level decoding and an encoder-decoder loss function based on the attention mechanism for label-level decoding; and optimizing and training the end-to-end speech detection model according to the fusion loss function and the adaptive moment estimation algorithm. In the above embodiments, by constructing the fusion loss function, the model can be optimized during the training process to improve the model accuracy.

[0098] Next, referring to Figure 4 shown, Figure 4 illustrates the Chinese speech recognition method in some application examples, which may specifically include the following steps:

[0099] Step S401, inputting the Chinese speech sampling as the target speech data;

[0100] Step S402, extracting FBANK filter bank features;

[0101] Step S403: Extract the voiceless / voiced metric features;

[0102] Step S404: Construct an end-to-end speech recognition model, including an end-to-end speech recognition model based on a conformer encoder, a CTC decoder, and an Attention decoder;

[0103] Step S405: Concatenate the FBANK filter bank features and the voiceless / voiced metric features, and input them into the trainable end-to-end speech recognition model;

[0104] Step S406: Train the end-to-end speech recognition model based on the fusion loss function and the Adam algorithm;

[0105] Step S407: Perform decoding and recognition based on the CTC prefix beam search algorithm, and merge the same ctc sequence prefixes;

[0106] Step S408: Rescore the n-best results of the CTC decoder through the Attention decoder to obtain the final speech recognition result.

[0107] It should be understood that although Figures 2 to 4 the steps in the flowchart are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figures 2 to 4 at least a part of the steps in

[0108] In some embodiments, as Figure 5 shown, a speech recognition device for a tonal language is provided, including: a speech data acquisition module 502, a voiceless / voiced metric feature extraction module 504, an acoustic frequency domain feature extraction module 506, a feature concatenation module 508, and a speech recognition module 510, where:

[0109] The speech data acquisition module 502 is configured to acquire target speech data, where the target speech data is speech data of a tonal language;

[0110] The voiceless / voiced metric feature extraction module 504 is used to extract voiceless / voiced metric features from the target speech data, where the voiceless / voiced metric features are features used to describe the pronunciation types of consonants in the target speech data;

[0111] The acoustic frequency domain feature extraction module 506 is used to extract acoustic frequency domain features from the target speech data;

[0112] The feature concatenation module 508 is used to concatenate the voiceless / voiced metric features and the acoustic frequency domain features to obtain concatenated features;

[0113] The speech recognition module 510 is used to input the concatenated features into an end-to-end speech detection model to obtain a speech recognition result output by the end-to-end speech detection model; wherein, the end-to-end speech detection model is pre-trained according to sample concatenated features obtained by concatenating sample voiceless / voiced metric features and sample acoustic frequency domains of sample speech data.

[0114] In some embodiments, the voiceless / voiced metric feature extraction module 504 performs preprocessing on the target speech data to obtain amplitude spectra corresponding to respective speech frames of the target speech data; generates a harmonic spectrum correlation function according to the respective amplitude spectra; and calculates voiceless / voiced metric features according to the harmonic spectrum correlation function.

[0115] In some embodiments, the voiceless / voiced metric feature extraction module 504 performs preprocessing on the target speech data to obtain amplitude spectra corresponding to respective speech frames of the target speech data, including: performing non-linear processing on the target speech data to obtain a non-linear speech signal; performing frame division and windowing processing on the non-linear speech signal to obtain a plurality of speech frame signals; and performing Fourier transform on the speech frame signals to obtain amplitude spectra corresponding to the respective speech frame signals.

[0116] In some embodiments, the voiceless / voiced metric feature extraction module 504 predicts a target frequency point when the harmonic spectrum correlation function takes the maximum value according to the argmax function; generates an intermediate parameter according to the sampling frequency, the total number of frequency points, and the minimum value of the fundamental frequency, and determines a value range of the frequency points according to the intermediate parameter; and generates voiceless / voiced metric features according to the harmonic spectrum correlation function, the target frequency point, the intermediate parameter, and the value range of the frequency points.

[0117] In some embodiments, the voiceless / voiced metric feature extraction module 504 performs cumulative processing and summation processing on the respective amplitude spectra according to the number of harmonics and the length of the frequency window to obtain a harmonic spectrum correlation function.

[0118] In some embodiments, the end-to-end speech detection model includes a Conformer encoder, a CTC decoder, and an attention decoder. The speech recognition module 510 performs causal convolution processing on the input concatenated features through the Conformer encoder to obtain the encoded concatenated features; performs decoding processing on the encoded concatenated features using the CTC decoder based on the prefix beam search method to obtain at least one first recognition result; re-scores each first recognition result using the attention decoder to obtain a second recognition result, and uses the second recognition result as the speech recognition result.

[0119] In some embodiments, the speech recognition module 510 is further configured to construct a fusion loss function, where the fusion loss function includes a CTC loss function for speech frame-level decoding and an encoder-decoder loss function based on the attention mechanism for label-level decoding; optimize and train the end-to-end speech detection model according to the fusion loss function and the adaptive moment estimation algorithm.

[0120] For the specific limitations of the speech recognition device for tonal languages, reference can be made to the limitations of the speech recognition method for tonal languages in the foregoing text, which will not be elaborated here. Each module in the above-mentioned speech recognition device for tonal languages can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above-mentioned modules.

[0121] In some embodiments, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 6 shown. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a speech recognition method for tonal languages. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0122] Those skilled in the art can understand, Figure 6The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.

[0123] In some embodiments, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented: obtaining target voice data, where the target voice data is voice data of a language with tones; extracting a voiceless / voiced metric feature from the target voice data, where the voiceless / voiced metric feature is a feature used to describe the pronunciation type of consonants in the target voice data; extracting an acoustic frequency domain feature from the target voice data; splicing the voiceless / voiced metric feature and the acoustic frequency domain feature to obtain a spliced feature; inputting the spliced feature into an end-to-end voice detection model to obtain a voice recognition result output by the end-to-end voice detection model; where the end-to-end voice detection model is pre-trained according to a sample spliced feature obtained by splicing a sample voiceless / voiced metric feature and a sample acoustic frequency domain of sample voice data.

[0124] In some embodiments, when the processor executes the computer program, the following steps are further implemented: preprocessing the target voice data to obtain an amplitude spectrum corresponding to each voice frame of the target voice data; generating a harmonic spectrum correlation function according to each amplitude spectrum; calculating a voiceless / voiced metric feature according to the harmonic spectrum correlation function.

[0125] In some embodiments, when the processor executes the computer program, the following steps are further implemented: performing nonlinear processing on the target voice data to obtain a non-linear voice signal; performing frame division and windowing processing on the non-linear voice signal to obtain a plurality of voice frame signals; performing Fourier transform on the voice frame signals to obtain an amplitude spectrum corresponding to each voice frame signal.

[0126] In some embodiments, when the processor executes the computer program, the following steps are further implemented: predicting a target frequency point when the harmonic spectrum correlation function takes the maximum value according to the argmax function; generating an intermediate parameter according to the sampling frequency, the total number of frequency points, and the minimum value of the fundamental frequency, and determining the value range of the frequency point according to the intermediate parameter; generating a voiceless / voiced metric feature according to the harmonic spectrum correlation function, the target frequency point, the intermediate parameter, and the value range of the frequency point.

[0127] In some embodiments, when the processor executes the computer program, the following steps are further implemented: performing cumulative processing and summation processing on each amplitude spectrum according to the number of harmonics and the length of the frequency window to obtain a harmonic spectrum correlation function.

[0128] In some embodiments, when the processor executes the computer program, the following steps are further implemented: performing causal convolution processing on the input concatenated features through a Conformer encoder to obtain encoded concatenated features; using a CTC decoder to perform decoding processing on the encoded concatenated features based on the prefix beam search method to obtain at least one first recognition result; using an attention decoder to re-score each first recognition result to obtain a second recognition result, and taking the second recognition result as the speech recognition result.

[0129] In some embodiments, when the processor executes the computer program, the following steps are further implemented: constructing a fusion loss function, where the fusion loss function includes a CTC loss function for speech frame-level decoding and an attention mechanism-based encoding and decoding loss function for label-level decoding; optimizing and training the end-to-end speech detection model according to the fusion loss function and the adaptive moment estimation algorithm.

[0130] In some embodiments, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented: obtaining target speech data, where the target speech data is speech data of a language with tones; extracting a voiceless / voiced metric feature from the target speech data, where the voiceless / voiced metric feature is a feature used to describe the pronunciation type of consonants in the target speech data; extracting an acoustic frequency domain feature from the target speech data; concatenating the voiceless / voiced metric feature and the acoustic frequency domain feature to obtain concatenated features; inputting the concatenated features into an end-to-end speech detection model to obtain a speech recognition result output by the end-to-end speech detection model; where the end-to-end speech detection model is pre-trained according to sample concatenated features obtained by concatenating sample voiceless / voiced metric features and sample acoustic frequency domains of sample speech data.

[0131] In some embodiments, when the computer program is executed by a processor, the following steps are further implemented: pre-processing the target speech data to obtain amplitude spectra respectively corresponding to each speech frame of the target speech data; generating a harmonic spectrum correlation function according to each amplitude spectrum; calculating the voiceless / voiced metric feature according to the harmonic spectrum correlation function.

[0132] In some embodiments, when the computer program is executed by a processor, the following steps are further implemented: performing non-linear processing on the target speech data to obtain a non-linear speech signal; performing frame division and windowing processing on the non-linear speech signal to obtain a plurality of speech frame signals; performing Fourier transform on the speech frame signals to obtain amplitude spectra respectively corresponding to each speech frame signal.

[0133] In some embodiments, when the computer program is executed by a processor, the following steps are further implemented: predicting a target frequency point when the harmonic spectrum correlation function takes the maximum value according to the argmax function; generating an intermediate parameter according to the sampling frequency, the total number of frequency points, and the minimum value of the fundamental frequency, and determining the value range of the frequency points according to the intermediate parameter; generating a voicing measure feature according to the harmonic spectrum correlation function, the target frequency point, the intermediate parameter, and the value range of the frequency points.

[0134] In some embodiments, when the computer program is executed by a processor, the following steps are further implemented: accumulating and summing each amplitude spectrum according to the number of harmonics and the length of the frequency window to obtain a harmonic spectrum correlation function.

[0135] In some embodiments, when the computer program is executed by a processor, the following steps are further implemented: performing causal convolution processing on the input concatenated feature through a Conformer encoder to obtain an encoded concatenated feature; using a CTC decoder to perform decoding processing on the encoded concatenated feature based on the prefix beam search method to obtain at least one first recognition result; using an attention decoder to re-score each first recognition result to obtain a second recognition result, and using the second recognition result as the speech recognition result.

[0136] In some embodiments, when the computer program is executed by a processor, the following steps are further implemented: constructing a fusion loss function, where the fusion loss function includes a CTC loss function for speech frame-level decoding and an encoder-decoder loss function based on the attention mechanism for label-level decoding; optimizing and training an end-to-end speech detection model according to the fusion loss function and the adaptive moment estimation algorithm.

[0137] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0138] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0139] The above-described embodiments merely represent several implementation manners of the present application. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. A method for speech recognition of a tonal language, the method comprising: Obtaining target speech data, where the target speech data is speech data of a tonal language; Extracting a voiceless / voiced metric feature from the target speech data, where the voiceless / voiced metric feature is a feature used to describe the pronunciation type of consonants in the target speech data; Extracting an acoustic frequency domain feature from the target speech data; Concatenating the voiceless / voiced metric feature and the acoustic frequency domain feature to obtain a concatenated feature; Inputting the concatenated feature into an end-to-end speech detection model to obtain a speech recognition result output by the end-to-end speech detection model; wherein, the end-to-end speech detection model is pre-trained according to a sample concatenated feature obtained by concatenating a sample voiceless / voiced metric feature and a sample acoustic frequency domain of sample speech data.

2. The method according to claim 1, wherein The extracting the voiceless / voiced metric feature from the target speech data includes: Performing pre-processing on the target speech data to obtain an amplitude spectrum corresponding to each speech frame of the target speech data; Generating a harmonic spectrum correlation function according to each amplitude spectrum; Calculating the voiceless / voiced metric feature according to the harmonic spectrum correlation function.

3. The method according to claim 2, wherein The performing pre-processing on the target speech data to obtain an amplitude spectrum corresponding to each speech frame of the target speech data includes: Performing non-linear processing on the target speech data to obtain a non-linear speech signal; Performing frame division and windowing processing on the non-linear speech signal to obtain a plurality of speech frame signals; Performing Fourier transform on the speech frame signals to obtain an amplitude spectrum corresponding to each speech frame signal.

4. The method according to claim 2, characterized in that, The calculating the voiceless / voiced metric feature according to the harmonic spectrum correlation function includes: Predicting a target frequency point when the harmonic spectrum correlation function takes the maximum value according to the argmax function; Generating an intermediate parameter according to the sampling frequency, the total number of frequency points, and the minimum value of the fundamental frequency, and determining a value range of the frequency point according to the intermediate parameter; Generating the voiceless / voiced metric feature according to the harmonic spectrum correlation function, the target frequency point, the intermediate parameter, and the value range of the frequency point.

5. The method according to claim 2, wherein The generating a harmonic spectrum correlation function according to each amplitude spectrum includes: Performing cumulative processing and summation processing on each amplitude spectrum according to the number of harmonics and the length of the frequency window to obtain the harmonic spectrum correlation function.

6. The method according to claim 1, wherein The end-to-end speech detection model includes a Conformer encoder, a CTC decoder, and an attention decoder. The inputting the concatenated feature into the end-to-end speech detection model to obtain the speech recognition result output by the end-to-end speech detection model includes: Performing causal convolution processing on the input concatenated feature through the Conformer encoder to obtain an encoded concatenated feature; Performing decoding processing on the encoded concatenated feature by using the CTC decoder based on the prefix beam search method to obtain at least one first recognition result; Re-scoring each first recognition result by using the attention decoder to obtain a second recognition result, and taking the second recognition result as the speech recognition result.

7. The method according to claim 6, wherein The method further includes: Construct a fusion loss function, where the fusion loss function includes a CTC loss function for frame-level decoding of speech and an encoder-decoder loss function based on the attention mechanism for label-level decoding; Optimize and train the end-to-end speech detection model according to the fusion loss function and the adaptive moment estimation algorithm.

8. A speech recognition device for a tonal language, characterized in that, The device includes: A speech data acquisition module, configured to acquire target speech data, where the target speech data is speech data of a language with tones; A voiceless / voiced metric feature extraction module, configured to extract voiceless / voiced metric features from the target speech data, where the voiceless / voiced metric features are features used to describe the pronunciation types of consonants in the target speech data; An acoustic frequency domain feature extraction module, configured to extract acoustic frequency domain features from the target speech data; A feature concatenation module, configured to concatenate the voiceless / voiced metric features and the acoustic frequency domain features to obtain concatenated features; A speech recognition module, configured to input the concatenated features into an end-to-end speech detection model to obtain a speech recognition result output by the end-to-end speech detection model; wherein, the end-to-end speech detection model is pre-trained according to sample concatenated features obtained by concatenating sample voiceless / voiced metric features and sample acoustic frequency domains of sample speech data.

9. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 7 are implemented.