Speech recognition method and electronic equipment

By adjusting the sampling rate to determine the target sampling parameters, second audio data is obtained for speech recognition, which solves the problem of insufficient speech recognition accuracy in the existing technology and improves the recognition accuracy.

CN121789651APending Publication Date: 2026-04-03LENOVO (BEIJING) LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-26
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

In existing speech recognition technologies, the accuracy of speech recognition results is poor, and there is still a significant gap even after post-processing and error correction.

Method used

By adjusting the sampling rate, the target sampling parameters are determined, and second audio data is obtained based on these parameters for speech recognition. The second audio data is then matched with the first audio data to improve recognition accuracy.

Benefits of technology

This improved the accuracy of speech recognition results and reduced the occurrence of recognition errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121789651A_ABST
    Figure CN121789651A_ABST
Patent Text Reader

Abstract

The invention discloses a voice recognition method and electronic equipment. The method comprises the steps of obtaining first audio data based on a first sampling rate; a target sampling parameter is determined based on the first audio data, the sampling rate corresponding to the target sampling parameter is different from the first sampling rate, and the second accuracy of voice recognition of the audio data obtained based on the sampling rate corresponding to the target sampling parameter is higher than the first accuracy; the first accuracy is the accuracy of voice recognition based on the audio data obtained based on the first sampling rate; obtaining second audio data based on the target sampling parameter, wherein the second audio data is matched with the first audio data; and performing voice recognition on the second audio data to obtain a voice recognition result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a speech recognition method and electronic device. Background Technology

[0002] When performing speech recognition on audio, the accuracy of the recognition results is poor. Currently, the accuracy of speech recognition results is usually improved through post-processing and error correction after speech recognition. For example, by combining contextual semantics and statistical rules, the number of speech recognition errors can be reduced. However, the speech recognition results obtained by using this method still have a large discrepancy with reality. Summary of the Invention

[0003] In view of this, this application provides a speech recognition method and an electronic device, the specific solution of which is as follows:

[0004] A speech recognition method, comprising:

[0005] The first audio data is obtained based on the first sampling rate;

[0006] Target sampling parameters are determined based on the first audio data. The sampling rate corresponding to the target sampling parameters is different from the first sampling rate. The second accuracy of speech recognition based on the audio data obtained based on the sampling rate corresponding to the target sampling parameters is higher than the first accuracy. The first accuracy is the accuracy of speech recognition based on the audio data obtained based on the first sampling rate.

[0007] Second audio data is obtained based on the target sampling parameters, and the second audio data is matched with the first audio data;

[0008] The second audio data is subjected to speech recognition to obtain the speech recognition result.

[0009] Furthermore, determining the target sampling parameters based on the first audio data includes:

[0010] The first audio data is sampled and transformed according to multiple different sampling transformation ratios to obtain multiple third audio data.

[0011] Target audio data is determined from the plurality of third audio data, wherein the speech recognition result of the target audio data satisfies the speech recognition condition;

[0012] The target sampling transformation ratio corresponding to the target audio data is determined as the target sampling parameter.

[0013] Furthermore, the second audio data and the first audio data are audio data of the same audio segment but with different sampling rates. Obtaining the second audio data based on the target sampling parameters includes:

[0014] The target audio data is identified as the second audio data.

[0015] Furthermore, the target sampling parameter characterizes the sampling transformation ratio, and obtaining the second audio data based on the target sampling parameter includes:

[0016] The initial audio data is sampled and transformed according to the target sampling parameters to obtain the second audio data after sampling transformation. The sampling rate of the initial audio data is the same as that of the first audio data, and the initial audio data and the first audio data are different parts of the same audio segment.

[0017] Furthermore, obtaining the second audio data based on the target sampling parameters includes:

[0018] The target sampling rate is determined based on the first sampling rate and the target sampling parameters;

[0019] The initial audio data is resampled according to the target sampling rate to obtain second audio data that matches the target sampling rate; the sampling rate of the initial audio data is the same as that of the first audio data, and the initial audio data and the first audio data are different parts of the same audio segment.

[0020] Furthermore, determining the target sampling parameters based on the first audio data includes:

[0021] The first audio data is resampled at multiple different sampling rates to obtain multiple third audio data.

[0022] Target audio data is determined from the plurality of third audio data, wherein the speech recognition result of the target audio data satisfies the speech recognition condition;

[0023] The sampling rate corresponding to the target audio data is determined as the target sampling parameter.

[0024] Furthermore, the target audio data is determined from multiple third-party audio data sources, including:

[0025] Speech recognition is performed on each third audio data separately to obtain the speech recognition result for each audio data;

[0026] The speech recognition results of multiple third-party audio data are compared to obtain the comparison results;

[0027] The target audio data is determined from the plurality of third audio data based on the comparison results.

[0028] Furthermore, the comparison of speech recognition results from multiple third audio data sets to obtain comparison results includes:

[0029] Determine the perplexity of each speech recognition result among the multiple third audio data;

[0030] The perplexity of the speech recognition results of the multiple third audio data is compared to obtain the comparison result.

[0031] Furthermore, determining the target sampling parameters based on the first audio data includes:

[0032] In response to obtaining target information, target sampling parameters are determined based on the first audio data, wherein the target information indicates that the first audio data is audio played at a variable speed.

[0033] An electronic device, comprising:

[0034] Audio acquisition device, used to acquire audio data;

[0035] The processor is configured to control the audio acquisition device to acquire first audio data based on a first sampling rate; determine a target sampling parameter based on the first audio data, wherein the sampling rate corresponding to the target sampling parameter is different from the first sampling rate, and the second accuracy of speech recognition based on the audio data acquired based on the sampling rate corresponding to the target sampling parameter is higher than the first accuracy, wherein the first accuracy is the accuracy of speech recognition based on the audio data acquired based on the first sampling rate; acquire second audio data based on the target sampling parameter, wherein the second audio data matches the first audio data; and perform speech recognition on the second audio data to obtain a speech recognition result. Attached Figure Description

[0036] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0037] Figure 1 This is a flowchart of a speech recognition method disclosed in an embodiment of this application;

[0038] Figure 2 This is a flowchart of a speech recognition method disclosed in an embodiment of this application;

[0039] Figure 3 This is a flowchart of a speech recognition method disclosed in an embodiment of this application;

[0040] Figure 4 This is a flowchart of a speech recognition method disclosed in an embodiment of this application;

[0041] Figure 5 This is a flowchart of a speech recognition method disclosed in an embodiment of this application;

[0042] Figure 6 This is a schematic diagram of the structure of an electronic device disclosed in an embodiment of this application;

[0043] Figure 7 This is a schematic diagram of the structure of a speech recognition system disclosed in an embodiment of this application. Detailed Implementation

[0044] The embodiments of this application are described below with reference to the accompanying drawings. The terminology used in the implementation section of this application is for explaining specific embodiments only and is not intended to limit the scope of this application.

[0045] The embodiments of this application will now be described with reference to the accompanying drawings. Those skilled in the art will recognize that, with technological advancements and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are equally applicable to similar technical problems.

[0046] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of elements is not necessarily limited to those elements, but may include other elements not explicitly listed or inherent to those processes, methods, products, or apparatuses.

[0047] This application discloses a speech recognition method, the flowchart of which is shown below. Figure 1 As shown, it includes:

[0048] Step S11: Obtain first audio data based on the first sampling rate;

[0049] Step S12: Determine the target sampling parameters based on the first audio data. The sampling rate corresponding to the target sampling parameters is different from the first sampling rate. The second accuracy of speech recognition based on the audio data obtained based on the sampling rate corresponding to the target sampling parameters is higher than the first accuracy. The first accuracy is the accuracy of speech recognition based on the audio data obtained based on the first sampling rate.

[0050] Step S13: Obtain the second audio data based on the target sampling parameters, and match the second audio data with the first audio data;

[0051] Step S14: Perform speech recognition on the second audio data to obtain the speech recognition result.

[0052] When performing speech recognition on audio, the accuracy of the recognition results is poor. Currently, the accuracy of speech recognition results is usually improved through post-processing and error correction after speech recognition. For example, by combining contextual semantics and statistical rules, the number of speech recognition errors can be reduced. However, the speech recognition results obtained by using this method still have a large discrepancy with reality.

[0053] Based on this, in this embodiment, a target sampling parameter is determined based on the first audio data. The sampling rate corresponding to the target sampling parameter is different from the sampling rate of the first audio data. Second audio data is obtained based on the target sampling parameter, and speech recognition is performed on the second audio data. This realizes the adjustment of the sampling parameter of the obtained second audio data based on the first audio data, so that the accuracy of speech recognition of the adjusted second audio data is higher than that of speech recognition of the first audio data. That is, the accuracy of speech recognition of the adjusted audio data is improved, and the occurrence of speech recognition errors is reduced.

[0054] The speech recognition method disclosed in this embodiment is applied to an electronic device. The electronic device determines target sampling parameters based on the obtained first audio data, and obtains second speech data based on the target sampling parameters. The electronic device performs speech recognition on the second audio data to obtain a speech recognition result.

[0055] The first audio data is obtained, which can be acquired by an electronic device using a first sampling rate. Alternatively, the first audio data can be transmitted to the electronic device from other devices. Regardless of the device transmitting the data, it is acquired using the first sampling rate. Furthermore, the electronic device can acquire the first audio data either during the execution of the speech recognition method disclosed in this embodiment or by pre-acquiring the pre-acquired first audio data during the execution of the speech recognition method disclosed in this embodiment.

[0056] After obtaining the first audio data, a target sampling parameter is determined based on the first audio data. That is, the first audio data is processed to obtain the target sampling parameter. The speech recognition result of the audio data corresponding to the target sampling parameter satisfies the speech recognition condition. Satisfying the speech recognition condition means that the recognition accuracy during speech recognition reaches a certain threshold. Only when the recognition accuracy of the speech recognition result of the audio data corresponding to a certain sampling parameter can reach the specific threshold is the sampling parameter determined as the target sampling parameter. The second audio data is then obtained based on the target sampling parameter so that speech recognition can be performed on the second audio data.

[0057] Wherein, the sampling rate corresponding to the target sampling parameter is different from the first sampling rate. If the sampling rate corresponding to the target sampling parameter is the second sampling rate, the accuracy of speech recognition based on the audio data obtained based on the second sampling rate is the second accuracy, and the accuracy of speech recognition based on the audio data obtained based on the first sampling rate is the first accuracy. In this case, the second accuracy is higher than the first accuracy.

[0058] After determining the target sampling parameters, the second audio data is obtained based on the target sampling parameters, and speech recognition is performed on the second audio data to obtain the speech recognition result. The accuracy of the speech recognition result is higher than that of the speech recognition of the first audio data.

[0059] Specifically, the matching of the second audio data with the first audio data can be achieved by the following: the second audio data and the first audio data are audio data of the same audio segment at different sampling rates, that is, the first audio data and the second audio data are obtained by collecting the same audio segment at different sampling rates; or, the second audio data and the first audio data can also be different parts of the same audio segment, that is, the same audio segment includes multiple different parts, the first audio data and the second audio data correspond to different parts of the audio segment, and the sampling rates corresponding to the first audio segment and the second audio segment are also different.

[0060] That is, the first audio data is obtained based on the first sampling rate. At this time, the second audio data is determined to be obtained according to the second sampling rate based on the first audio data. The accuracy of speech recognition of the second audio data is higher than that of speech recognition of the first audio data. Therefore, speech recognition is performed on the second audio data instead of the first audio data, so as to improve the accuracy of audio data recognition and reduce the occurrence of recognition errors.

[0061] The speech recognition method disclosed in this embodiment, after obtaining first audio data based on a first sampling rate, determines target sampling parameters based on the first audio data. The sampling rate corresponding to the target sampling parameters is different from the first sampling rate. The second accuracy of speech recognition based on the audio data obtained based on the sampling rate corresponding to the target sampling parameters is higher than the first accuracy, where the first accuracy is the accuracy of speech recognition based on the audio data obtained based on the first sampling rate. Then, second audio data is obtained based on the target sampling parameters, matched with the first audio data, and speech recognition is performed on the second audio data to obtain a speech recognition result. This scheme, after determining the target sampling parameters based on the first audio data, obtains the second audio data based on the target sampling parameters and performs speech recognition on the second audio data. Since the second accuracy of speech recognition based on the audio data obtained based on the sampling rate corresponding to the target sampling parameters is higher than the accuracy of speech recognition based on the audio data obtained based on the first sampling rate, the accuracy of speech recognition on the second audio data is higher than the accuracy of speech recognition on the first audio data. This achieves the goal of improving the accuracy of speech recognition of adjusted audio data by adjusting the target sampling parameters, effectively reducing speech recognition errors.

[0062] This embodiment discloses a speech recognition method, the flowchart of which is as follows: Figure 2 As shown, it includes:

[0063] Step S21: Obtain first audio data based on the first sampling rate;

[0064] Step S22: Perform sampling transformation on the first audio data according to multiple different sampling transformation ratios to obtain multiple third audio data;

[0065] Step S23: Determine the target audio data from multiple third-party audio data, and the speech recognition result of the target audio data satisfies the speech recognition conditions;

[0066] Step S24: Determine the target sampling transformation ratio corresponding to the target audio data as the target sampling parameter. The sampling rate corresponding to the target sampling parameter is different from the first sampling rate. The second accuracy of speech recognition based on the audio data obtained based on the sampling rate corresponding to the target sampling parameter is higher than the first accuracy. The first accuracy is the accuracy of speech recognition based on the audio data obtained based on the first sampling rate.

[0067] Step S25: Obtain the second audio data based on the target sampling parameters, and match the second audio data with the first audio data;

[0068] Step S26: Perform speech recognition on the second audio data to obtain the speech recognition result.

[0069] First audio data is obtained based on a first sampling rate. Target sampling parameters are determined based on the first audio data. The sampling rate corresponding to the target sampling parameters is different from the first sampling rate. The second accuracy of speech recognition based on the audio data obtained based on the sampling rate corresponding to the target sampling parameters is higher than the first accuracy of speech recognition based on the audio data obtained based on the first sampling rate. Second audio data matching the first audio data is obtained based on the target sampling parameters, and speech recognition is performed on it. The accuracy of the resulting speech recognition is higher than the accuracy of speech recognition based on the first audio data, thus improving the accuracy of speech recognition based on audio data.

[0070] Specifically, determining the target sampling parameters based on the first audio data can be achieved by: performing sampling transformations on the first audio data according to multiple different sampling transformation ratios to obtain multiple third audio data; determining the target audio data from the multiple third audio data, wherein the speech recognition result of the target audio data satisfies the speech recognition conditions; and determining the target sampling transformation ratio corresponding to the target audio data as the target sampling parameters.

[0071] After obtaining the first audio data, in order to determine the optimal sampling parameters (i.e. target sampling parameters) corresponding to the first audio data, the first audio data can be sampled and transformed according to different sampling transformation ratios to obtain multiple third audio data. The multiple third audio data are compared to determine one third audio data from the multiple third audio data, which is used as the target audio data. The target sampling transformation ratio corresponding to the target audio data is the target sampling parameter.

[0072] The optimal sampling parameter is the third audio data obtained by sampling and transforming the first audio data according to this optimal parameter. This third audio data has the highest recognition accuracy during speech recognition. For example, given the first audio data, multiple different sampling transformation ratios are available: sampling transformation ratio 1, sampling transformation ratio 2, and sampling transformation ratio 3. Sampling and transforming the first audio data according to sampling transformation ratio 1 yields third audio data 1; sampling and transforming the first audio data according to sampling transformation ratio 2 yields third audio data 2; and sampling and transforming the first audio data according to sampling transformation ratio 3 yields third audio data 3. At this point, third audio data 1, third audio data 2, and third audio data 3 need to be compared to determine the target audio data. For example, if the target audio data is third audio data 2, then the sampling transformation ratio 2 corresponding to third audio data 2 is the target sampling parameter.

[0073] If the third audio data 2 is determined as the target audio data, it can be shown that the accuracy of speech recognition on the third audio data 2 is higher than that on the third audio data 1 and the third audio data 3. Therefore, the sampling transformation ratio 2 corresponding to the third audio data 2 is the optimal sampling parameter.

[0074] Specifically, comparing the third audio data 1, third audio data 2, and third audio data 3 can be done as follows: performing speech recognition on third audio data 1 to obtain recognition result 1; performing speech recognition on third audio data 2 to obtain recognition result 2; performing speech recognition on third audio data 3 to obtain recognition result 3; determining and comparing the recognition accuracy of recognition result 1, recognition result 2, and recognition result 3; and selecting the one with the highest recognition accuracy. For example, if the recognition accuracy of recognition result 2 is higher than that of recognition result 1 and also higher than that of recognition result 3, then the third audio data 2 corresponding to recognition result 2 is determined as the target audio data.

[0075] The method for determining the target audio from multiple third-party audio data can be as follows: perform speech recognition on each third-party audio data separately to obtain the speech recognition result of each audio data, compare the speech recognition results of multiple third-party audio data to obtain the comparison result, and determine the target audio data from multiple third-party audio data based on the comparison result.

[0076] Specifically, comparing the speech recognition results of multiple third audio data to obtain the comparison result can be achieved by: determining the perplexity of each speech recognition result among the multiple third audio data, comparing the perplexity of the speech recognition results of the multiple third audio data, and obtaining the comparison result.

[0077] The perplexity of each speech recognition result is used to evaluate the fluency of the corresponding text sequence. A lower perplexity indicates a more fluent text sequence, thus further confirming the accuracy of the recognition result. The perplexity of each speech recognition result can be determined through a pre-trained language model to ensure the accuracy of the perplexity, thereby guaranteeing the overall accuracy of speech recognition.

[0078] For example, if the first sampling rate is 16K, after obtaining the first audio data at a sampling rate of 16K, to determine the target sampling parameters, upsampling and downsampling can be performed on the first audio data. This involves determining multiple different sampling transformation ratios, such as 1 / 4, 1 / 2, 2, 3, and 4. Downsampling the first audio data at a 1 / 4 sampling transformation ratio yields the third audio data 1 with a sampling rate of 4K; downsampling the first audio data at a 1 / 2 sampling transformation ratio yields the third audio data 2 with a sampling rate of 8K; upsampling the first audio data at twice the sampling transformation ratio yields the third audio data 3 with a sampling rate of 32K; and so on. An audio data set is upsampled by a factor of 3 to obtain a third audio data set 4 with a sampling rate of 48kHz. The first audio data set is upsampled by a factor of 4 to obtain a third audio data set 5 with a sampling rate of 64kHz. Speech recognition is then performed on the third audio data sets 1, 2, 3, 4, and 5 respectively, yielding corresponding speech recognition results. The perplexity of these five results is calculated. Based on the comparison of perplexity, the speech recognition result with the lowest perplexity is determined to be the one with a sampling rate of 2. This 2x sampling rate is then set as the target sampling parameter, and the second audio data set is determined accordingly.

[0079] The speech recognition method disclosed in this embodiment obtains first audio data based on a first sampling rate, then performs sampling transformations on the first audio data according to multiple different sampling transformation ratios to obtain multiple third audio data. From these multiple third audio data, a target audio data is determined. The speech recognition result of the target audio data satisfies the speech recognition conditions. The target sampling transformation ratio corresponding to the target audio data is determined as the target sampling parameter. This allows for the acquisition of second audio data based on the target sampling parameter, and subsequent speech recognition, thereby achieving the goal of improving speech recognition accuracy through adjusting the sampling parameters. This scheme, after obtaining the first audio data, performs sampling transformations on it according to different sampling transformation ratios, and then determines the target audio data that satisfies the speech recognition conditions from the multiple sampled and transformed third audio data. This determines the target sampling parameter without complex processing or resampling. Furthermore, selecting one third audio data as the target audio data and determining the target sampling parameter based on it ensures the accuracy of the determined target sampling parameter.

[0080] Furthermore, in the speech recognition method disclosed in this embodiment, after determining the target sampling parameters, it is necessary to obtain the second audio data based on the target sampling parameters in order to perform speech recognition on the second audio data. The second audio data is matched with the first audio data. At this time, the second audio data and the first audio data may be audio data of the same audio segment with different sampling rates, or they may be different parts corresponding to the same audio segment.

[0081] If the second audio data and the first audio data are audio data of the same audio segment but with different sampling rates, then the target audio data corresponding to the target sampling parameters can be directly determined as the second audio data, and the speech recognition result corresponding to the second audio data can be directly obtained.

[0082] For example, if speech recognition is needed for a certain audio segment, the audio segment can first be sampled at a first sampling rate to obtain first audio data. This first sampling rate can be the default sampling rate of the electronic device or a sampling rate set by the electronic device. After obtaining the first audio data, the first audio data is sampled and transformed according to multiple different sampling transformation ratios to obtain multiple third audio data. One third audio data is determined from the multiple third audio data, and this determined third audio data is designated as the target audio data. The target audio data is the second audio data, which is the audio data of the audio segment determined according to the target sampling rate. Speech recognition of the second audio data can then be performed to achieve speech recognition of the audio segment.

[0083] In this process, the first audio data is obtained, but speech recognition is not performed on the first audio data. Instead, speech recognition is performed on the second audio data obtained after the sampling transformation ratio of the first audio data, so as to improve the recognition accuracy.

[0084] In addition, if the second audio data and the first audio data are different parts of the same audio segment, the obtained initial audio data needs to be sampled and transformed according to the target sampling parameters to obtain the sampled and transformed second audio data. The sampling rate of the obtained initial audio data is the same as the sampling rate of the first audio data, and the initial audio data and the first audio data are different parts of the same audio segment.

[0085] To ensure the accuracy of speech recognition for a given audio segment, this embodiment first determines the sampling parameters that match the audio segment, i.e., the target sampling parameters. When the audio segment is sampled and transformed according to the target sampling parameters, the resulting second audio data will have higher accuracy in speech recognition than audio data obtained by transforming it according to other sampling parameters. Therefore, in this embodiment, the audio segment can be divided into different parts, and the target sampling parameters can be determined from one part. Then, each part of the audio segment is sampled and transformed according to the target sampling parameters to ensure the accuracy of speech recognition for that audio segment.

[0086] The audio segment may include at least a first audio segment. The target sampling parameter, i.e., the target sampling transformation ratio, is determined through the first audio segment, such as 2 times. After determining the target sampling transformation ratio, each part of the audio segment is sampled and transformed according to the target sampling transformation ratio, and speech recognition is performed on the audio data obtained after sampling transformation. If the audio segment also includes an initial audio segment, the initial audio segment is sampled and transformed according to the target sampling transformation ratio to obtain the second audio data after sampling transformation. Speech recognition is performed on the second audio data, and the accuracy of the speech recognition result is higher than the accuracy of the speech recognition result obtained by directly performing speech recognition on the first audio data that has not undergone sampling transformation.

[0087] In this case, if the initial audio data and the first audio data are different parts of the same audio segment, then the complete audio segment can be obtained directly. After obtaining the complete audio segment, a portion of it (such as the first audio segment) is sampled and transformed to determine the target sampling transformation ratio. Then, the other parts of the complete audio segment are sampled and transformed according to the target sampling transformation ratio. Therefore, since the complete audio segment is obtained directly, it can be determined that the sampling rate for obtaining the initial audio segment is the same as the sampling rate for obtaining the first audio data, both being the first sampling rate.

[0088] Alternatively, the electronic device may not directly acquire the complete audio segment, but rather acquire different parts of the audio segment sequentially. For example, if the audio segment includes part 1, part 2, and part 3, part 1 is acquired first to obtain the first audio data. Then, part 2 is acquired to obtain the first initial audio data. Finally, part 3 is acquired to obtain the second initial audio data. The step of determining the target sampling transformation ratio can be performed after all the audio segments are acquired. In this case, one audio data (e.g., the first audio data) is directly selected from the three acquired audio data (the first audio data, the first initial audio data, and the second initial audio data) to determine the target sampling transformation ratio. Then, the other audio data (e.g., the first and second initial audio data) are sampled and transformed according to the target sampling transformation ratio.

[0089] It should be noted that even if the complete audio segment is not obtained directly in this embodiment, but different parts of the audio segment are obtained sequentially, the sampling rate on which different parts are collected is the same, that is, the first sampling rate. This ensures that after determining the target sampling transformation ratio, other audio data can be sampled and transformed directly according to the target sampling transformation ratio, without having to determine different target sampling transformation ratios for different audio data. This ensures the efficiency of audio data processing and thus achieves the goal of ensuring the efficiency of speech recognition.

[0090] This embodiment discloses a speech recognition method, the flowchart of which is as follows: Figure 3 As shown, it includes:

[0091] Step S31: Obtain first audio data based on the first sampling rate;

[0092] Step S32: Resample the first audio data according to multiple different sampling rates to obtain multiple third audio data;

[0093] Step S33: Determine the target audio data from multiple third-party audio data, and the speech recognition result of the target audio data meets the speech recognition conditions;

[0094] Step S34: Determine the sampling rate corresponding to the target audio data as the target sampling parameter. The sampling rate corresponding to the target sampling parameter is different from the first sampling rate. The second accuracy of speech recognition based on the audio data obtained based on the sampling rate corresponding to the target sampling parameter is higher than the first accuracy. The first accuracy is the accuracy of speech recognition based on the audio data obtained based on the first sampling rate.

[0095] Step S35: Obtain the second audio data based on the target sampling parameters, and match the second audio data with the first audio data;

[0096] Step S36: Perform speech recognition on the second audio data to obtain the speech recognition result.

[0097] First audio data is obtained based on a first sampling rate. Target sampling parameters are determined based on the first audio data. The sampling rate corresponding to the target sampling parameters is different from the first sampling rate. The second accuracy of speech recognition based on the audio data obtained based on the sampling rate corresponding to the target sampling parameters is higher than the first accuracy of speech recognition based on the audio data obtained based on the first sampling rate. Second audio data matching the first audio data is obtained based on the target sampling parameters, and speech recognition is performed on it. The accuracy of the resulting speech recognition is higher than the accuracy of speech recognition based on the first audio data, thus improving the accuracy of speech recognition based on audio data.

[0098] Specifically, determining the target sampling parameter based on the first audio data can be achieved by: resampling the first audio data at multiple different sampling rates to obtain multiple third audio data; determining the target audio data from the multiple third audio data, wherein the speech recognition result of the target audio data satisfies the speech recognition conditions; and determining the sampling rate corresponding to the target audio data as the target sampling parameter.

[0099] After obtaining the first audio data based on the first sampling rate, in order to determine the target sampling parameters corresponding to the first audio data, the first audio data can be resampled at different sampling rates to obtain multiple third audio data. The multiple third audio data correspond to different sampling rates, and the multiple third audio data correspond to the same audio segment as the first audio data. That is, the same audio segment is collected at different sampling rates to obtain multiple third audio data and the first audio data.

[0100] When resampling the first audio data at multiple different sampling rates, the first sampling rate may not be included among the multiple different sampling rates. For example, if the first sampling rate is 16K, then 16K is not included among the multiple different sampling rates. Therefore, the first audio data will not be included among the multiple third audio data. In this case, when determining the target audio data, it is necessary to select one audio data from the multiple third audio data and the first audio data as the target audio data to ensure that the selected target audio data has the highest speech recognition accuracy among the audio data (multiple third audio data and first audio data) obtained after collecting the audio segment.

[0101] For example, the first sampling rate is 16K, and the first audio data is obtained based on the 16K sampling rate. Multiple different sampling rates are available, including 8K, 32K, 48K, and 64K. Resampling is performed based on these different sampling rates to obtain the third audio data 1 corresponding to 8K, the third audio data 2 corresponding to 32K, the third audio data 3 corresponding to 48K, and the third audio data 4 corresponding to 64K. Then, one of the first audio data, the third audio data 1, the third audio data 2, the third audio data 3, and the third audio data 4 is selected as the target audio data. If the target audio data is the third audio data 3, then the sampling rate of 48K corresponding to the third audio data 3 can be determined as the target sampling rate. In this process, there is no need to resample the audio data with the 16K sampling rate, and a relatively complete audio data with a higher sampling rate can be obtained, thereby ensuring the accuracy of the finally determined target audio data.

[0102] Of course, when resampling the first audio data at multiple different sampling rates, the first sampling rate can also be included among these multiple different sampling rates. That is, after obtaining the first audio data based on the first sampling rate, it needs to be resampled at multiple different sampling rates. In this case, the multiple different sampling rates are pre-set and will not change. Regardless of the first sampling rate when the first audio data was obtained, these multiple different sampling rates remain fixed. In this example, obtaining the first audio data is only to determine the audio segment that needs to be used for speech recognition. Once the first audio data is obtained, it can be determined which audio segment the first audio data was obtained from, thereby determining the corresponding audio segment, and then resampling the audio segment at multiple different sampling rates.

[0103] Furthermore, determining the target audio data from multiple third-party audio data can be specifically as follows:

[0104] Speech recognition is performed on each third audio data separately to obtain the speech recognition result of each audio data; the speech recognition results of multiple third audio data are compared to obtain the comparison result; the target audio data is determined from multiple third audio data based on the comparison result.

[0105] The comparison results are obtained by comparing the speech recognition results of multiple third audio data. Specifically, this can be achieved by: determining the perplexity of each speech recognition result among the multiple third audio data; and comparing the perplexity of the speech recognition results of the multiple third audio data to obtain the comparison results.

[0106] The perplexity of each speech recognition result is used to evaluate the fluency of the corresponding text sequence. A lower perplexity indicates a more fluent text sequence, thus further confirming the accuracy of the recognition result. The perplexity of each speech recognition result can be determined through a pre-trained language model to ensure the accuracy of the perplexity, thereby guaranteeing the overall accuracy of speech recognition.

[0107] For example, if the first sampling rate is 16K, after obtaining the first audio data at a sampling rate of 16K, to determine the target sampling parameters, the first audio data can be resampled, i.e., multiple different sampling rates can be determined, such as 4K, 8K, 32K, 48K, and 64K. Therefore, acquiring the first audio data at a sampling rate of 4K yields the third audio data 1 with a sampling rate of 4K; acquiring the first audio data at a sampling rate of 8K yields the third audio data 2 with a sampling rate of 8K; acquiring the first audio data at a sampling rate of 32K yields the third audio data 3 with a sampling rate of 32K; and so on. Audio is acquired at a sampling rate of 48K to obtain the third audio data 4 with a sampling rate of 48K; the first audio data is acquired at a sampling rate of 64K to obtain the third audio data 5 with a sampling rate of 64K. Speech recognition is performed on the third audio data 1, third audio data 2, third audio data 3, third audio data 4, third audio data 5 and the first audio data respectively to obtain the corresponding speech recognition results. The perplexity of the six speech recognition results is calculated. Based on the comparison of perplexity, it is determined that the speech recognition result with the lowest perplexity is the one with a sampling rate of 32K. Therefore, the sampling rate of 32K is determined as the target sampling parameter, and the second audio data is determined accordingly.

[0108] Furthermore, in the speech recognition method disclosed in this embodiment, after determining the target sampling parameter (target sampling rate), it is necessary to obtain second audio data based on the target sampling parameter in order to perform speech recognition on the second audio data and obtain the speech recognition result. Specifically, obtaining the second audio data can be achieved by resampling the initial audio data (audio data that needs to be performed on speech recognition) according to the target sampling rate to obtain second audio data that matches the target sampling rate.

[0109] Alternatively, instead of resampling, the initial audio data can be directly adjusted according to the target sampling rate. The resulting audio data is the second audio data, with the sampling rate of the second audio data being the target sampling rate. For example, if the target sampling rate is 32K and the initial audio data has a sampling rate of 16K, then the sampling rate of the initial audio data is increased to 32K to obtain the second audio data with a sampling rate of 32K, which can then be used for speech recognition.

[0110] The speech recognition method disclosed in this embodiment, after obtaining first audio data based on a first sampling rate, resamples the first audio data according to multiple different sampling rates to obtain multiple third audio data. Target audio data is then determined from these third audio data, and the sampling rate corresponding to the target audio data is used as the target sampling parameter. This method improves the accuracy of speech recognition by obtaining second audio data according to the target sampling parameter and performing speech recognition on it. This solution ensures the accuracy of the target sampling parameter determination by resampling the first audio data according to multiple different sampling rates, thereby guaranteeing the accuracy of speech recognition of the audio data.

[0111] This embodiment discloses a speech recognition method, the flowchart of which is as follows: Figure 4 As shown, it includes:

[0112] Step S41: Obtain first audio data based on the first sampling rate;

[0113] Step S42: Determine the target sampling parameters based on the first audio data. The sampling rate corresponding to the target sampling parameters is different from the first sampling rate. The second accuracy of speech recognition based on the audio data obtained based on the sampling rate corresponding to the target sampling parameters is higher than the first accuracy. The first accuracy is the accuracy of speech recognition based on the audio data obtained based on the first sampling rate.

[0114] Step S43: Determine the target sampling rate based on the first sampling rate and the target sampling parameters;

[0115] Step S44: Resample the initial audio data according to the target sampling rate to obtain second audio data that matches the target sampling rate. The sampling rate of the initial audio data is the same as that of the first audio data. The initial audio data and the first audio data are different parts of the same audio segment.

[0116] Step S45: Perform speech recognition on the second audio data to obtain the speech recognition result.

[0117] First audio data is obtained based on a first sampling rate. Target sampling parameters are determined based on the first audio data. The sampling rate corresponding to the target sampling parameters is different from the first sampling rate. The second accuracy of speech recognition based on the audio data obtained based on the sampling rate corresponding to the target sampling parameters is higher than the first accuracy of speech recognition based on the audio data obtained based on the first sampling rate. Second audio data matching the first audio data is obtained based on the target sampling parameters, and speech recognition is performed on it. The accuracy of the resulting speech recognition is higher than the accuracy of speech recognition based on the first audio data, thus improving the accuracy of speech recognition based on audio data.

[0118] Specifically, obtaining the second audio data based on the target sampling parameters can be achieved by: determining the target sampling rate based on the first sampling rate and the target sampling parameters; resampling the initial audio data according to the target sampling rate to obtain the second audio data that matches the target sampling rate, wherein the sampling rate of the initial audio data is the same as that of the first audio data, and the initial audio data and the first audio data are different parts of the same audio segment.

[0119] The target sampling parameter is the target sampling transformation ratio. When obtaining the first audio data and determining the target sampling transformation ratio based on the first audio data, it can be done by: performing sampling transformation on the first audio data according to multiple different sampling transformation ratios to obtain multiple third audio data, determining the target audio data from the multiple third audio data, and determining the target audio data as the target sampling parameter when the speech recognition result of the target audio data meets the speech recognition conditions.

[0120] In this embodiment, after obtaining the target sampling transformation ratio, the second audio data can be obtained by resampling. That is, the current sampling rate and the target sampling transformation ratio are determined. Then, the target sampling rate is determined based on the current sampling rate and the target sampling transformation ratio. The audio data to be recognized is resampled according to the target sampling rate to obtain the second audio data. The second audio data matches the target sampling rate. The accuracy of the speech recognition result obtained by performing speech recognition on the second audio data is higher than the accuracy of the speech recognition result obtained by performing speech recognition on the first audio data obtained according to the first sampling rate.

[0121] The audio data that needs to be resampled can be: the audio data that needs to be used for speech recognition, such as: the first audio data, audio data that belongs to a different part of the same audio segment as the first audio data, audio data that does not belong to the same audio segment as the first audio data, etc.

[0122] If the audio data to be recognized for speech is the first audio data, it is obtained based on the first sampling rate. In order to improve the accuracy of the speech recognition result obtained from the first audio data, it is necessary to resample the first audio data. In this case, the initial audio data is the first audio data. If the first audio data is obtained by sampling a specific audio segment, the specific audio segment is directly resampled according to the target sampling rate to obtain the second audio data. The sampling rate of the second audio data is the target sampling rate. Speech recognition can be performed directly on the second audio data.

[0123] If the audio data to be recognized for speech is a different part of the same audio segment as the first audio data, the electronic device directly obtains the complete audio segment. This complete audio segment includes at least two parts: the first audio data and the initial audio data. In this case, the sampling rate of the initial audio data is the same as that of the first audio data, both being the first sampling rate. After determining the target sampling transformation ratio based on the first audio data and the target sampling rate based on the first sampling rate and the target sampling transformation ratio, the initial audio data needs to be resampled according to the target sampling rate to obtain the second audio data. The sampling rate of the second audio data is the target sampling rate. Of course, in this case, any part of the complete audio segment needs to be resampled according to the target sampling rate to ensure the accuracy of speech recognition for the complete audio segment. In this case, multiple parts of the complete audio segment can be resampled separately according to the target sampling rate, or the complete audio segment can be resampled directly according to the target sampling rate to ensure that the second audio data obtained after resampling corresponds to the complete audio segment.

[0124] If the audio data to be recognized for speech is not in the same audio segment as the first audio data, the electronic device, after determining the target sampling rate, directly determines the audio segment corresponding to the audio data (i.e., the initial audio data) that is not in the same audio segment as the first audio data, and samples the audio segment according to the target sampling rate to obtain the second audio data corresponding to the audio segment. The accuracy of the speech recognition result obtained by performing speech recognition on the second audio data is higher than the accuracy of the speech recognition result obtained by performing speech recognition on the corresponding initial audio data directly.

[0125] Of course, after obtaining the first audio data, determining the target sampling parameter based on the first audio data can also be done as follows: resample the first audio data according to multiple different sampling rates to obtain multiple third audio data; determine the target audio data from the multiple third audio data, and the speech recognition result of the target audio data satisfies the speech recognition conditions. At this time, the sampling rate corresponding to the target audio data is directly determined as the target sampling parameter. The target sampling parameter determined in this way does not need to go through the first sampling rate and the sampling transformation ratio to determine the target sampling rate. Instead, the target sampling rate can be obtained directly. After determining the target sampling rate, the initial audio data is directly resampled according to the target sampling rate to obtain the second audio data that matches the target sampling rate. After obtaining the second audio data, speech recognition is performed on the second audio data to obtain the speech recognition result.

[0126] The speech recognition method disclosed in this embodiment, after obtaining first audio data based on a first sampling rate and determining target sampling parameters based on the first audio data, needs to determine a target sampling rate based on the first sampling rate and the target sampling parameters, so as to resample the initial audio data according to the target sampling rate to obtain second audio data that matches the target sampling rate. Then, speech recognition is performed on the second audio data to ensure that the accuracy of the obtained speech recognition result is higher than the accuracy of the speech recognition result obtained by directly performing speech recognition on the first audio data that has not been resampled. This realizes the improvement of the accuracy of speech recognition of the adjusted audio data by adjusting the sampling rate, and effectively reduces the situation of speech recognition errors.

[0127] This embodiment discloses a speech recognition method, the flowchart of which is as follows: Figure 5 As shown, it includes:

[0128] Step S51: Obtain first audio data based on the first sampling rate;

[0129] Step S52: In response to obtaining target information, determine target sampling parameters based on the first audio data. The target information indicates that the first audio data is audio played at a variable speed. The sampling rate corresponding to the target sampling parameters is different from the first sampling rate. The second accuracy of speech recognition based on the audio data obtained based on the sampling rate corresponding to the target sampling parameters is higher than the first accuracy. The first accuracy is the accuracy of speech recognition based on the audio data obtained based on the first sampling rate.

[0130] Step S53: Obtain the second audio data based on the target sampling parameters, and match the second audio data with the first audio data;

[0131] Step S54: Perform speech recognition on the second audio data to obtain the speech recognition result.

[0132] First audio data is obtained based on a first sampling rate. Target sampling parameters are determined based on the first audio data. The sampling rate corresponding to the target sampling parameters is different from the first sampling rate. The second accuracy of speech recognition based on the audio data obtained based on the sampling rate corresponding to the target sampling parameters is higher than the first accuracy of speech recognition based on the audio data obtained based on the first sampling rate. Second audio data matching the first audio data is obtained based on the target sampling parameters, and speech recognition is performed on it. The accuracy of the resulting speech recognition is higher than the accuracy of speech recognition based on the first audio data, thus improving the accuracy of speech recognition based on audio data.

[0133] Specifically, determining the target sampling parameters based on the first audio data can be achieved by: in response to obtaining target information, determining the target sampling parameters based on the first audio data, wherein the target information indicates that the first audio data is variable speed playback audio.

[0134] If the target sampling parameter is the target sampling rate, then after the target sampling rate is determined, the electronic device will acquire audio data according to the target sampling rate when it subsequently obtains audio data. That is, as long as the target sampling rate is determined, the electronic device will directly obtain audio data according to the target sampling rate and directly perform speech recognition on the obtained audio data to ensure the accuracy of speech recognition of the obtained audio data.

[0135] If the target sampling parameter is the target sampling transformation ratio, then after determining the target sampling transformation ratio, when the electronic device obtains audio data, it still obtains audio data according to the initial sampling rate (e.g., the first sampling rate). After obtaining the audio data, it needs to perform sampling transformation according to the target sampling transformation ratio to obtain the audio data after sampling transformation, and then perform speech recognition on the audio data after sampling transformation to ensure the accuracy of speech recognition.

[0136] When an electronic device collects audio data, if the audio data collected is obtained by collecting audio segments played by other devices (such as electronic devices or other devices), then if the playback speed of the audio segment changes, it will affect the accuracy of the electronic device in collecting audio data and performing speech recognition on the audio data.

[0137] For example, if a playback device plays an audio clip at 1x speed and an electronic device collects audio data at a sampling rate of 16kHz, the accuracy of the audio data obtained at this time is relatively high (i.e., it meets the conditions for speech recognition). In this case, the sampling rate of 16kHz is the target sampling rate. When the playback device changes the playback speed, such as playing the audio clip at 2x speed, the electronic device still collects audio data at the current target sampling rate of 16kHz. However, the accuracy of the audio data obtained at this time will decrease when used for speech recognition. That is, at the same sampling rate, different audio playback speeds will affect the accuracy of speech recognition of the collected audio.

[0138] Therefore, when the playback speed of the playback device changes, the target sampling parameters (target sampling rate or target sampling transformation ratio) need to be adjusted to ensure the accuracy of speech recognition based on the audio data obtained by the electronic device.

[0139] Specifically, after the electronic device obtains the first audio data, it needs to determine whether the target information has been obtained. Only if the target information is determined will the target sampling parameters be determined based on the first audio data, thereby obtaining the second audio data and performing speech recognition on the second audio data. If the target information is not obtained, it means that the playback speed of the first audio data obtained at the time of playback has not been switched. In this case, the first sampling rate at which the first audio data was obtained is the current optimal sampling rate (target sampling rate), and speech recognition can be performed directly on the first audio data. Alternatively, if the target information is not obtained, the first audio data is obtained according to the first sampling rate, and the sampling rate is still switched according to the current target sampling transformation ratio. That is, the target sampling transformation ratio at this time is the current optimal sampling transformation ratio, and there is no need to switch the current sampling transformation ratio.

[0140] The target information represents the first audio data as variable-speed playback audio, that is, the first audio data is the audio collected after the playback speed changes. In other words, whether the playback speed of the playback device changes can be represented by the target information. If the playback speed of the playback device changes, the target information is generated and output to the electronic device when playing the audio segment. If the playback speed of the playback device does not change, there is no need to generate target information. In this case, the electronic device collects audio data, but cannot obtain target information, so there is no need to redetermine the target sampling parameters.

[0141] The speech recognition method disclosed in this embodiment, after obtaining first audio data based on a first sampling rate, only determines target sampling parameters based on the first audio data after obtaining target information, and further obtains second audio data based on the target sampling parameters, so as to perform speech recognition on the second audio data and obtain a speech recognition result. In this scheme, when the first audio data is obtained and it is determined that the first audio data is audio played at a variable speed, in order to avoid the problem of reduced accuracy of speech recognition of audio played at a variable speed due to the variable speed playback, when it is determined that the obtained first audio data is audio played at a variable speed, it is necessary to determine target sampling parameters based on the first audio data, so as to obtain second audio data based on the target sampling parameters, thereby ensuring the accuracy of speech recognition and avoiding the problem of reduced accuracy when performing speech recognition on audio data due to variable speed playback.

[0142] This embodiment discloses an electronic device, the structural schematic diagram of which is shown below. Figure 6 As shown, it includes:

[0143] Audio acquisition device 61 and processor 62.

[0144] The audio acquisition device 61 is used to acquire audio data;

[0145] The processor 62 is used to control the audio acquisition device to obtain first audio data based on a first sampling rate; determine target sampling parameters based on the first audio data, wherein the sampling rate corresponding to the target sampling parameters is different from the first sampling rate, and the second accuracy of speech recognition based on the audio data obtained based on the sampling rate corresponding to the target sampling parameters is higher than the first accuracy, wherein the first accuracy is the accuracy of speech recognition based on the audio data obtained based on the first sampling rate; obtain second audio data based on the target sampling parameters, wherein the second audio data is matched with the first audio data; and perform speech recognition on the second audio data to obtain a speech recognition result.

[0146] The electronic device disclosed in this embodiment is implemented based on the speech recognition method disclosed in the above embodiments, and will not be described again here.

[0147] The electronic device disclosed in this embodiment, after obtaining first audio data based on a first sampling rate, determines target sampling parameters based on the first audio data. The sampling rate corresponding to the target sampling parameters is different from the first sampling rate. The second accuracy of speech recognition based on the audio data obtained based on the sampling rate corresponding to the target sampling parameters is higher than the first accuracy. The first accuracy is the accuracy of speech recognition based on the audio data obtained based on the first sampling rate. Then, second audio data is obtained based on the target sampling parameters, and the second audio data is matched with the first audio data. Speech recognition is then performed on the second audio data to obtain a speech recognition result. This solution, after determining the target sampling parameters based on the first audio data, obtains the second audio data based on the target sampling parameters and performs speech recognition on the second audio data. Since the second accuracy of speech recognition based on the audio data obtained based on the sampling rate corresponding to the target sampling parameters is higher than the accuracy of speech recognition based on the audio data obtained based on the first sampling rate, the accuracy of speech recognition based on the second audio data is higher than the accuracy of speech recognition based on the first audio data. This achieves the goal of improving the accuracy of speech recognition of adjusted audio data by adjusting the target sampling parameters, effectively reducing speech recognition errors.

[0148] This embodiment discloses a speech recognition system, the structural diagram of which is shown below. Figure 7 As shown, it includes:

[0149] The system comprises a first acquisition unit 71, a determination unit 72, a second acquisition unit 73, and a speech recognition unit 74.

[0150] The first obtaining unit 71 is used to obtain first audio data based on the first sampling rate;

[0151] The determining unit 72 is used to determine the target sampling parameters based on the first audio data. The sampling rate corresponding to the target sampling parameters is different from the first sampling rate. The second accuracy of speech recognition based on the audio data obtained based on the sampling rate corresponding to the target sampling parameters is higher than the first accuracy. The first accuracy is the accuracy of speech recognition based on the audio data obtained based on the first sampling rate.

[0152] The second acquisition unit 73 is used to obtain second audio data based on the target sampling parameters, and the second audio data is matched with the first audio data.

[0153] The speech recognition unit 74 is used to perform speech recognition on the second audio data to obtain the speech recognition result.

[0154] Furthermore, the determining unit is used for:

[0155] The first audio data is sampled and transformed according to multiple different sampling transformation ratios to obtain multiple third audio data; the target audio data is determined from the multiple third audio data, and the speech recognition result of the target audio data satisfies the speech recognition conditions; the target sampling transformation ratio corresponding to the target audio data is determined as the target sampling parameter.

[0156] Furthermore, the second audio data and the first audio data are audio data of the same audio segment at different sampling rates, and the second acquisition unit is used for:

[0157] The target audio data is identified as the second audio data.

[0158] Furthermore, the target sampling parameter characterizes the sampling transformation ratio, and the second obtaining unit is used for:

[0159] The initial audio data is sampled and transformed according to the target sampling parameters to obtain the second audio data after sampling transformation. The sampling rate of the initial audio data is the same as that of the first audio data, and the initial audio data and the first audio data are different parts of the same audio segment.

[0160] Furthermore, the second obtaining unit is used for:

[0161] Based on the first sampling rate and the target sampling parameters, the target sampling rate is determined; the initial audio data is resampled according to the target sampling rate to obtain the second audio data that matches the target sampling rate; the sampling rate of the initial audio data is the same as that of the first audio data, and the initial audio data and the first audio data are different parts of the same audio segment.

[0162] Furthermore, the determining unit is used for:

[0163] The first audio data is resampled at multiple different sampling rates to obtain multiple third audio data; the target audio data is determined from the multiple third audio data, and the speech recognition result of the target audio data meets the speech recognition conditions; the sampling rate corresponding to the target audio data is determined as the target sampling parameter.

[0164] Furthermore, the determining unit is used for:

[0165] Speech recognition is performed on each third audio data separately to obtain the speech recognition result of each audio data; the speech recognition results of multiple third audio data are compared to obtain the comparison result; the target audio data is determined from multiple third audio data based on the comparison result.

[0166] Furthermore, the determining unit is used for:

[0167] Determine the perplexity of each speech recognition result among multiple third audio data; compare the perplexity of the speech recognition results of multiple third audio data to obtain the comparison result.

[0168] Furthermore, the determining unit is used for:

[0169] In response to obtaining target information, target sampling parameters are determined based on the first audio data, and the target information indicates that the first audio data is audio played at a variable speed.

[0170] The speech recognition system disclosed in this embodiment is implemented based on the speech recognition method disclosed in the above embodiments, and will not be described again here.

[0171] The speech recognition system disclosed in this embodiment, after obtaining first audio data based on a first sampling rate, determines target sampling parameters based on the first audio data. The sampling rate corresponding to the target sampling parameters is different from the first sampling rate. The second accuracy of speech recognition based on the audio data obtained based on the sampling rate corresponding to the target sampling parameters is higher than the first accuracy, where the first accuracy is the accuracy of speech recognition based on the audio data obtained based on the first sampling rate. Then, second audio data is obtained based on the target sampling parameters, matched with the first audio data, and speech recognition is performed on the second audio data to obtain a speech recognition result. This solution, after determining the target sampling parameters based on the first audio data, obtains the second audio data based on the target sampling parameters and performs speech recognition on the second audio data. Since the second accuracy of speech recognition based on the audio data obtained based on the sampling rate corresponding to the target sampling parameters is higher than the accuracy of speech recognition based on the audio data obtained based on the first sampling rate, the accuracy of speech recognition on the second audio data is higher than the accuracy of speech recognition on the first audio data. This achieves the goal of improving the accuracy of speech recognition of adjusted audio data by adjusting the target sampling parameters, effectively reducing speech recognition errors.

[0172] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the device embodiment drawings provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.

[0173] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CPUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, training equipment, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0174] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.

[0175] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).

Claims

1. A speech recognition method, comprising: The first audio data is obtained based on the first sampling rate; Target sampling parameters are determined based on the first audio data. The sampling rate corresponding to the target sampling parameters is different from the first sampling rate. The second accuracy of speech recognition based on the audio data obtained based on the sampling rate corresponding to the target sampling parameters is higher than the first accuracy. The first accuracy is the accuracy of speech recognition based on the audio data obtained based on the first sampling rate. Second audio data is obtained based on the target sampling parameters, and the second audio data is matched with the first audio data; The second audio data is subjected to speech recognition to obtain the speech recognition result.

2. The method according to claim 1, wherein determining the target sampling parameters based on the first audio data includes: The first audio data is sampled and transformed according to multiple different sampling transformation ratios to obtain multiple third audio data. Target audio data is determined from the plurality of third audio data, wherein the speech recognition result of the target audio data satisfies the speech recognition condition; The target sampling transformation ratio corresponding to the target audio data is determined as the target sampling parameter.

3. The method according to claim 2, wherein the second audio data and the first audio data are audio data of the same audio segment at different sampling rates, and the step of obtaining the second audio data based on the target sampling parameter includes: The target audio data is identified as the second audio data.

4. The method according to claim 2, wherein the target sampling parameter characterizes the sampling transformation ratio, and the step of obtaining the second audio data based on the target sampling parameter includes: The initial audio data is sampled and transformed according to the target sampling parameters to obtain the second audio data after sampling transformation. The sampling rate of the initial audio data is the same as that of the first audio data, and the initial audio data and the first audio data are different parts of the same audio segment.

5. The method according to claim 1, wherein obtaining the second audio data based on the target sampling parameters comprises: The target sampling rate is determined based on the first sampling rate and the target sampling parameters; The initial audio data is resampled according to the target sampling rate to obtain second audio data that matches the target sampling rate; the sampling rate of the initial audio data is the same as that of the first audio data, and the initial audio data and the first audio data are different parts of the same audio segment.

6. The method according to claim 1, wherein determining the target sampling parameters based on the first audio data includes: The first audio data is resampled at multiple different sampling rates to obtain multiple third audio data. Target audio data is determined from the plurality of third audio data, wherein the speech recognition result of the target audio data satisfies the speech recognition condition; The sampling rate corresponding to the target audio data is determined as the target sampling parameter.

7. The method according to claim 2 or 6, determining target audio data from a plurality of third audio data, comprising: Speech recognition is performed on each third audio data separately to obtain the speech recognition result for each audio data; The speech recognition results of multiple third-party audio data are compared to obtain the comparison results; The target audio data is determined from the plurality of third audio data based on the comparison results.

8. The method according to claim 7, wherein comparing the speech recognition results of multiple third audio data to obtain a comparison result includes: Determine the perplexity of each speech recognition result among the multiple third audio data; The perplexity of the speech recognition results of the multiple third audio data is compared to obtain the comparison result.

9. The method according to claim 1, wherein determining the target sampling parameters based on the first audio data comprises: In response to obtaining target information, target sampling parameters are determined based on the first audio data, wherein the target information indicates that the first audio data is audio played at a variable speed.

10. An electronic device, comprising: An audio acquisition device used to obtain audio data; The processor is used to control the audio acquisition device to obtain first audio data based on a first sampling rate; Target sampling parameters are determined based on the first audio data. The sampling rate corresponding to the target sampling parameters is different from the first sampling rate. The second accuracy of speech recognition based on the audio data obtained based on the sampling rate corresponding to the target sampling parameters is higher than the first accuracy. The first accuracy is the accuracy of speech recognition based on the audio data obtained based on the first sampling rate. Second audio data is obtained based on the target sampling parameters. The second audio data is matched with the first audio data. Speech recognition is performed on the second audio data to obtain a speech recognition result.