Speech recognition result screening method and device, electronic equipment and medium

By screening and analyzing candidate texts and their varying recognition frequency parameters from the speech recognition model, the text to be labeled is determined, solving the problem of labor-intensive manual correction and achieving efficient screening and labeling of speech recognition results.

CN119049470BActive Publication Date: 2025-11-21BEIJING XIAOMI MOBILE SOFTWARE CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310620490.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-29
Publication Date
2025-11-21
Estimated Expiration
2043-05-29

AI Technical Summary

Technical Problem

Existing methods for correcting speech recognition results are labor-intensive, especially the manual annotation of large amounts of data, which cannot be routinely implemented due to time and labor costs.

Method used

By acquiring candidate recognition texts and their recognition counts output by a preset speech recognition model within a historical time period, target recognition texts identical to those in the preset corpus are selected, and the texts to be labeled are determined based on the parameters of changes in the number of recognition counts, thus reducing the amount of manual labeling.

Benefits of technology

This reduces the number of annotations required for subsequent speech recognition results, saves human resources, and accurately locates the text to be annotated, thus improving recognition accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119049470B_ABST
    Figure CN119049470B_ABST
Patent Text Reader

Abstract

The method, device, electronic equipment and medium for screening a voice recognition result provided by the present disclosure, the method comprising: obtaining candidate recognition texts output by a preset voice recognition model in a first historical time period, and first recognition times corresponding to the candidate recognition texts respectively; screening the plurality of candidate recognition texts to obtain first target recognition texts identical to texts in a preset corpus; for any one of the first target recognition texts, obtaining a recognition times change parameter corresponding to the first target recognition text based on the first recognition times corresponding to the first target recognition text and second recognition times corresponding to the first target recognition text; and determining a to-be-labeled recognition text from the first target recognition texts based on the recognition times change parameters corresponding to the first target recognition texts respectively. The method can reduce the number of voice recognition results to be labeled.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of speech recognition, and particularly relates to a speech recognition result screening method and device, electronic equipment and medium. BACKGROUND

[0002] In the process of recognizing speech through a speech recognition model, there will inevitably be some inaccurate recognition. By correcting the inaccurate recognition in the speech recognition result, the subsequent speech recognition model recognition accuracy can be improved.

[0003] In related technologies, there is a method of correcting the recognition result by manual labeling. However, the method of correcting by manual labeling has the problem of consuming manpower. SUMMARY

[0004] To overcome the problems in related technologies, the present disclosure provides a speech recognition result screening method, device, electronic equipment and medium.

[0005] According to a first aspect of an embodiment of the present disclosure, a speech recognition result screening method is provided, which comprises:

[0006] obtaining a plurality of candidate recognition texts output by a preset speech recognition model in a first historical time period, and a first recognition number corresponding to each of the plurality of candidate recognition texts, the first recognition number corresponding to a candidate recognition text representing the number of times that the preset speech recognition model outputs the candidate recognition text in the first historical time period;

[0007] screening the plurality of candidate recognition texts to obtain a first target recognition text identical to a text in a preset corpus, the text in the preset corpus being a recognition-accurate text corresponding to to-be-recognized speech data;

[0008] For each of the first target recognition texts, based on the first recognition number corresponding to the first target recognition text and a second recognition number corresponding to the first target recognition text, a recognition number change parameter corresponding to the first target recognition text is obtained, the second recognition number corresponding to a first target recognition text representing the number of times that the preset speech recognition model outputs the first target recognition text in a time period before the first historical time period;

[0009] Based on the recognition number change parameter corresponding to each of the first target recognition texts, a to-be-labeled recognition text is determined from each of the first target recognition texts.

[0010] Optionally, the time period before the first historical time period comprises a second historical time period adjacent to the first historical time period and having the same time length as the first historical time period, the first historical time period being before the second historical time period, the obtaining of the recognition frequency change parameter corresponding to the first target recognized text based on the first recognition frequency corresponding to the first target recognized text and the second recognition frequency corresponding to the first target recognized text in the second historical time period comprises:

[0011] the obtaining of the recognition frequency change parameter corresponding to the first target recognized text based on the first recognition frequency corresponding to the first target recognized text and the second recognition frequency corresponding to the first target recognized text in the second historical time period.

[0012] Optionally, the obtaining of the recognition frequency change parameter corresponding to the first target recognized text based on the first recognition frequency corresponding to the first target recognized text and the second recognition frequency corresponding to the first target recognized text in the second historical time period comprises:

[0013] determining a difference between the first recognition frequency corresponding to the first target recognized text and the second recognition frequency corresponding to the first target recognized text in the second historical time period;

[0014] determining the difference as the recognition frequency change parameter corresponding to the first target recognized text, and / or determining a ratio of the difference to the second recognition frequency corresponding to the first target recognized text in the second historical time period as the recognition frequency change parameter corresponding to the first target recognized text;

[0015] the determining of the to-be-labeled recognized text from each of the first target recognized texts based on the recognition frequency change parameter corresponding to each of the first target recognized texts comprises:

[0016] the obtaining of a preset number of first target recognized texts corresponding to the recognition frequency change parameter in a front rank as the to-be-labeled recognized text from each of the first target recognized texts, and / or the obtaining of a preset number of first target recognized texts corresponding to the recognition frequency change parameter in a rear rank as the to-be-labeled recognized text.

[0017] Optionally, the first historical time period corresponds to a preset time length, the time period before the first historical time period comprises a second historical time period adjacent to the first historical time period and having the preset time length, and a third historical time period adjacent to the first historical time period and greater than the preset time length, the obtaining of the recognition frequency change parameter corresponding to the first target recognized text based on the first recognition frequency corresponding to the first target recognized text and the second recognition frequency corresponding to the first target recognized text in the second historical time period comprises:

[0018] The first target recognition text is obtained based on the first recognition frequency corresponding to the first target recognition text, the second recognition frequency corresponding to the first target recognition text in the second historical time period, and the third recognition frequency corresponding to the first target recognition text in the third historical time period.

[0019] Optionally, the first target recognition text is obtained based on the first recognition frequency corresponding to the first target recognition text, the second recognition frequency corresponding to the first target recognition text in the second historical time period, and the third recognition frequency corresponding to the first target recognition text in the third historical time period.

[0020] The product of the first value and the second value is determined as the recognition frequency variation parameter corresponding to the first target recognition text, the first value is the ratio of the first recognition frequency corresponding to the first target recognition text and the third recognition frequency, and the second value is obtained based on the second recognition frequency corresponding to the first target recognition text and a preset logarithmic function.

[0021] The to-be-labeled recognition text is determined from each of the first target recognition texts based on the recognition frequency variation parameter corresponding to each of the first target recognition texts.

[0022] The first target recognition text corresponding to the recognition frequency variation parameter in the front is obtained from each of the first target recognition texts as the to-be-labeled recognition text.

[0023] Optionally, the method further comprises:

[0024] The third target recognition text is determined from the plurality of candidate recognition texts based on the first recognition frequency corresponding to each of the plurality of candidate recognition texts.

[0025] The to-be-processed recognition text is determined based on the third target recognition text.

[0026] Optionally, the method further comprises:

[0027] The plurality of candidate recognition texts are screened to obtain second target recognition texts different from the texts in the preset corpus, and the second target recognition texts are used for artificial recognition and labeling.

[0028] Optionally, the method further comprises:

[0029] The recognition accurate text of the voice data corresponding to the second target recognition text is obtained.

[0030] add the recognized accurate text to the preset corpus.

[0031] According to a second aspect of the embodiments of the present disclosure, a speech recognition result screening device is provided, which comprises:

[0032] The obtaining module is configured to obtain a plurality of candidate recognition texts output by a preset speech recognition model in a first historical time period, and a first recognition frequency corresponding to each of the plurality of candidate recognition texts, the first recognition frequency corresponding to a candidate recognition text representing a number of times that the preset speech recognition model outputs the candidate recognition text in the first historical time period;

[0033] The screening module is configured to screen the plurality of candidate recognition texts to obtain a first target recognition text identical to a text in a preset corpus, the text in the preset corpus being a recognized accurate text corresponding to the to-be-recognized speech data;

[0034] The computing module is configured to, for any one of the first target recognition texts, obtain a recognition frequency change parameter corresponding to the first target recognition text based on the first recognition frequency corresponding to the first target recognition text and a second recognition frequency corresponding to the first target recognition text, the second recognition frequency corresponding to a first target recognition text representing a number of times that the preset speech recognition model outputs the first target recognition text in a time period before the first historical time period;

[0035] The determining module is configured to determine a to-be-labeled recognition text from the first target recognition texts based on the recognition frequency change parameters corresponding to the first target recognition texts.

[0036] According to a third aspect of the embodiments of the present disclosure, an electronic device is provided, which comprises a processor, and a memory for storing processor-executable instructions, wherein the processor is configured to implement the steps of the method of the first aspect when executed.

[0037] According to a fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, which stores computer program instructions, the program instructions being executed by a processor to implement the steps of the method provided by the first aspect of the present disclosure.

[0038] The method, device, electronic device and medium for screening a voice recognition result provided by the present disclosure screen a plurality of candidate recognition texts output by a preset voice recognition model in a first historical time period and a first recognition frequency corresponding to each of the plurality of candidate recognition texts, obtain a first target recognition text identical to a text in a preset corpus, obtain a recognition frequency change parameter corresponding to each of the first target recognition texts based on the first recognition frequency corresponding to the first target recognition text and a second recognition frequency corresponding to the first target recognition text, and determine a to-be-labeled recognition text from the first target recognition texts based on the recognition frequency change parameter corresponding to each of the first target recognition texts. The to-be-labeled recognition text is obtained by mining the recognition frequency change parameter corresponding to the first target recognition text, which can reduce the number of voice recognition results to be labeled in the subsequent process, save human resources, and accurately locate the recognition text to be labeled.

[0039] It should be understood that the foregoing general description and the following detailed description are only exemplary and explanatory, and are not limiting to the present disclosure. BRIEF DESCRIPTION OF DRAWINGS

[0040] The accompanying drawings, which are incorporated into and form part of the specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0041] Figure 1 FIG. 1 is a flowchart of a method for screening a voice recognition result according to an example embodiment of the present disclosure;

[0042] Figure 2 FIG. 2 is a flowchart of a method for screening a voice recognition result according to an example embodiment of the present disclosure;

[0043] Figure 3 FIG. 3 is a visual display diagram of a TOP text according to an example embodiment of the present disclosure;

[0044] Figure 4 FIG. 4 is a block diagram of a screening device for a voice recognition result according to an example embodiment of the present disclosure;

[0045] Figure 5 FIG. 5 is a structural diagram of an electronic device according to an example embodiment of the present disclosure. DETAILED DESCRIPTION

[0046] The exemplary embodiments will be described in detail below with reference to the accompanying drawings. In the following description, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present disclosure. Instead, they only describe example devices and methods consistent with some aspects of the present disclosure, as detailed in the appended claims.

[0047] It should be noted that all actions of obtaining signals, information or data in this application are carried out in accordance with the corresponding data protection regulations and policies of the country where the device is located, and with the authorization given by the owner of the corresponding device.

[0048] The speech recognition model needs to continuously find recognition errors in its recognition results in order to focus on optimization in subsequent iterations to improve the recognition accuracy of subsequent speech recognition models.

[0049] In the related art, there is a method of correcting the recognition result by an NLP (Natural Language Processing) model. The NLP model can determine whether the recognition result is correct and correct the recognition result with recognition errors. However, since the NLP model can only correct the text result trained, it cannot be fully covered, for example, some newly added words.

[0050] In addition to correcting the recognition result by the NLP model, the recognition result can also be corrected by a human. However, since the speech recognition model deployed in the application collects a large amount of data, the manual annotation of a large amount of data cannot be normalized in terms of time cost and labor cost.

[0051] To solve the above problems, the present disclosure provides a speech recognition result screening method. Figure 1 is a flowchart of a speech recognition result screening method according to an exemplary embodiment, as Figure 1 shown, the speech recognition result screening method is applied in an electronic device, which can be a mobile phone, a computer, a tablet device, etc. The speech recognition result screening method can include the following steps:

[0052] Step S110, obtaining a plurality of candidate recognition texts output by a preset speech recognition model in a first historical time period, and a first recognition frequency corresponding to each of the plurality of candidate recognition texts.

[0053] The first recognition frequency corresponding to a candidate recognition text represents the number of times that the preset speech recognition model outputs the candidate recognition text in the first historical time period.

[0054] In the embodiments of the present disclosure, the speech recognition model can be an online or offline deployed speech recognition model, for example, a "Xiaoice" speech recognition model.

[0055] The first historical time period can be a time period of one day, one week, one month, etc. before the current time.

[0056] It can be understood that, in order to improve the recognition accuracy of the preset speech recognition model, the speech recognition model can collect the speech recognition input data and the output recognition result of each user each time under the condition of obtaining the authorization of the user. Thus, the electronic device in the embodiments of the present disclosure can obtain the input data and the recognition result of the preset speech recognition model in the first historical time period, so as to obtain the number of times of the input data and the recognition result.

[0057] Further, the electronic device can count the recognition results output by the preset speech recognition model in the first historical time period, and obtain the number of times of recognition corresponding to each recognition result, that is, once the recognition result is recognized as a certain text, the number of times of recognition is counted once.

[0058] For example, in the process of speech recognition, Xiaoice outputs the recognition result "play Tashanhe" for a certain input speech, and accumulates the number of times of recognition corresponding to the recognition result "play Tashanhe" once.

[0059] In some embodiments, for some small-range application speech recognition models, considering that the data amount is small, all the recognition results of the speech recognition model in the first historical time period can be obtained, and the obtained all the recognition results are taken as the obtained multiple candidate recognition texts, and the number of times of recognition corresponding to each candidate recognition result is taken as the first number of times of recognition.

[0060] In some other embodiments, for some large-range application speech recognition models, for example, the "Xiaoice" speech recognition model, the data amount generated every day is huge, thus, all the recognition results of the speech recognition model in the first historical time period can be obtained, and the recognition results are sorted according to the number of times of recognition in the first historical time period, so as to obtain a preset number of recognition results with a high ranking of the number of times of recognition as the obtained multiple candidate recognition texts, and the number of times of recognition corresponding to each candidate recognition result is taken as the first number of times of recognition.

[0061] For example, 10,000 recognition results with a high ranking of the number of times of recognition in the recognition results output by the "Xiaoice" speech recognition model are obtained as the obtained multiple candidate recognition texts.

[0062] In step S120, the plurality of candidate recognition texts are filtered to obtain a first target recognition text identical to the text in the preset corpus. The text in the preset corpus is a recognition-accurate text corresponding to the to-be-recognized speech data.

[0063] In the embodiments of the present disclosure, a corpus, i.e., a preset corpus, can be pre-set to store the recognition-accurate texts corresponding to the to-be-recognized speech data. The to-be-recognized speech data can be input speech data of the preset speech recognition model.

[0064] Optionally, the text in the preset corpus can be a recognition-accurate text corresponding to manual annotation. Optionally, the text in the preset corpus can also be a recognition-accurate text output by a model for correcting the speech recognition result.

[0065] In the embodiments of the present disclosure, after obtaining the plurality of candidate recognition texts, the plurality of candidate recognition texts can be filtered by judging whether each candidate recognition text exists in the preset corpus to obtain the first target recognition text identical to the text in the preset corpus.

[0066] That is, if a candidate recognition text exists in the preset corpus, the candidate recognition text can be determined as the first target recognition text.

[0067] In step S130, for any one of the first target recognition texts, a recognition frequency change parameter corresponding to the first target recognition text is obtained based on a first recognition frequency corresponding to the first target recognition text and a second recognition frequency corresponding to the first target recognition text in a time period before the first historical time period.

[0068] The second recognition frequency corresponding to a first target recognition text represents the number of times that the preset speech recognition model outputs the first target recognition text in the time period before the first historical time period.

[0069] It can be understood that after the foregoing steps, a plurality of first target recognition texts can be obtained. At this time, for any one of the plurality of first target recognition texts, a recognition frequency change parameter corresponding to the first target recognition text can be obtained based on a first recognition frequency corresponding to the first target recognition text and a second recognition frequency corresponding to the first target recognition text.

[0070] The recognition frequency change parameter corresponding to the first target recognition text can reflect the change of the number of times that the first target recognition text has been recognized in history. For example, whether the first target recognition text has been steadily increasing or decreasing in history, whether the first target recognition text has been explosively increasing or decreasing in history, etc.

[0071] Optionally, the recognition frequency change parameter can include at least one of an absolute change value, a relative burst value, and a frequency mutation value.

[0072] In step S140, the to-be-labeled recognition text is determined from the first target recognition texts based on the respective recognition frequency change parameters corresponding to the first target recognition texts.

[0073] The to-be-labeled recognition text can be understood as a recognition text that needs to be labeled additionally, which can be correct or incorrect, and the incorrect recognition text can be labeled with the correct text through labeling.

[0074] In the embodiments of the present disclosure, after obtaining the recognition frequency change parameters corresponding to the first target recognition texts, the to-be-labeled recognition text can be determined from the first target recognition texts based on the respective recognition frequency change parameters corresponding to the first target recognition texts.

[0075] In some embodiments, after the to-be-labeled recognition text is determined from the first target recognition texts, the to-be-labeled recognition text can be labeled. For example, the to-be-labeled recognition text can be labeled through artificial labeling, and it is labeled whether the to-be-labeled recognition text is correct or incorrect, and the correct text is labeled for the incorrect to-be-labeled recognition text.

[0076] In the embodiments of the present disclosure, the to-be-labeled recognition text is determined from the first target recognition texts based on the respective recognition frequency change parameters corresponding to the first target recognition texts. Through the recognition frequency change parameter corresponding to the first target recognition text, the to-be-labeled recognition text is obtained, which can reduce the number of subsequent speech recognition results for labeling, for example, reduce the number of subsequent artificial labeling, save human resources, and on the other hand, the recognition frequency of the correct recognition text is usually stable, if it is not stable, it can be incorrect, therefore, the to-be-labeled recognition text is determined from the first target recognition texts through the recognition frequency change parameter of the correct first target recognition text, which can accurately locate the to-be-labeled recognition text.

[0077] In some embodiments, the time period before the first historical time period comprises a second historical time period before the first historical time period and adjacent to the first historical time period, and the second historical time period has the same length of time as the first historical time period. For example, the first historical time period is yesterday, and the second historical time period is the day before yesterday.

[0078] In this case, in step S130, the identification frequency change parameter corresponding to the first target identified text can be obtained based on the first identification frequency corresponding to the first target identified text and the second identification frequency corresponding to the first target identified text in the second historical time period.

[0079] In this case, in step S130, the identification frequency change parameter corresponding to the first target identified text can be obtained based on the first identification frequency corresponding to the first target identified text and the second identification frequency corresponding to the first target identified text in the second historical time period.

[0080] In the embodiments of the present disclosure, for any one of the first target identified texts, the first identification frequency and the second identification frequency corresponding to the first target identified text can be obtained, and then the identification frequency change parameter corresponding to the first target identified text can be obtained based on the first identification frequency and the second identification frequency corresponding to the first target identified text.

[0081] For example, the first historical time period is yesterday, and the second historical time period is the day before yesterday, i.e., the day before yesterday. At this time, the identification frequency change parameter corresponding to the first target identified text can be obtained based on the identification frequency of the first target identified text in yesterday and the identification frequency of the first target identified text in the day before yesterday.

[0082] In some embodiments, the identification frequency change parameter corresponding to the first target identified text can be obtained based on the first identification frequency corresponding to the first target identified text and the second identification frequency corresponding to the first target identified text in the second historical time period, which can include the following steps:

[0083] determining the difference between the first identification frequency corresponding to the first target identified text and the second identification frequency corresponding to the first target identified text in the second historical time period;

[0084] determining the difference as the identification frequency change parameter corresponding to the first target identified text, and / or determining the ratio of the difference to the second identification frequency corresponding to the first target identified text in the second historical time period as the identification frequency change parameter corresponding to the first target identified text;

[0085] Correspondingly, in step S140, the identified text to be labeled can be determined from the first target identified texts based on the identification frequency change parameters corresponding to the respective first target identified texts, which can include the following steps:

[0086] From each first target recognition text, a preset number of first target recognition texts with a top ranking of the corresponding recognition frequency change parameter are obtained as the to-be-labeled recognition text, and / or a preset number of first target recognition texts with a bottom ranking of the corresponding recognition frequency change parameter are obtained as the to-be-labeled recognition text.

[0087] In the embodiments of the present disclosure, after the first recognition frequency corresponding to the first target recognition text and the second recognition frequency corresponding to the first target recognition text are determined, the difference between the first recognition frequency corresponding to the first target recognition text and the corresponding second recognition frequency can be calculated.

[0088] Here, after the difference between the first recognition frequency corresponding to the first target recognition text and the corresponding second recognition frequency is calculated, the recognition frequency change parameter corresponding to the first target recognition text can be obtained in different ways.

[0089] Alternatively, the difference value can be determined as the recognition frequency change parameter corresponding to the first target recognition text. At this time, the calculated recognition frequency change parameter can also be referred to as an absolute change value. If the absolute change value is greater than 0, the corresponding first target recognition text can be referred to as an absolute increase text. If the absolute change value is less than 0, the corresponding first target recognition text can be referred to as an absolute decrease text.

[0090] Alternatively, the ratio of the difference value to the second recognition frequency corresponding to the first target recognition text in the second historical time period can be determined as the recognition frequency change parameter corresponding to the first target recognition text. At this time, the calculated recognition frequency change parameter can also be referred to as a relative burst value. If the relative burst value is greater than 0, the corresponding first target recognition text can be referred to as a relative high burst text. If the relative burst value is less than 0, the corresponding first target recognition text can be referred to as a relative low burst text.

[0091] Alternatively, the difference value and the ratio of the difference value to the second recognition frequency corresponding to the first target recognition text in the second historical time period are both determined as the recognition frequency change parameter corresponding to the first target recognition text.

[0092] In the embodiments of the present disclosure, when the difference value is determined as the recognition frequency change parameter corresponding to the first target recognition text, and / or the ratio of the difference value to the second recognition frequency corresponding to the first target recognition text in the second historical time period is determined as the recognition frequency change parameter corresponding to the first target recognition text, further, based on the respective recognition frequency change parameters corresponding to each first target recognition text, the to-be-labeled recognition text can be determined from each first target recognition text in multiple different ways.

[0093] Optionally, a preset number of first target recognition texts with the highest corresponding recognition frequency change parameters can be selected from each first target recognition text as the recognition texts to be labeled.

[0094] Optionally, a preset number of first target recognition texts with the corresponding recognition frequency change parameter sorted last can be selected from each first target recognition text as the recognition text to be labeled.

[0095] Optionally, a preset number of first target recognition texts with the highest corresponding recognition frequency change parameters and a preset number of first target recognition texts with the lowest corresponding recognition frequency change parameters can be selected from each first target recognition text as the recognition texts to be labeled.

[0096] In other words, in this embodiment of the disclosure, the text to be labeled determined from each first target identification text may include at least one of the following:

[0097] After calculating the difference between the first recognition count and the second recognition count for each first target recognition text, a preset number of first target recognition texts with the highest corresponding difference counts can be selected from the differences for each first target recognition text as the texts to be labeled.

[0098] From the differences corresponding to each first target recognition text, a preset number of first target recognition texts with the lowest corresponding difference can be selected as the texts to be labeled.

[0099] After calculating the ratio of the difference between each first target recognition text and its corresponding second recognition count, a preset number of first target recognition texts with the highest ratios can be selected from the ratios corresponding to each first target recognition text as the texts to be labeled.

[0100] From the ratios corresponding to each first target recognition text, a preset number of first target recognition texts with the lowest ratios can be selected as the texts to be labeled.

[0101] It should be noted that, considering the possibility that the second recognition count might be 0, in this case, the sum of the second recognition count and a preset value can be obtained, and then the ratio of the difference to the sum can be determined as the recognition count change parameter corresponding to the first target text. For example, the preset value could be 1.

[0102] In some embodiments, the first historical time period corresponds to a preset time length, the time period before the first historical time period includes a second historical time period before the first historical time period and adjacent to the first historical time period, and the second historical time period has a preset time length, and the time period before the first historical time period includes a third historical time period before the first historical time period and adjacent to the first historical time period, and the third historical time period is greater than the preset time length. For example, the preset time length can be one day, one week, one month, etc., which is determined according to actual needs.

[0103] For example, assuming that the preset time length is one day and the first historical time period is yesterday, the second historical time period is the day before yesterday, and the third historical time period can be five days, ten days, thirty days, etc. including yesterday, or all historical time periods including yesterday, which is determined according to actual needs.

[0104] In this case, in step S130, the identification frequency change parameter corresponding to the first target recognition text can be obtained based on the first identification frequency corresponding to the first target recognition text and the second identification frequency corresponding to the first target recognition text.

[0105] The identification frequency change parameter corresponding to the first target recognition text can be obtained based on the first identification frequency corresponding to the first target recognition text, the second identification frequency corresponding to the first target recognition text in the second historical time period, and the third identification frequency corresponding to the first target recognition text in the third historical time period.

[0106] In the embodiments of the present disclosure, for any one of the first target recognition texts, the first identification frequency, the second identification frequency, and the third identification frequency corresponding to the first target recognition text can be obtained, and then the identification frequency change parameter corresponding to the first target recognition text can be obtained based on the first identification frequency, the second identification frequency, and the third identification frequency corresponding to the first target recognition text.

[0107] In some embodiments, the identification frequency change parameter corresponding to the first target recognition text can be obtained based on the first identification frequency corresponding to the first target recognition text, the second identification frequency corresponding to the first target recognition text in the second historical time period, and the third identification frequency corresponding to the first target recognition text in the third historical time period.

[0108] The product of the first value and the second value is determined as the identification frequency change parameter corresponding to the first target recognition text, the first value is the ratio of the first identification frequency corresponding to the first target recognition text and the third identification frequency, and the second value is obtained based on the second identification frequency corresponding to the first target recognition text and a preset logarithmic function.

[0109] Correspondingly, in step S140, determining the to-be-labeled recognition text from each first target recognition text based on the respective recognition frequency change parameter corresponding to each first target recognition text can include the following steps:

[0110] From each first target recognition text, a preset number of first target recognition texts with a high ranking in the corresponding recognition frequency change parameter are obtained as the to-be-labeled recognition text.

[0111] In the embodiments of the present disclosure, the ratio of the first recognition frequency corresponding to each first target recognition text to the third recognition frequency corresponding thereto, i.e., the first value, can be obtained, and in addition, the second value corresponding to each first target recognition text can be obtained based on the second recognition frequency corresponding to each first target recognition text and a preset logarithmic function. Then, the product of the first value and the second value corresponding to each first target recognition text can be determined as the recognition frequency change parameter corresponding to each first target recognition text. At this time, the calculated recognition frequency change parameter can also be referred to as a frequency mutation value.

[0112] In some embodiments, the preset logarithmic function can be, for example, a log function, or an lg function, or an ln function.

[0113] In the embodiments of the present disclosure, after obtaining the respective recognition frequency change parameter corresponding to each first target recognition text, a preset number of first target recognition texts with a high ranking in the corresponding recognition frequency change parameter can be obtained from each first target recognition text as the to-be-labeled recognition text.

[0114] In addition, in some embodiments, the method of the embodiments of the present disclosure can further include the following steps:

[0115] According to the respective first recognition frequency corresponding to each candidate recognition text, a preset number of third target recognition texts with a high ranking in the corresponding first recognition frequency are determined from the plurality of candidate recognition texts;

[0116] According to the third target recognition text, a to-be-processed recognition text is determined.

[0117] The to-be-processed recognition text can be understood as a recognition text that needs to be focused on.

[0118] In the embodiments of the present disclosure, the respective first recognition frequency corresponding to each candidate recognition text can be obtained, and then a preset number of third target recognition texts with a high ranking in the corresponding first recognition frequency can be determined from the plurality of candidate recognition texts, and then a to-be-processed recognition text can be determined according to the third target recognition text.

[0119] With the foregoing example, assuming that the top 10,000 recognition results in the recognition result output by the Xiaoai speech recognition model are obtained as the plurality of candidate recognition texts, a preset number of third target recognition texts corresponding to the top recognition results can be selected from the 10,000 candidate recognition texts, and the third target recognition texts are also used as the to-be-processed texts.

[0120] For example, for the Xiaoai speech recognition model, the recognition text Xiaoai is a high-frequency text requested by an online sound box user, and if the text is not in the third target recognition text in the data in the first historical time period, it needs to be focused on.

[0121] In the embodiments of the present disclosure, the to-be-processed recognition text is determined according to the third target recognition text, which can stably monitor the high-frequency data in the first historical time period and help quickly mine the anomaly.

[0122] In some embodiments, the method of the embodiments of the present disclosure can further include the following steps:

[0123] The plurality of candidate recognition texts are screened to obtain second target recognition texts that are different from the texts in the preset corpus, and the second target recognition texts are used for manual recognition annotation.

[0124] In the embodiments of the present disclosure, considering that the texts in the preset corpus are accurate recognition texts corresponding to the to-be-recognized speech data, if the candidate recognition text is not in the preset corpus, it may be an incorrect recognition text. Therefore, after obtaining the plurality of candidate recognition texts, the plurality of candidate recognition texts can be screened to obtain second target recognition texts that are different from the texts in the preset corpus, and the obtained second target recognition texts can be used for manual recognition annotation. Compared with manually annotating all candidate recognition texts, the workload of manual annotation can be reduced.

[0125] In some embodiments, considering the situation of insufficient manpower, a preset number of second target recognition texts corresponding to the top recognition times can be selected from the second target recognition texts for manual screening. In the case of insufficient manpower, selecting second target recognition texts with more recognition times for manual annotation can improve the efficiency of manual annotation while ensuring high correction effect.

[0126] In some embodiments, the method of the embodiments of the present disclosure can further include the following steps:

[0127] An accurate recognition text corresponding to the speech data of the second target recognition text is obtained;

[0128] The accurate recognition text is added to the preset corpus.

[0129] In this embodiment of the disclosure, as can be seen from the foregoing, after obtaining the second target recognition text, the second target recognition text can be manually identified and annotated to obtain the accurately recognized text of the speech data corresponding to the second target recognition text. Then, the accurately recognized text can be added to the preset corpus to continuously update the preset corpus. By continuously updating the preset corpus, the preset corpus can be continuously enriched and expanded, thereby filtering the candidate recognition text collected subsequently, reducing the amount of speech data that needs to be manually annotated, and saving human resources.

[0130] It should be noted that the method in this embodiment can be executed continuously over time, thereby continuously enriching the preset corpus and continuously mining the text to be labeled and recognized in the process.

[0131] It should be noted that, in the embodiments of this disclosure, the preset quantity can be set according to actual needs, and the preset quantities can be the same or different.

[0132] Below, in conjunction with Figure 2 The flowchart shown illustrates a specific embodiment of the method for filtering speech recognition results according to this disclosure.

[0133] like Figure 2 As shown, the electronic device can acquire the "Xiao Ai" voice recognition of the user's voice throughout yesterday. The recognized texts are sorted according to the number of times each text has been recognized, and the top 10,000 texts are selected as candidate texts. From these 10,000 candidate texts, the top 100 texts with the highest recognition frequency are selected as the third target text, also known as the full TOP text. For ease of description, yesterday will be referred to as Day 1.

[0134] In addition, the currently accumulated preset corpus can be obtained, and based on whether the candidate recognition text exists in the preset corpus, the 10,000 candidate recognition texts can be divided into first target recognition texts that are in the preset corpus and second target recognition texts that are not in the preset corpus.

[0135] For the second target recognition texts obtained through screening, the top 100 second target recognition texts ranked by recognition frequency can be selected as TOP texts not in the preset corpus and manually annotated.

[0136] For the selected target texts, we can obtain the number of times each target text was identified the day before yesterday, and the number of times each target text was identified the previous 7 days, including yesterday. For ease of description, we will refer to the day before yesterday as day 2, and the previous 7 days, including yesterday, as last week.

[0137] Subsequently, various recognition frequency change parameters corresponding to each first target text can be calculated according to the obtained various recognition frequencies. Taking the various recognition frequency change parameters corresponding to each first target text as an example, the calculation method is as follows:

[0138] Absolute change value = recognition frequency on the first day - recognition frequency on the second day;

[0139] Relative burst value = [(recognition frequency on the first day - recognition frequency on the second day) / (recognition frequency on the second day + 1)]%;

[0140] Frequency mutation value = (recognition frequency on the first day / last week's recognition frequency)*log(recognition frequency on the second day);

[0141] Through the above calculation process, the absolute change value, the relative burst value, and the frequency mutation value corresponding to each first target text can be obtained. Then, the top 100 first target recognition texts in the front of the sorting according to the absolute change value from large to small can be selected as the absolute increase TOP texts, and the top 100 first target recognition texts in the back of the sorting can be selected as the absolute decrease TOP texts.

[0142] The top 100 first target recognition texts in the front of the sorting according to the relative burst value from large to small can be selected as the relative high TOP texts, and the top 100 first target recognition texts in the back of the sorting can be selected as the relative low TOP texts.

[0143] The top 100 first target recognition texts in the front of the sorting according to the frequency mutation value from large to small can be selected as the frequency mutation TOP texts.

[0144] Through the above process, 100 full amount TOP texts, 100 TOP texts not in the preset corpus, 100 absolute increase TOP texts, 100 absolute decrease TOP texts, 100 relative high TOP texts, 100 relative low TOP texts, and 100 frequency mutation TOP texts can be obtained. The recognition frequency change parameters can include the absolute change value, the relative burst value, and the frequency mutation value.

[0145] Finally, the obtained 700 texts can be manually annotated. It should be noted that there can be repeated texts in the obtained 700 texts.

[0146] In addition, in some embodiments, as shown in FIG. 13, Figure 3 the various TOP texts obtained can be visually displayed in an interface. It should be noted that, due to the length of the paper, Figure 3 the first 12 of the various TOP texts are displayed in FIG. 13.

[0147] Exemplarily, Figure 3The data on November 08, 2022 is shown as Figure 3 As shown, the full amount of TOP text shown corresponds to the number of recognitions on November 08, 2022, the number of recognitions on November 08, 2022 is shown for the text not in the preset corpus TOP, the absolute change value (positive integer) on November 08, 2022 is shown for the absolute increase TOP text, the absolute change value (negative integer) on November 08, 2022 is shown for the absolute decrease TOP text, the relative sudden value (positive percentage) on November 08, 2022 is shown for the relative sudden TOP text, the relative sudden value (negative percentage) on November 08, 2022 is shown for the relative sudden TOP text, and the frequency mutation value on November 08, 2022 is shown for the frequency mutation TOP text.

[0148] Figure 4 is a block diagram of a voice recognition result screening device according to an exemplary embodiment. Please refer to Figure 4 The voice recognition result screening device 400 is applied to an electronic device, and the voice recognition result screening device 400 comprises:

[0149] The acquisition module 410 is configured to acquire a plurality of candidate recognition texts output by a preset voice recognition model in a first historical time period, and a first recognition number corresponding to each of the plurality of candidate recognition texts. The first recognition number corresponding to a candidate recognition text represents the number of times that the preset voice recognition model outputs the candidate recognition text in the first historical time period.

[0150] The screening module 420 is configured to screen the plurality of candidate recognition texts to obtain a first target recognition text identical to a text in a preset corpus. The text in the preset corpus is a recognition-accurate text corresponding to the to-be-recognized voice data.

[0151] The calculation module 430 is configured to, for any one of the first target recognition texts, obtain a recognition number change parameter corresponding to the first target recognition text based on the first recognition number corresponding to the first target recognition text and a second recognition number corresponding to the first target recognition text. The second recognition number corresponding to a first target recognition text represents the number of times that the preset voice recognition model outputs the first target recognition text in a time period before the first historical time period.

[0152] The determination module 440 is configured to determine a to-be-labeled recognition text from each of the first target recognition texts based on the recognition number change parameter corresponding to each of the first target recognition texts.

[0153] Optionally, the time period before the first historical time period comprises a second historical time period before the first historical time period and adjacent to the first historical time period, and the second historical time period has the same time length as the first historical time period. The computing module 430 comprises:

[0154] The first computing submodule is configured to obtain, based on the first recognition frequency corresponding to the first target recognized text and the second recognition frequency corresponding to the first target recognized text within the second historical time period, a recognition frequency change parameter corresponding to the first target recognized text.

[0155] Optionally, the computing submodule comprises:

[0156] The first computing unit is configured to determine a difference between the first recognition frequency corresponding to the first target recognized text and the second recognition frequency corresponding to the first target recognized text within the second historical time period.

[0157] The second computing unit is configured to determine, as the recognition frequency change parameter corresponding to the first target recognized text, the difference and / or a ratio of the difference to the second recognition frequency corresponding to the first target recognized text within the second historical time period.

[0158] Correspondingly, the determining module 440 is further configured to obtain, from each of the first target recognized texts, a preset number of first target recognized texts with a higher ranking of the recognition frequency change parameter as the to-be-labeled recognized text, and / or a preset number of first target recognized texts with a lower ranking of the recognition frequency change parameter as the to-be-labeled recognized text.

[0159] Optionally, the first historical time period corresponds to a preset time length, the time period before the first historical time period comprises a second historical time period before the first historical time period and adjacent to the first historical time period, and the second historical time period has the preset time length, and comprises a third historical time period before the first historical time period and adjacent to the first historical time period, and the third historical time period is longer than the preset time length. The computing module 430 comprises:

[0160] The second computing submodule is configured to obtain, based on the first recognition frequency corresponding to the first target recognized text, the second recognition frequency corresponding to the first target recognized text within the second historical time period, and the third recognition frequency corresponding to the first target recognized text within the third historical time period, a recognition frequency change parameter corresponding to the first target recognized text.

[0161] Optionally, the second computing submodule comprises:

[0162] The third computing unit is configured to determine a product of a first value and a second value as the identification frequency change parameter corresponding to the first target identified text, the first value being a ratio of the first identification frequency corresponding to the first target identified text and a third identification frequency, and the second value being obtained based on the second identification frequency corresponding to the first target identified text and a preset logarithmic function.

[0163] Correspondingly, the determining module 440 is further configured to obtain a preset number of first target identified texts with a higher ranking of the identification frequency change parameter from the first target identified texts as the to-be-labeled identified text.

[0164] Optionally, the apparatus 400 further comprises:

[0165] The third target identified text determining module is configured to determine a preset number of third target identified texts with a higher ranking of the first identification frequency from the plurality of candidate identified texts according to the first identification frequency corresponding to each of the plurality of candidate identified texts.

[0166] The to-be-processed identified text determining module is configured to determine the to-be-processed identified text according to the third target identified text.

[0167] Optionally, the apparatus 400 further comprises:

[0168] The second target identified text screening module is configured to screen the plurality of candidate identified texts to obtain second target identified texts different from texts in a preset corpus, the second target identified texts being used for manual identification and labeling.

[0169] Optionally, the apparatus 400 further comprises:

[0170] The identified accurate text obtaining module is configured to obtain identified accurate texts of the voice data corresponding to the second target identified texts.

[0171] The adding module is configured to add the identified accurate texts to the preset corpus.

[0172] As to the screening apparatus of the voice recognition result in the above-mentioned embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments of the method, and will not be described in detail here.

[0173] The present disclosure further provides a computer readable storage medium having computer program instructions stored thereon, the program instructions being executed by a processor to implement the steps of the screening method of the voice recognition result provided by the present disclosure.

[0174] Figure 5is a block diagram of an electronic device for a screening method of a speech recognition result according to an exemplary embodiment. The electronic device 500 can be, for example, a mobile phone, a computer, a tablet device, etc.

[0175] Referring to Figure 5 The electronic device 500 can include one or more of the following components: a processing component 502, a memory 504, a power supply component 506, a multimedia component 508, an audio component 510, an input / output (I / O) interface 512, a sensor component 514, and a communication component 516.

[0176] The processing component 502 usually controls overall operations of the electronic device 500, such as operations associated with displaying, making phone calls, data communications, camera operations, and recording operations. The processing component 502 can include one or more processors 520 to execute instructions to complete all or part of steps of the above-described methods. In addition, the processing component 502 can include one or more modules to facilitate interaction between the processing component 502 and other components. For example, the processing component 502 can include a multimedia module to facilitate the interaction between the multimedia component 508 and the processing component 502.

[0177] The memory 504 is configured to store various types of data to support operations of the electronic device 500. Examples of these data include instructions for any application or method operating on the electronic device 500, contact data, phonebook data, messages, pictures, videos, etc. The memory 504 can be implemented by any type of volatile or non-volatile storage devices or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.

[0178] The power supply component 506 supplies electric power for various components of the electronic device 500. The power supply component 506 can include a power management system, one or more power supplies, and other components associated with generating, managing and distributing electric power for the electronic device 500.

[0179] The multimedia component 508 includes a screen to provide an output interface between the electronic device 500 and a user. In some embodiments, the screen can include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive an input signal from a user. The touch panel includes one or more touch sensors to sense a touch, a slide, and a gesture on the touch panel. The touch sensor can not only sense a boundary of a touching or a sliding action, but also detect duration and intensity of the touching or sliding action. In some embodiments, the multimedia component 508 includes a front camera and / or a rear camera. When the electronic device 500 is in an operating mode, such as a camera mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zooming capability.

[0180] The audio component 510 is configured to output and / or input an audio signal. For example, the audio component 510 includes a microphone (MIC) to receive an external audio signal when the electronic device 500 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in the memory 504 or transmitted via the communication component 516. In some embodiments, the audio component 510 further includes a speaker to output an audio signal.

[0181] The input / output interface 512 provides an interface between the processing component 502 and peripheral interface modules, which can be a keypad, a click wheel, buttons, and the like. The buttons can include, but are not limited to, a home button, a volume button, a start button, and a lock button.

[0182] The sensor component 514 includes one or more sensors to provide various state assessments for the electronic device 500. For example, the sensor component 514 can detect an open / closed state of the electronic device 500, relative positioning of components, such as a display and a keypad of the electronic device 500, a change in position of the electronic device 500 or a component of the electronic device 500, presence or absence of user contact with the electronic device 500, an orientation or acceleration / deceleration of the electronic device 500, and a temperature change of the electronic device 500. The sensor component 514 can include a proximity sensor to detect presence of an object within a proximity range of the electronic device 500 without any physical contact. The sensor component 514 can further include a light sensor, such as a CMOS or CCD image sensor, to use in an imaging application. In some embodiments, the sensor component 514 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0183] The communication component 516 is configured to facilitate wired or wireless communication between the electronic device 500 and other devices. The electronic device 500 can access a wireless network based on a communication standard, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 516 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 516 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technology.

[0184] In an exemplary embodiment, the electronic device 500 can be implemented with one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, micro-controllers, microprocessors, or other electronic elements, for performing the above-described method of filtering a speech recognition result.

[0185] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions, such as the memory 504 including instructions, is also provided, which can be executed by the processor 520 of the electronic device 500 to implement the above-described method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disc, and an optical data storage device, etc.

[0186] In another exemplary embodiment, a computer program product is also provided, which contains a computer program capable of being executed by a programmable device, and the computer program has code portions for executing the above-described method of filtering a speech recognition result when executed by the programmable device.

[0187] Other embodiments of the present disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the present disclosure. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure following the general principles thereof and including modifications and equivalents of the present disclosure. The specification and examples are to be regarded as illustrative only, and the true scope and spirit of the present disclosure are indicated by the following claims.

[0188] It should be understood that the present disclosure is not limited to the precise structures as set forth above and shown in the attached drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is indicated by the appended claims, rather than by the description.

Claims

1. A method for filtering speech recognition results, characterized in that, The method includes: Obtain multiple candidate recognition texts output by the preset speech recognition model within a first historical time period, and the first recognition counts corresponding to each of the multiple candidate recognition texts. The first recognition counts corresponding to a candidate recognition text represent the number of times the preset speech recognition model outputs the candidate recognition text within the first historical time period. The multiple candidate texts are filtered to obtain the first target text that is the same as the text in the preset corpus, where the text in the preset corpus is the accurately recognized text corresponding to the speech data to be recognized. For any one of the first target recognition texts, based on the first recognition count corresponding to the first target recognition text and the second recognition count corresponding to the first target recognition text, the recognition count change parameter corresponding to the first target recognition text is obtained. The second recognition count corresponding to a first target recognition text represents the number of times the preset speech recognition model outputs the first target recognition text in the time period before the first historical time period. Based on the recognition frequency variation parameters corresponding to each of the first target recognition texts, the text to be labeled is determined from each of the first target recognition texts.

2. The method according to claim 1, characterized in that, The time periods preceding the first historical time period include the second historical time periods that are adjacent to the first historical time period and have the same duration. The step of obtaining the change parameter of the number of recognitions corresponding to the first target recognition text based on the first recognition count and the second recognition count corresponding to the first target recognition text includes: Based on the first recognition count corresponding to the first target recognition text and the second recognition count corresponding to the first target recognition text within the second historical time period, the change parameter of the recognition count corresponding to the first target recognition text is obtained.

3. The method according to claim 2, characterized in that, The step of obtaining the change parameter of the number of recognitions corresponding to the first target recognition text based on the first recognition count corresponding to the first target recognition text and the second recognition count corresponding to the first target recognition text within the second historical time period includes: Determine the difference between the first recognition count corresponding to the first target recognition text and the second recognition count corresponding to the first target recognition text within the second historical time period; The difference is determined as the change parameter of the number of recognitions corresponding to the first target recognition text, and / or the ratio of the difference to the second number of recognitions corresponding to the first target recognition text in the second historical time period is determined as the change parameter of the number of recognitions corresponding to the first target recognition text; The step of determining the text to be labeled from each of the first target recognition texts based on the recognition count variation parameter corresponding to each of the first target recognition texts includes: From each of the first target recognition texts, a preset number of first target recognition texts with the highest sorting order of the recognition count change parameter are obtained as the recognition texts to be labeled, and / or a preset number of first target recognition texts with the lowest sorting order of the recognition count change parameter are obtained as the recognition texts to be labeled.

4. The method according to claim 1, characterized in that, The first historical time period corresponds to a preset duration. The time periods preceding the first historical time period include a second historical time period preceding the first historical time period that is adjacent to the first historical time period and has the preset duration, and a third historical time period preceding the first historical time period that is adjacent to the first historical time period and has a duration longer than the preset duration. The step of obtaining the recognition count change parameter corresponding to the first target recognition text based on the first recognition count corresponding to the first target recognition text and the second recognition count corresponding to the first target recognition text includes: Based on the first recognition count corresponding to the first target recognition text, the second recognition count corresponding to the first target recognition text in the second historical time period, and the third recognition count corresponding to the first target recognition text in the third historical time period, the recognition count change parameter corresponding to the first target recognition text is obtained.

5. The method according to claim 4, characterized in that, The method of obtaining the recognition frequency change parameter corresponding to the first target recognition text based on the first recognition frequency corresponding to the first target recognition text, the second recognition frequency corresponding to the first target recognition text in the second historical time period, and the third recognition frequency corresponding to the first target recognition text in the third historical time period includes: The product of the first value and the second value is determined as the parameter for the change in the number of recognitions corresponding to the first target recognition text. The first value is the ratio of the first recognition count and the corresponding third recognition count to the first target recognition text. The second value is obtained based on the second recognition count to the first target recognition text and a preset logarithmic function. The step of determining the text to be labeled from each of the first target recognition texts based on the recognition count variation parameter corresponding to each of the first target recognition texts includes: From each of the first target recognition texts, a preset number of first target recognition texts that are ranked first in the corresponding recognition count change parameter are obtained as the recognition texts to be labeled.

6. The method according to claim 1, characterized in that, The method further includes: Based on the first recognition count corresponding to each of the plurality of candidate recognition texts, a preset number of third target recognition texts with the highest first recognition counts are determined from the plurality of candidate recognition texts; Based on the third target identified text, the text to be processed is determined.

7. The method according to claim 1, characterized in that, The method further includes: The multiple candidate texts are filtered to obtain a second target text that is different from the text in the preset corpus. The second target text is used for manual identification and annotation.

8. The method according to claim 7, characterized in that, The method further includes: Obtain the accurately recognized text from the speech data corresponding to the second target text; The accurately identified text is added to the preset corpus.

9. A device for filtering speech recognition results, characterized in that, The device includes: The acquisition module is configured to acquire multiple candidate recognition texts output by a preset speech recognition model within a first historical time period, and the first recognition counts corresponding to the multiple candidate recognition texts respectively. The first recognition counts corresponding to a candidate recognition text represent the number of times the preset speech recognition model outputs the candidate recognition text within the first historical time period. The filtering module is configured to filter the plurality of candidate recognition texts to obtain a first target recognition text that is the same as the text in the preset corpus, wherein the text in the preset corpus is the accurately recognized text corresponding to the speech data to be recognized; The calculation module is configured to, for any one of the first target recognition texts, obtain a recognition count change parameter corresponding to the first target recognition text based on the first recognition count corresponding to the first target recognition text and the second recognition count corresponding to the first target recognition text. The second recognition count corresponding to a first target recognition text represents the number of times the preset speech recognition model outputs the first target recognition text in the time period before the first historical time period. The determination module is configured to determine the text to be labeled from each of the first target recognition texts based on the recognition count change parameter corresponding to each of the first target recognition texts.

10. An electronic device, characterized in that, The electronic device includes: processor; Memory used to store processor-executable instructions; The processor is configured to implement the steps of the method according to any one of claims 1 to 8 when executing.

11. A computer-readable storage medium storing computer program instructions thereon, characterized in that, When executed by a processor, the program instructions implement the steps of the method described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Voice recognition model generation method and device, storage medium and electronic equipment

    CN108847222A

  • Text recognition method and device, electronic equipment and storage medium

    CN115291791A