Data sampling inspection method and device, equipment and storage medium

By dividing text groups based on the degree of difference and drawing sample texts from each group, the problem that the sampling samples in the prior art cannot accurately represent the overall data quality, and more accurate data sampling results are achieved.

CN119962528APending Publication Date: 2025-05-09IFLYTEK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411861959.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-17
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

When selecting samples, existing data sampling methods may not accurately represent the overall data quality, resulting in inaccurate sampling results.

Method used

By obtaining the recognition text of the target speech, the text proofreading process is performed to obtain the proofreading text, and then the proofreading text is divided into multiple first proofreading text groups based on the degree of difference between the recognition text and the proofreading text, and the text text is sampled from each group.

Benefits of technology

This method can effectively reduce the situation where the sample text is concentrated in a first proofreading text group, ensuring that the quality of the sample text drawn can accurately represent the quality of several proofreading texts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119962528A_ABST
    Figure CN119962528A_ABST
Patent Text Reader

Abstract

The invention discloses a data sampling inspection method and device, equipment and a storage medium. The method comprises the following steps: acquiring recognition texts corresponding to a plurality of target voices respectively; performing text proofreading processing on each recognition text to obtain a corresponding proofreading text; dividing the plurality of proofreading texts into a plurality of first proofreading text groups based on the difference degree between each recognition text and the corresponding proofreading text; and extracting a proofreading text from each first proofreading text group to serve as a sample text. In this way, the sampling accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a data sampling method, device, equipment and storage medium. Background Art

[0002] Currently, when faced with a large amount of data, a small number of samples are usually selected from the large amount of data through random sampling, and the quality test results of the small number of samples are used as the quality test results of the large amount of data. In some cases, improper random sampling methods may cause the sample to fail to represent the quality of the overall data. For example, the quality of the sample is qualified, but the quality of the large amount of data is unqualified. Therefore, how to improve the random sampling method to improve the accuracy of the sample is a technical problem that needs to be solved urgently. Summary of the invention

[0003] The main technical problem solved by the present application is to provide a data sampling method, device, equipment and storage medium, which can make the quality of the extracted sample text accurately represent the quality of several proofreading texts.

[0004] In order to solve the above technical problems, a technical solution adopted in the present application is: to provide a data sampling method, the method comprising: obtaining recognition texts corresponding to several target voices respectively; performing text proofreading processing on each recognition text to obtain a corresponding proofreading text; based on the degree of difference between each recognition text and the corresponding proofreading text, dividing the several proofreading texts into multiple first proofreading text groups; extracting proofreading text from each first proofreading text group as a sample text.

[0005] In order to solve the above technical problems, another technical solution adopted in the present application is: to provide a data sampling device, including: an acquisition module, a proofreading module, a division module and an extraction module; the acquisition module is used to obtain recognition texts corresponding to several target voices respectively; the proofreading module is used to perform text proofreading processing on each recognition text to obtain a corresponding proofreading text; the division module is used to divide the several proofreading texts into multiple first proofreading text groups based on the degree of difference between each recognition text and the corresponding proofreading text; the extraction module is used to extract proofreading texts from each first proofreading text group as sample text.

[0006] To solve the above technical problems, another technical solution adopted in the present application is: to provide an electronic device, comprising a memory and a processor coupled to each other, the memory storing program instructions; the processor is used to execute the program instructions stored in the memory to implement the above method.

[0007] In order to solve the above technical problem, another technical solution adopted by the present application is: providing a computer-readable storage medium for storing program instructions, which can be executed to implement the above method.

[0008] The above scheme first divides a number of proofreading texts into a number of first proofreading text groups based on the degree of difference between each recognized text and the corresponding proofreading text, and then extracts sample texts from each of the divided first proofreading text groups. Compared with the method of directly sampling from a number of proofreading texts without dividing the text groups based on the degree of difference, the present application extracts sample texts from each of the first proofreading text groups obtained by the division, which can effectively reduce the situation where the sample texts are concentrated in a certain first proofreading text group. Therefore, the quality of the sample texts extracted by the present application can accurately represent the quality of a number of proofreading texts.

[0009] Furthermore, after dividing the text groups based on the degree of difference, proofreading texts with similar degree of difference are divided into the same group, so the sample text extracted from each group can better represent the overall text of the group. Compared with the method of directly sampling from a number of proofreading texts without dividing the text groups based on the degree of difference, the above method of the present application can reduce the situation where the sample text extracted cannot accurately reflect a number of proofreading texts due to the large degree of difference. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 It is a flow chart of an embodiment of a data sampling method provided by the present application;

[0011] Figure 2 yes Figure 1 The flowchart of step S14 is shown as an embodiment;

[0012] Figure 3 yes Figure 1 The flowchart of another embodiment of step S14 is shown;

[0013] Figure 4 It is a schematic diagram of the framework of an embodiment of a data sampling device provided by the present application;

[0014] Figure 5 It is a schematic diagram of a framework of an embodiment of an electronic device provided by the present application;

[0015] Figure 6 It is a schematic diagram of the framework of the computer-readable storage medium provided by this application. DETAILED DESCRIPTION

[0016] In order to make the purpose, technical solution and effect of the present application clearer and more specific, the present application is further described in detail below with reference to the accompanying drawings and examples.

[0017] In addition, if there are descriptions involving "first", "second", etc. in the embodiments of the present application, the descriptions of "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or suggesting their relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In addition, the technical solutions between the various embodiments can be combined with each other, but they must be based on the ability of ordinary technicians in the field to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such combination of technical solutions does not exist and is not within the scope of protection required by this application.

[0018] See also Figure 1 , Figure 1 It is a flow chart of an embodiment of the data sampling method provided by the present application. It should be noted that if there are substantially the same results, this embodiment does not Figure 1 The process sequence shown is limited. Figure 1 As shown, this embodiment includes:

[0019] S11: Obtain recognition texts corresponding to a number of target voices.

[0020] This embodiment is used to divide a number of proofreading texts into a plurality of first proofreading text groups based on the degree of difference between each recognized text and the corresponding proofreading text, so as to divide the proofreading texts with similar degrees of difference into the same group, and then extract sample texts that can better represent the overall text of the group from each first proofreading text group. Furthermore, each first proofreading text group has proofreading text as sample text, which can effectively reduce the situation where sample texts are concentrated in a certain first proofreading text group, so that the quality of the sample text extracted by this application can accurately represent the quality of a number of proofreading texts.

[0021] The target speech in this embodiment can be a short audio with an audio duration of a few seconds, or a long audio with an audio duration of several minutes or even more than ten minutes.

[0022] In one implementation scenario, considering that the text corresponding to a long speech with a long speech duration is long, when the proofreading text of the recognition text corresponding to the entire long speech is sampled as a single proofreading text, there may be long speech that is not sampled, resulting in a narrow coverage of the sample text. Based on this, the long speech can be divided into multiple short speech segments, and each short speech segment is used as each target speech.

[0023] Therefore, in this implementation scenario, the speech duration of each target speech is a speech segment less than or equal to the preset duration, and when the original speech is a long audio, the original speech is divided into multiple short speech segments with a speech duration less than or equal to the preset duration. The specific preset duration can be adjusted according to the actual sampling effect or actual needs.

[0024] Among them, a speech recognition model (or transcription engine) or manual labor can be used to divide a long speech into multiple speech segments according to the pause position of the speech or the switching position of the speaker, and each speech segment is used as a target speech.

[0025] Furthermore, considering that the speech recognition model or transcription engine may have inaccurate segmentation, when the segmentation position of each speech segment is manually adjusted, the manually adjusted segmentation position is directly used as the actual segmentation position, and the long speech is segmented to obtain each speech segment.

[0026] In this embodiment, the recognized text is obtained by performing speech recognition on the target speech using a speech recognition model or a transcription engine.

[0027] S12: Perform text proofreading on each recognized text to obtain a corresponding proofread text.

[0028] In one embodiment, performing text proofreading on each recognized text includes: performing at least one of text replacement, text deletion, and new text insertion on each recognized text.

[0029] Character replacement is to replace a character in the recognized text with another character, character deletion is to delete the existing characters in the recognized text, and new character insertion is to add new characters between two characters in the recognized text.

[0030] For example, the recognized text is "a review article with relatively good machine learning", and the proofread text after proofreading is "a review article with relatively good machine learning". For another example, the recognized text is "this is a book about someone who has maintained strong hearing under the oppression of the northern wind and snow", and the proofread text after proofreading is "this is a tree that has remained stubbornly upright under the oppression of the northern wind and snow".

[0031] In one implementation scenario, specifically in the annotation scenario, the proofreading text is the annotated text about the target speech, that is, text proofreading is performed on the basis of the recognized text to identify and modify the erroneous content in the recognized text, and the proofreading text after the proofreading is used as the annotated text.

[0032] Generally, the annotated text is the annotated text of the training sample used for model training. It is understandable that the higher the quality of the annotated text, the more conducive it is to training a high-quality model. Therefore, this application needs to select sample texts that can accurately reflect the quality of several annotated texts from each proofreading text (annotated text).

[0033] In this embodiment, the proofreading of the recognized text may be performed by proofreading (modifying) erroneous characters in the recognized text using a relevant model (such as a language / semantic model), or by proofreading (modifying) erroneous characters in the recognized text using manual proofreading.

[0034] S13: Based on the difference between each recognized text and the corresponding proofreading text, the plurality of proofreading texts are divided into a plurality of first proofreading text groups.

[0035] In one embodiment, the calculated semantic similarity between the recognized text and the corresponding proofread text may be used to measure the degree of difference between the two texts.

[0036] In another embodiment, the degree of difference can be determined based on the change rate between the recognized text and the corresponding proofread text. The change rate is determined based on the amount of change of the proofread text relative to the recognized text. The amount of change can be the number of different characters or words between the recognized text and the corresponding proofread text. The specific method of measuring the degree of difference can be determined according to the specific scenario. For example, in the annotation scenario, the degree of difference between the two texts can be measured by the change rate.

[0037] In one embodiment, the change rate is the ratio of the number of changed characters to the total number of characters in the recognized text, as shown in Table 1 below.

[0038] Table 1

[0039]

[0040] In this embodiment, a plurality of proofreading texts may be divided into a plurality of first proofreading text groups based on the change rate intervals of the change rates of the respective proofreading texts; different first proofreading text groups correspond to different change rate intervals.

[0041] The change rate interval is, for example, [0.0, 0.1], (0.1, 0.2], (0.2, 0.3], (0.3, 0.4], (0.4, 0.5], (0.5, 0.6], (0.6, 0.7], (0.7, 0.8], (0.8, 0.9], (0.9, 1.0], or [0.0, 0.2], (0.2, 0.4], (0.4, 0.6], (0.6, 0.8], (0.8, 1.0]. Specifically, multiple change rate intervals can be divided according to actual needs or actual sampling effects.

[0042] For example, among several proofreading texts, the proofreading texts with a change rate of (0.3, 0.4] include proofreading text A and proofreading text B. Then, proofreading text A and proofreading text B can be classified into the first proofreading text group with a change rate of (0.3, 0.4].

[0043] S14: Extract proofreading texts from each first proofreading text group as sample texts.

[0044] In one embodiment, see Figure 2 , Figure 2 yes Figure 1 The flowchart of an embodiment of step S14 is shown. In this embodiment, step S14 further includes:

[0045] S21: Obtain a target sampling quantity of a number of proofreading texts.

[0046] S22: Sampling proofreading texts from each first proofreading text group according to the target sampling quantity to obtain sample texts.

[0047] The target sampling quantity can be preset according to the actual situation (such as manpower and physical consumption), and is not specifically limited here. For example, if the total amount of several proofreading texts is 1000, and 30% of the proofreading texts need to be extracted as sample texts, the target sampling quantity is 300.

[0048] In one embodiment, after determining the target sampling quantity, a text group sequence of each first proofreading text group can be obtained first; then, according to the arrangement order of each first proofreading text group in the text group sequence, at least one round of sampling of proofreading texts is performed on each first proofreading text group until the number of proofreading texts extracted reaches the target sampling quantity, and each proofreading text extracted is used as a sample text.

[0049] In a specific embodiment, during each round of sampling, the expected sampling numbers of different first proofreading text groups are the same. There may be first proofreading text groups whose proofreading text number is less than the expected sampling number, for example, the expected sampling number of each group in each round is 2, but there may be first proofreading text groups whose sampling number is 1 or 0.

[0050] When there is a first proofreading text group in each first proofreading text group whose number of proofreading texts is less than the expected sampling number (for distinction, the first proofreading text group that meets this condition is called the second proofreading text group), all the proofreading texts in the second proofreading text group are used as sample texts for this round of sampling. That is, if there are proofreading texts in the second proofreading text group, all the proofreading texts in the second proofreading text group are used as sample texts for this round of sampling. Of course, it can be understood that if there are no proofreading texts in the second proofreading text group, that is, the number of proofreading texts is 0, no proofreading texts will be sampled.

[0051] In one implementation scenario, for sampling uniformity, in each round of sampling, a proofreading text may be extracted from each first proofreading text group as a sample text. If there is no first proofreading text group in the sampling process, it will be skipped. If at the end of one round, the number of sample texts extracted does not reach the target sampling quantity, the next round of extraction will be carried out until the number of sample texts extracted reaches the target sampling quantity.

[0052] Of course, at least two proofreading texts can be extracted from each text group as sample texts each time, and the specific number of sample texts extracted from each group each time can be adjusted according to actual conditions. For example, when the target sampling amount is multiple times the number of the first proofreading text groups, multiple proofreading texts can be extracted from each first proofreading text group each time as sample texts in the first round. If the absolute difference between the number of sample texts extracted in the first round and the target sampling amount is less than the number of the first proofreading text groups that still have proofreading texts, it can be adjusted to extract one annotated text from each first proofreading text group each time as the sample text.

[0053] It should be noted that the text group sequence of the first proofreading text groups can be formed by randomly sorting the first proofreading text groups, or can be arranged in order of difference from large to small, or from small to large.

[0054] In another embodiment, see Figure 3 , Figure 3 yes Figure 1 The flowchart of another embodiment of step S14 is shown. In this embodiment, step S14 further includes:

[0055] S31: Obtaining the sampling ratio corresponding to each first proofreading text group; wherein the sampling ratio is determined based on the difference degree and the error probability of the corresponding proofreading text, and the error probability is positively correlated with the sampling ratio.

[0056] S32: Extracting proofreading texts from each first proofreading text group as sample texts according to the sampling ratio corresponding to each first proofreading text group.

[0057] In this embodiment, sample texts are extracted from each first proofreading text group according to the sampling ratio corresponding to each first proofreading text group.

[0058] In this embodiment, the sampling ratio is determined based on the difference degree and the error probability of the corresponding proofreading text, and the error probability is positively correlated with the sampling ratio. That is, for the first proofreading text group corresponding to the difference degree (or change rate interval) with a larger error probability, a higher sampling ratio can be set, and conversely, for the first proofreading text group corresponding to the difference degree (or change rate interval) with a smaller error probability, a lower sampling ratio can be set.

[0059] The difference degree and the error probability of the corresponding proofreading text are obtained based on the statistics of the difference degree and the corresponding error amount of a large number of proofreading texts.

[0060] In one implementation scenario, if statistics show that the proportion of proofread texts that are consistent with the recognized text is high, and the probability of errors in a large number of proofread texts that are consistent with the recognized text is very low, in order to increase the amount of erroneous texts extracted, and to avoid the situation where the accuracy of the sampled texts meets the requirements, but the accuracy of the actual proofread texts does not meet the requirements, a higher sampling ratio can be set for the first text group with a higher probability of errors.

[0061] The above scheme first divides a number of proofreading texts into a number of first proofreading text groups based on the degree of difference between each recognized text and the corresponding proofreading text, and then extracts sample texts from each of the divided first proofreading text groups. Compared with the method of directly sampling from a number of proofreading texts without dividing the text groups based on the degree of difference, the present application extracts sample texts from each of the first proofreading text groups obtained by the division, which can effectively reduce the situation where the sample texts are concentrated in a certain first proofreading text group. Therefore, the quality of the sample texts extracted by the present application can accurately represent the quality of a number of proofreading texts.

[0062] Furthermore, after dividing the text groups based on the degree of difference, proofreading texts with similar degree of difference are divided into the same group, so the sample text extracted from each group can better represent the overall text of the group. Compared with the method of directly sampling from a number of proofreading texts without dividing the text groups based on the degree of difference, the above method of the present application can reduce the situation where the sample text extracted cannot accurately reflect a number of proofreading texts due to the large degree of difference.

[0063] In some embodiments, in order to facilitate the search of sample texts randomly selected, after the proofreading texts are extracted from each first proofreading text group as sample texts, the extracted sample texts can also be highlighted, and / or the remaining texts other than the sample texts in several proofreading texts can be hidden, and / or a function of selectively displaying any one or several first proofreading text groups can be provided, for example, relevant buttons or options are provided to display the selected first proofreading text group and hide the unselected first proofreading text groups.

[0064] To facilitate understanding of the data sampling method provided by the present application, the following briefly summarizes the above-mentioned implementation methods of the present application by taking the annotation scenario as an example:

[0065] First, with the development of artificial intelligence technology, supervised learning technology based on machine learning has been applied to more and more fields. In supervised learning, machine learning models are trained using training data, so the data annotation quality of training data is crucial to the training of the model. At present, manual annotation is usually used to generate training data. In order to make the annotation results of training data more accurate, multiple rounds of annotation are often used, such as annotation + inspection. However, multiple rounds of annotation lead to a significant increase in costs, so the number of rounds after the first round can be sampled. Random sampling directly from a large number of annotated texts may cause the sampled samples to not accurately represent the overall quality of the annotated text, or even a large number of erroneous texts to be missed, so improving the accuracy of the sampled samples is conducive to improving the quality of the annotated text.

[0066] Among them, the proofread text mentioned in the present application can be an annotated text obtained by proofreading the recognized text. Proofreading the recognized text can be obtained by proofreading (modifying) the erroneous characters in the recognized text using relevant models (such as language / semantic models), or by proofreading (modifying) the erroneous characters in the recognized text using manual annotation. The modification can include but is not limited to at least one of text replacement, text deletion, and insertion of new text; the recognized text is obtained by recognizing the target speech using a relevant speech recognition model or transcription engine.

[0067] Among them, based on the change rate between the recognized text and the annotated text, the change rate interval corresponding to the annotated text can be determined, and then several annotated texts can be divided into multiple text groups according to the change rate interval in which the change rate of each annotated text lies, wherein different text groups correspond to different change rate intervals.

[0068] Furthermore, when extracting sample texts from each text group, one of the following methods may be used.

[0069] Method 1: Obtain a target sampling quantity of the annotated text, and obtain sample texts by sampling from each text group according to the target sampling quantity.

[0070] The target sampling quantity can be determined according to the total amount of annotated texts and the set sampling ratio, or it can be determined directly. Before sampling, the sampling order (sampling sequence) of each text group can be determined first, for example, by the size order of each change rate interval, such as from large to small or from small to large, or by random sorting, to obtain the sampling sequence of each text group. Then, according to the sampling sequence of each text group, at least one round of sampling is performed on each text group until the number of sample texts extracted reaches the target sampling quantity.

[0071] In one implementation scenario, in order to ensure sampling uniformity, in each round of sampling, an annotated text can be extracted from each text group as a sample text. If there is no annotated text in the text group during the sampling process, it will be skipped. If at the end of one round, the number of sample texts extracted still does not reach the target sampling quantity, the next round of extraction will be carried out until the number of sample texts extracted reaches the target sampling quantity.

[0072] Of course, at least two annotated texts may be extracted from each text group as sample texts each time, and the specific number of sample texts extracted from each group each time may be adjusted according to actual conditions. For example, when the target sampling amount is multiple times the number of text groups, multiple annotated texts may be extracted from each text group each time as sample texts in the first round. If the absolute difference between the number of sample texts extracted in the first round and the target sampling amount is less than the number of text groups that still have annotated texts, one annotated text may be extracted from each text group each time as sample text.

[0073] Method 2: First, obtain the sampling ratio corresponding to each text group; wherein the sampling ratio is determined based on the change rate of the annotated text and the corresponding error probability, and the error probability is positively correlated with the sampling ratio; then, according to the sampling ratio corresponding to each text group (text group with a change rate range), extract the annotated text from each text group as the sample text.

[0074] The change rate of the annotated text and the corresponding error probability are obtained by statistically analyzing the error amounts based on a large number of annotated texts with different change rates.

[0075] It is understandable that if the error probability of the annotated text in a certain change rate interval is high, the sampling ratio of the annotated text in the change rate interval can be increased, otherwise the sampling ratio of the annotated text in the change rate interval can be reduced, so as to increase the extraction amount of erroneous annotated text in this way, so as to facilitate the subsequent strict control of the quality of the annotated text. For example, if the quality of the extracted sample text still meets the requirements after increasing the extraction amount of erroneous annotated text, it means that the quality of the overall annotated text is higher.

[0076] In summary, the above method of dividing text groups based on the change rate and sampling from each text group can effectively improve the accuracy of the sampled samples, that is, the quality of the extracted sample text can accurately reflect the overall quality of all annotated texts.

[0077] To elaborate, in general, the amount of annotated text contained in each change rate interval is different. Some intervals may contain a large amount of annotated text, while some intervals may contain only a small amount of annotated text. If a random sampling method is adopted from a large amount of annotated text, the sampled text may mostly come from the same change rate interval, while the annotated text in some change rate intervals is not sampled. Correspondingly, the accuracy of the sampled text is low and the sampled text cannot fully represent the annotated text in all change rate intervals.

[0078] However, the present application is different. In the sampling method of the sampling division interval of the present application, the corresponding sample text will be drawn in each text group, so the sample text finally drawn can fully represent the annotated text of all change rate intervals, and the accuracy of the sample text is high.

[0079] See also Figure 4 , Figure 4 It is a schematic diagram of the framework of an embodiment of a data sampling device provided by the present application. In this embodiment, the data sampling device 40 includes an acquisition module 41, a proofreading module 42, a division module 43 and an extraction module 44. The acquisition module 41 is used to acquire recognition texts corresponding to a plurality of target voices; the proofreading module 42 is used to perform text proofreading processing on each recognition text to obtain a corresponding proofreading text; the division module 43 is used to divide a plurality of proofreading texts into a plurality of first proofreading text groups based on the degree of difference between each recognition text and the corresponding proofreading text; the extraction module 44 is used to extract proofreading texts from each first proofreading text group as sample texts.

[0080] In some embodiments, the degree of difference is determined based on the change rate between the recognized text and the corresponding proofreading text; the division module 43 divides the plurality of proofreading texts into a plurality of first proofreading text groups based on the degree of difference between each recognized text and the corresponding proofreading text, including: dividing the plurality of proofreading texts into a plurality of first proofreading text groups based on the change rate interval in which the change rate of each proofreading text lies; different first proofreading text groups correspond to different change rate intervals.

[0081] In some embodiments, the extraction module 44 extracts proofreading texts from each first proofreading text group as sample texts, including: obtaining a target sampling amount of a number of proofreading texts; sampling proofreading texts from each first proofreading text group according to the target sampling amount to obtain sample texts.

[0082] In some embodiments, sampling of proofreading texts is performed from each first proofreading text group according to a target sampling quantity to obtain sample texts, including: obtaining a text group sequence of each first proofreading text group; performing at least one round of sampling of proofreading texts on each first proofreading text group according to an arrangement order of each first proofreading text group in the text group sequence, until the number of proofreading texts sampled reaches the target sampling quantity, and using each sampled proofreading text as a sample text.

[0083] In some embodiments, during each round of sampling, the expected sampling numbers of different first proofreading text groups are the same; wherein the first proofreading text group whose proofreading text number is less than the expected sampling number is the second proofreading text group, and all the proofreading texts in the second proofreading text group are used as sample texts for this round of sampling.

[0084] In some embodiments, the extraction module 44 extracts proofreading texts from each first proofreading text group as sample text, and also includes: obtaining a sampling ratio corresponding to each first proofreading text group; wherein the sampling ratio is determined based on the degree of difference and the error probability of the corresponding proofreading text, and the error probability is positively correlated with the sampling ratio; according to the sampling ratio corresponding to each first proofreading text group, extracting proofreading texts from each first proofreading text group as sample text.

[0085] In some embodiments, the text proofreading processing performed by the proofreading module 42 includes at least one of text replacement, text deletion and insertion of new text; and / or, the proofreading text is annotated text about the target voice; and / or, after the extraction module 44 extracts the proofreading text from each first proofreading text group as a sample text, the method further includes at least one of the following steps: highlighting the sample text; hiding the remaining text in a number of proofreading texts except the sample text.

[0086] In some embodiments, the speech duration of each target speech acquired by the acquisition module 41 is less than or equal to a preset duration; and / or, each target speech is obtained by segmenting at least one original speech; and / or, the recognized text is obtained by performing speech recognition on the target speech using a speech recognition model.

[0087] See also Figure 5 , Figure 5 1 is a schematic diagram of a framework of an electronic device according to an embodiment of the present application. In this embodiment, the electronic device 50 includes a memory 51 and a processor 52 coupled to each other.

[0088] The memory 51 stores program instructions, and the processor 52 is used to execute the program instructions stored in the memory 51 to implement the steps of any of the above method implementations. In a specific implementation scenario, the electronic device 50 may include, but is not limited to: a microcomputer, a server, and in addition, the electronic device 50 may also include a mobile device such as a laptop computer and a tablet computer, which is not limited here.

[0089] Specifically, the processor 52 is used to control itself and the memory 51 to implement the steps of any of the above-mentioned embodiments. The processor 52 can also be called a CPU (Central Processing Unit). The processor 52 may be an integrated circuit chip with signal processing capabilities. The processor 52 can also be a general-purpose processor, a digital signal processor (Digital Signal Processor, DSP), an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), a field-programmable gate array (Field-Programmable Gate Array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. In addition, the processor 52 can be implemented by an integrated circuit chip.

[0090] See also Figure 6 , Figure 6 It is a schematic diagram of the framework of the computer-readable storage medium provided by the present application. The computer-readable storage medium 60 of the embodiment of the present application stores a program instruction 61, and when the program instruction 61 is executed, the method provided by any embodiment of the above method and any non-conflicting combination is implemented. Among them, the program instruction 61 can form a program file and be stored in the above-mentioned computer-readable storage medium 60 in the form of a software product, so that a computer device (which can be a personal computer, a server, or a network device, etc.) executes all or part of the steps of each implementation method of the present application. The aforementioned computer-readable storage medium 60 includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a disk or an optical disk, or a terminal device such as a computer, a server, a mobile phone, and a tablet.

[0091] The above scheme first divides a number of proofreading texts into a number of first proofreading text groups based on the degree of difference between each recognized text and the corresponding proofreading text, and then extracts sample texts from each of the divided first proofreading text groups. Compared with the method of directly sampling from a number of proofreading texts without dividing the text groups based on the degree of difference, the present application extracts sample texts from each of the first proofreading text groups obtained by the division, which can effectively reduce the situation where the sample texts are concentrated in a certain first proofreading text group. Therefore, the quality of the sample texts extracted by the present application can accurately represent the quality of a number of proofreading texts.

[0092] Furthermore, after dividing the text groups based on the degree of difference, proofreading texts with similar degree of difference are divided into the same group, so the sample text extracted from each group can better represent the overall text of the group. Compared with the method of directly sampling from a number of proofreading texts without dividing the text groups based on the degree of difference, the above method of the present application can reduce the situation where the sample text extracted cannot accurately reflect a number of proofreading texts due to the large degree of difference.

[0093] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.

[0094] The above description of various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced to each other, and for the sake of brevity, they will not be repeated herein.

[0095] In the several embodiments provided in the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation described above is only schematic. For example, the division of modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical or other forms.

[0096] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0097] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0098] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) or a processor (processor) to perform all or part of the steps of each implementation method of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program code.

[0099] The above description is only an implementation method of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly used in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A data sampling method, characterized in that: The method comprises: Obtain recognition texts corresponding to several target voices respectively; Performing text proofreading on each recognized text to obtain a corresponding proofread text; Based on the degree of difference between each of the recognized texts and the corresponding proofreading texts, the plurality of proofreading texts are divided into a plurality of first proofreading text groups; The proofreading texts are extracted from each of the first proofreading text groups as sample texts.

2. The method according to claim 1, characterized in that The degree of difference is determined based on the change rate between the recognized text and the corresponding proofread text; The step of dividing the plurality of proofreading texts into a plurality of first proofreading text groups based on the degree of difference between each of the recognized texts and the corresponding proofreading texts comprises: Based on the change rate interval of each of the proofreading texts, the proofreading texts are divided into a plurality of first proofreading text groups; different first proofreading text groups correspond to different change rate intervals.

3. The method according to claim 1 or 2, characterized in that: The step of extracting the proofreading texts from each of the first proofreading text groups as sample texts includes: Obtaining a target sampling quantity of a number of the proofreading texts; The proofreading texts are sampled from each of the first proofreading text groups according to the target sampling quantity to obtain the sample texts.

4. The method according to claim 3, characterized in that The sampling of the proofreading texts from each of the first proofreading text groups according to the target sampling amount to obtain the sample texts includes: Obtaining a text group sequence of each of the first proofreading text groups; According to the arrangement order of each of the first proofreading text groups in the text group sequence, at least one round of sampling of the proofreading texts is performed on each of the first proofreading text groups until the number of the proofreading texts sampled reaches the target sampling amount, and each of the proofreading texts sampled is used as the sample text.

5. The method according to claim 4, characterized in that In each round of sampling, the expected number of samples of different first proofreading text groups is the same; The first proofreading text group in which the number of proofreading texts is less than the expected sampling number is the second proofreading text group, and all the proofreading texts in the second proofreading text group are used as the sample texts for this round of sampling.

6. The method according to claim 1 or 2, characterized in that: The step of extracting the proofreading texts from each of the first proofreading text groups as sample texts further includes: Obtaining a sampling ratio corresponding to each of the first proofreading text groups; wherein the sampling ratio is determined based on the difference degree and the error probability of the corresponding proofreading text, and the error probability is positively correlated with the sampling ratio; According to the sampling ratio corresponding to each of the first proofreading text groups, proofreading texts are extracted from each of the first proofreading text groups as sample texts.

7. The method according to claim 1, characterized in that The text proofreading process includes at least one of replacing text, deleting text, and inserting new text; And / or, the proofreading text is an annotated text about the target speech; And / or, after extracting the proofreading texts from each of the first proofreading text groups as sample texts, the method further comprises at least one of the following steps: highlighting the sample text; Hide the remaining texts in the proofreading texts except the sample texts.

8. The method according to claim 1, characterized in that The speech duration of each of the target speech is less than or equal to a preset duration; And / or, each of the target speech is obtained by segmenting at least one original speech; And / or, the recognized text is obtained by performing speech recognition on the target speech using a speech recognition model.

9. A data sampling device, characterized in that: The device comprises: An acquisition module is used to acquire recognition texts corresponding to a plurality of target voices; A proofreading module is used to perform text proofreading on each recognized text to obtain a corresponding proofread text; A division module, used for dividing the plurality of proofreading texts into a plurality of first proofreading text groups based on the degree of difference between each of the recognized texts and the corresponding proofreading texts; The extraction module is used to extract the proofreading texts from each of the first proofreading text groups as sample texts.

10. An electronic device, characterized in that: comprising a memory and a processor coupled to each other, The memory stores program instructions; The processor is used to execute the program instructions stored in the memory to implement the method according to any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program instructions that can be run by a processor, and the program instructions can be executed by the processor to implement the method according to any one of claims 1 to 8.