Data Screening Method, Device, Storage Medium and Electronic Device

Through the recognition model, the speech data is initially annotated and trained, and the speech data is generated with confusion filtered, which solves the problem of data randomness and uneven distribution in the speech recognition model training, and improves the generalization performance of the model.

CN114970880BActive Publication Date: 2025-06-20BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210350894.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-02
Publication Date
2025-06-20
Estimated Expiration
2042-04-02

AI Technical Summary

Technical Problem

In the prior art, due to the randomness of speech data and uneven domain distribution during training, the model after training is low reliability and poor generalization performance.

Method used

By using the recognition model to identify the data to be recognized, the initial annotation is generated, the data is added to the training set for model training, the confusion of the data is generated, and the data is determined as the target data when the confusion is greater than the preset threshold.

Benefits of technology

Efficient screening of candidate samples is achieved, the repetition rate of the data training set is reduced, the uniformity of the data distribution is improved, and the generalization performance of the identification model is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114970880B_ABST
    Figure CN114970880B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a data screening method, apparatus, storage medium and electronic device. The method includes: using an identification model to identify data to be identified, generating a first initial annotation, adding the data to be identified and the corresponding first initial annotation to the data training set corresponding to the identification model, training the identification model based on the data training set, generating a perplexity of the data to be screened through the trained identification model, and determining the data to be screened as target data when the perplexity is greater than a preset perplexity threshold. Thereby, it is ensured that the repetition rate of the data screened in each iteration with the existing data training set is low, and the distribution of the data training set is more uniform, and finally the trained identification model has better generalization performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computers, and in particular, to a data screening method, apparatus, storage medium, and electronic device. Background Art

[0002] As a machine learning task, speech recognition requires a large amount of speech data for model training. In the prior art, the sources of speech data are extensive and the collection difficulty is low. When performing model training, it is necessary to randomly sample the collected speech data, and then put the sampled speech data into the model for recognition training. However, due to problems such as strong randomness and uneven domain distribution of the obtained speech data, the repetition rate of the speech data used for training is relatively high, resulting in problems of low reliability and poor generalization performance of the obtained recognition model after training. Summary of the Invention

[0003] This part of the content is provided to introduce the concepts in a brief form, and these concepts will be described in detail in the following specific implementation part. This part of the content is not intended to identify the key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0004] In a first aspect, the present disclosure provides a data screening method, including:

[0005] Using a recognition model to recognize data to be recognized, and generating a first initial annotation;

[0006] Adding the data to be recognized and the corresponding first initial annotation to the data training set corresponding to the recognition model, and training the recognition model based on the data training set;

[0007] Generating the perplexity of the data to be screened through the trained recognition model;

[0008] When the perplexity is greater than a preset perplexity threshold, determining the data to be screened as target data.

[0009] In a second aspect, the present embodiment provides a data screening apparatus, the apparatus includes:

[0010] A first generation module, configured to use a recognition model to recognize data to be recognized, and generate a first initial annotation;

[0011] A training module, configured to add the data to be recognized and the corresponding first initial annotation to the data training set corresponding to the recognition model, and train the recognition model based on the data training set;

[0012] A second generation module, configured to generate the perplexity of the data to be screened through the trained recognition model;

[0013] A determination module, configured to determine the data to be screened as target data when the perplexity is greater than a preset perplexity threshold.

[0014] In a third aspect, the present disclosure provides a computer-readable medium, on which a computer program is stored, and when the program is executed by a processing device, the steps of the method described in the first aspect are implemented.

[0015] In a fourth aspect, the present disclosure provides an electronic device, including:

[0016] A storage device, on which a computer program is stored;

[0017] A processing device, configured to execute the computer program in the storage device to implement the steps of the method described in the first aspect.

[0018] Through the above technical solutions, the data to be recognized is recognized by using a recognition model to generate a first initial annotation. The data to be recognized and the corresponding first initial annotation are added to the data training set corresponding to the recognition model. The recognition model is trained based on the data training set. The perplexity of the data to be screened is generated by the trained recognition model. When the perplexity is greater than the preset perplexity threshold, the data to be screened is determined as target data. Therefore, the data to be recognized is first used as training data, and the recognition model is recognized and trained based on this training data. Then, the perplexity of the data to be screened is determined based on the trained recognition model, and it is determined whether the data to be screened is target data according to this perplexity, so as to realize the screening of candidate samples. The above solution can quantify the fitting degree of the sample with the training set by using the perplexity of the candidate sample by the recognition model, can select sample data with low perplexity from the candidate samples, ensure that the repetition rate of the data screened each time with the existing data training set is low, and make the distribution of the data training set more uniform, and finally the trained recognition model has better generalization performance.

[0019] Other features and advantages of the present disclosure will be described in detail in the subsequent specific implementation part. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Combined with the drawings and referring to the following specific implementation manners, the above and other features, advantages and aspects of the embodiments of the present disclosure will become more obvious. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic, and the original components and elements are not necessarily drawn to scale. In the drawings:

[0021] Figure 1 is a flowchart of a data screening method shown according to an exemplary embodiment.

[0022] Figure 2 is a flowchart of a method for generating perplexity shown according to an exemplary embodiment.

[0023] Figure 3 is a flowchart of another data screening method shown according to an exemplary embodiment.

[0024] Figure 4 is a flowchart of yet another data screening method shown according to an exemplary embodiment.

[0025] Figure 5 is a structural block diagram of a data screening device shown according to an exemplary embodiment.

[0026] Figure 6 is a schematic structural diagram of an electronic device shown according to an exemplary embodiment. Detailed implementation manners

[0027] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0028] It should be understood that the various steps recorded in the method embodiments of the present disclosure can be executed in a different order and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.

[0029] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependent relationships.

[0030] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".

[0031] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0032] All actions of obtaining signals, information or data in the present disclosure are carried out on the premise of complying with the corresponding data protection regulations and policies of the country where it is located and obtaining the authorization given by the owner of the corresponding device.

[0033] It should be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to users and the authorization of users should be obtained through appropriate means in accordance with relevant laws and regulations.

[0034] For example, when responding to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, application program, server, or storage medium that executes the operation of the technical solution of the present disclosure according to the prompt message.

[0035] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving an active request from the user may be, for example, in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0036] It should be understood that the above process of notifying and obtaining user authorization is only illustrative and does not constitute a limitation on the implementation manner of the present disclosure. Other manners that meet relevant laws and regulations can also be applied to the implementation manner of the present disclosure.

[0037] At the same time, it should be understood that the data involved in the present technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of corresponding laws, regulations and related regulations.

[0038] For the training of existing recognition models, a large amount of sample data is required. The recognition model is used to label the sample data, and then the corresponding labels are calibrated to make the recognition model more accurate. However, when collecting sample data, due to the wide range of data sources and the low collection difficulty, the distribution of the sample data in the corresponding fields is uneven, and there is a problem of high repetition rate, resulting in weak recognition ability and poor generalization performance of the finally obtained recognition model through training. Therefore, before training the recognition model, it is necessary to screen the sample data to obtain sample data with uniform distribution fields and low repetition rate. Existing screening methods are all to classify and screen the fields to which the sample data belongs after manually reading the sample data. When training the recognition model, the required amount of sample data is large. The manual screening method is too inefficient, has a high error rate, consumes a large amount of human costs, and obtains fewer target samples.

[0039] In view of this, in order to improve the generalization performance of the recognition model and enhance the recognition ability of the recognition model, the embodiments of the present disclosure provide a data screening method. It can be understood that the data involved in the technical solution (including but not limited to the data itself, data acquisition or use) should comply with the requirements of relevant laws, regulations and related provisions.

[0040] Figure 1 is a flowchart of a data screening method shown according to an exemplary embodiment. Refer to Figure 1 This data screening method includes:

[0041] Step S11: Use the recognition model to recognize the data to be recognized and generate a first initial annotation.

[0042] It can be understood that in this embodiment, the recognition model is used to recognize the sample data, and according to the recognition result, the sample data can be annotated to determine relevant information such as the attributes, content, and data volume of the sample data. The recognition model can be any data model for machine recognition, which converts the sample data into recognition data in a target form through a built-in model conversion algorithm. The target form can be a data form such as text form, audio form, video form, etc. that can be used for intuitive query. For example, the recognition model can be a speech recognition model, and the corresponding annotation information can be the semantic information of the sample data, etc.; the recognition model can also be a barcode recognition model. By scanning the corresponding barcode, the obtained annotation information can be the item information recorded in the barcode (including: quantity, origin, production date, etc.).

[0043] The data to be recognized in this embodiment is a part of the sample database. The preset number of sample data can be obtained as the data to be recognized by random sampling in the sample database. Using the data to be recognized as the basic data, the recognition model recognizes the recognition data and generates a corresponding first initial annotation.

[0044] It should be noted that this is only a screening process for the sample data in this recognition stage. Therefore, in order to improve the screening efficiency, when the recognition model recognizes the sample data in this stage, only part of the recognition algorithms can be called for preliminary recognition operations to improve the model recognition efficiency. The corresponding first initial annotation generated can be a pseudo-label, and this pseudo-label can be used to distinguish the types and fields of each sample data, etc. For example, the sample data is language data, and the recognition model is a language recognition model. After obtaining the preset number of data to be recognized from the sample data by random sampling, the first initial annotation obtained through the language recognition model can be the first initial annotation obtained by the recognition model through preliminary recognition without considering relevant language habits such as dialects, modal particles, and rhotacized syllables. This first initial annotation is a pseudo-label and is not the final recognition result of the recognition model.

[0045] Optionally, before the above step S11, the screening method may further include:

[0046] Obtain initial data from the initial database.

[0047] In the case where there is no corresponding annotation information for the initial data, the initial data is used as the data to be recognized.

[0048] It should be noted that the data sources in the initial database are relatively extensive, including both labeled sample data and unlabeled sample data. To avoid the influence of annotation information on the recognition training of the recognition model, in this embodiment, it is necessary to preliminarily screen the sample data in the initial database to determine the unlabeled data in the initial database as the data to be recognized. For example, in the present disclosure, the data in the initial database is preliminarily screened, and a corresponding sample database is generated based on the unlabeled data.

[0049] Step S12, add the data to be recognized and the corresponding first initial annotation to the data training set corresponding to the recognition model, and train the recognition model based on the data training set.

[0050] It can be understood that in this embodiment, the recognition model needs to have a certain data recognition ability to perform primary recognition on the data to be recognized. Therefore, when training the recognition model with data, it is necessary to establish a corresponding data training set. This data training set includes a certain number of preset training data and corresponding annotation information. Training the recognition model through this data training set can enable the recognition model to learn how to obtain accurate annotation information from the training data, thereby improving the recognition ability. Adding the labeled sample data to this data training set and training the recognition model based on the updated data training set can thus improve the recognition ability of the recognition model. In this embodiment, the data to be recognized and the corresponding first initial annotation are added to the data training set corresponding to the recognition model, and the recognition model is trained based on the updated data training set, so that the recognition model updates the corresponding recognition algorithm based on the newly added data to be recognized and the corresponding first initial standard.

[0051] Step S13, generate the perplexity of the data to be screened through the trained recognition model.

[0052] It should be noted that the data to be screened in this embodiment is sample data randomly sampled from the sample database. By identifying the data to be screened, the perplexity of the identification model for the data to be screened is determined. Among them, the perplexity represents the identification ability of the identification model for the data to be screened based on the data in the data training set with the corresponding identification algorithm. For example, the higher the perplexity, the weaker the identification ability of the identification model for the data to be screened through the data training set, and at the same time, the lower the similarity between the data to be screened and other data in the data training set; the lower the perplexity, the more accurately the identification model can identify the corresponding annotation information of the data to be screened based on the data training set, and the higher the similarity between the data to be screened and the corresponding data in the data training set. It can be understood that when determining the perplexity of the data to be screened by the identification model, it is necessary to compare the data to be screened with the data in the data training set. Among them, the comparison method can be confirmed according to the type of the identification model. For example, when the identification model is a language identification model, it is necessary to use the identification algorithm of the identification model to determine the audio curve of the data to be screened and compare it with the audio curve of the preset language data in the data training set to judge whether there is data information similar to the data to be screened in the preset language data, as well as the similarity degree of the data information, so as to determine the perplexity of the data to be screened.

[0053] Figure 2 is a flowchart of a method for generating perplexity shown according to an exemplary embodiment. Refer to Figure 2 The above step S13 includes:

[0054] Step S131, identify the data to be screened through the trained identification model to generate a second initial annotation.

[0055] Step S132, compare the second initial annotation with the data training set to confirm the fitting degree of the second initial annotation.

[0056] Step S133, determine the perplexity corresponding to the data to be screened according to the fitting degree.

[0057] It can be understood that before screening the data to be screened, it is necessary to preliminarily identify the data to be screened through the trained recognition model to determine the second initial annotation corresponding to the data to be screened. Compare the second initial annotation with the relevant annotation data in the data training set to determine the fitting degree of the second initial annotation with the data training set. Among them, the fitting degree can be the similarity degree between the second initial annotation and the relevant annotation data in the data training set. For example, the similarity degree between each byte can be obtained by comparing the bytes of the second initial annotation with the relevant annotation data, so as to determine the fitting degree of the second initial annotation. According to this fitting degree, determine the perplexity corresponding to the data to be screened. For example, when comparing the second initial annotation with the data training set, the relevant information such as the field, byte capacity, and meaning representation to which the second initial annotation belongs can also be determined by analyzing the second initial annotation, and based on this relevant information, it can be determined whether there is other annotation data in the data training set that is the same as this relevant information, so as to determine the fitting degree of the second initial annotation. It should be noted that in this embodiment, the fitting degree is inversely proportional to the perplexity. The higher the fitting degree of the second initial annotation, the lower the corresponding perplexity.

[0058] Optionally, the above step S132 includes:

[0059] Parse the second initial annotation to determine the corresponding entry data.

[0060] Compare the entry data with the annotation data corresponding to the data training set to generate the word frequency corresponding to the entry data.

[0061] Determine the fitting degree corresponding to the second initial annotation according to the word frequency.

[0062] It can be understood that in this embodiment, by parsing the second initial annotation, the entry data included in the second initial annotation is determined. For example, when the recognition model is a language recognition model, the data to be screened is language data. By using the recognition model to recognize the language data, the semantic information of the language data is determined as the second initial annotation. Generally, language data consists of words, prepositions, and conjunctions. By parsing the second initial annotation, the entry data that constitutes the second initial annotation is obtained. The entry data is compared with the corresponding annotation data in the data training set to determine the word frequency of each entry data in the data training set. According to the word frequency of each entry data, the fitting degree corresponding to the second initial annotation is determined. For example, when determining the word frequency corresponding to the second initial annotation, it is necessary to generate the corresponding word frequency according to the frequency of each entry data in the second initial flag in the data training set. At the same time, it is also necessary to perform a full permutation and combination of the entry data according to each entry data corresponding to the second initial annotation, and compare the newly generated entry data after the permutation and combination with the relevant data in the data training set to generate the corresponding word frequency; the fitting degree corresponding to the second initial annotation is determined according to the generated word frequency.

[0063] Step S14, when the perplexity is greater than the preset perplexity threshold, determine the data to be screened as the target data.

[0064] It can be understood that in this embodiment, it is necessary to screen out the relevant data with a low repetition rate from the sample data as the target data for model training of the recognition model, so that the recognition model can adapt to more various types of sample data and improve the generalization ability of the recognition model. After determining the perplexity of the data to be screened through the above steps, the perplexity is compared with the preset perplexity threshold. When the perplexity is greater than the preset perplexity threshold, it indicates that the recognition ability of the recognition model for the data to be screened is low, and the annotation information corresponding to the data to be screened cannot be accurately recognized through the existing data training set and the corresponding recognition algorithm. Then, it is determined that the data to be screened is the target data, and this data to be screened can be added to the data training set to continue training the recognition model, thereby improving the generalization ability of the recognition model. For example, the perplexity can be in numerical form, and the preset perplexity threshold can be set to 100. If the perplexity of the data to be screened determined by the recognition model exceeds 100, then it is determined that the data to be screened is the target data; when the perplexity corresponding to the data to be screened is less than 100, the data to be screened is discarded and removed from the sample database.

[0065] Figure 3 is a flowchart of another data screening method shown according to an exemplary embodiment. Refer to Figure 3 After the above step S14, the screening method further includes:

[0066] Step S141: Use the target data as the data to be recognized, and repeatedly perform the step of using the recognition model to recognize the data to be recognized until the recognition model is trained based on the data training set, so as to generate a trained recognition model.

[0067] Step S142: Screen other data to be screened based on the trained recognition model.

[0068] It can be understood that after determining that the data to be screened is the target data through the above steps, use the data to be screened as the data to be recognized, re-recognize based on the recognition model to determine the corresponding annotation information, add the data to be recognized and the corresponding annotation information to the data training set, and re-train the recognition model based on the updated data training set, so as to generate a trained recognition model, and screen other data to be screened in the sample database based on the trained recognition model. It should be noted that during the process of continuing to screen other data to be screened, after determining the target data through the recognition model, the target data needs to be re-added to the data training set, and the recognition model is trained by the updated data training set, and then other data to be screened are screened based on the trained recognition model. The screening process is a cyclic process. After determining the target data, first use the target data as the data to be recognized to train the recognition model, and then perform perplexity scoring on other data to be screened based on the trained recognition model.

[0069] Optionally, step S142 includes:

[0070] Divide other data to be screened into multiple data subsets with a preset number. Optionally, evenly divide other data to be screened into multiple data subsets with a preset number.

[0071] Use each data subset as the data to be recognized in turn, and repeatedly perform the above step of using the recognition model to recognize the data to be recognized until a trained recognition model is generated.

[0072] Obtain the data training set corresponding to the trained recognition model.

[0073] Generate a target data set based on the data in the data training set.

[0074] It can be understood that during the training process of the recognition model, there are a large number of sample databases in the sample database that need to be screened. By taking out the data to be screened from the sample database one by one through the above-mentioned implementation method and screening and recognizing them through the recognition model, it takes a long time and the efficiency is low. Therefore, in this embodiment, other data to be screened in the sample database can be divided into a preset number of data subsets. Starting from the first data subset, the relevant data in the first data subset is used as the data to be recognized, and according to the recognition method in the above-mentioned embodiment, the data to be recognized in the first data subset is initially recognized based on the recognition model to generate a corresponding first initial annotation set. The first data subset and the corresponding first initial annotation set are added to the data training set corresponding to the recognition model, and the recognition model is trained based on the data training set. The second data subset is used as the data to be screened, and the trained recognition model is used to perform perplexity recognition on the second data subset. Optionally, the perplexity of the second data subset can be the average of the sub-perplexities corresponding to each piece of data to be screened in the data subset, or it can also be the sum of the sub-perplexities corresponding to each piece of data to be screened. This embodiment does not make a limitation on this.

[0075] When it is determined that the perplexity corresponding to the second data subset is greater than the preset perplexity threshold, it is determined that the relevant data corresponding to the second data subset is the target data; and the second data subset is used as the data to be recognized, the second data subset is recognized to determine the initial annotation, and the second data subset and the corresponding initial annotation are added to the data training set corresponding to the recognition model. The recognition model is retrained based on the updated data training set, and the perplexity score of the third data subset is performed according to the trained recognition model. Traverse the sample database to screen out the target data set from the sample database, and determine that the newly added data set in the data training set is the target data set. When it is determined that the perplexity corresponding to the second data subset is less than or equal to the preset perplexity threshold, it is determined to discard the second data subset, and the perplexity score of the third data subset is performed based on the recognition model, and the scores and screens of other data subsets are performed in turn until the sample database is traversed, and the newly added data in the data training set is determined as the target data set.

[0076] Figure 4 is a flowchart of another data screening method shown according to an exemplary embodiment. See Figure 4 and this data screening method includes:

[0077] (1) Collect sample data without annotation information from the initial database through preliminary screening;

[0078] (2) Use the speech recognition model to recognize the sample data to generate corresponding pseudo-labels;

[0079] (3) Divide the sample data and the corresponding pseudo-labels into multiple subsets of data;

[0080] (4) Add the first subset of data to the data training set of the speech recognition model;

[0081] (5) Perform speech training on the speech recognition model using the updated data training set;

[0082] (6) Score the sample perplexity of the second subset of data using the trained speech recognition model;

[0083] (7) If the perplexity is greater than the preset perplexity threshold, add the second subset of data and the corresponding pseudo-labels to the data training set; if the perplexity is less than or equal to the preset perplexity threshold, discard the second subset of data;

[0084] (8) Traverse the sample database and repeat the steps in (5) - (7) for subsequent subsets;

[0085] (9) Use the data in the updated data training set as the target data.

[0086] Through the above technical solution, use the recognition model to recognize the data to be screened, generate the first initial annotation, add the data to be screened and the corresponding first initial annotation to the data training set corresponding to the recognition model, train the recognition model based on the data training set, generate the perplexity of the data to be screened using the trained recognition model, and determine the data to be screened as the target data when the perplexity is greater than the preset perplexity threshold. Thus, first use the data to be screened as the training data, recognize and train the recognition model based on this training data, then determine the perplexity of the data to be screened based on the trained recognition model, and determine whether the data to be screened is the target data according to this perplexity, realizing the screening of candidate samples. Using the perplexity of the candidate samples by the recognition model to quantify the fitting degree of the sample with the training set can select sample data with low perplexity from the candidate samples, ensure that the repetition rate of the data screened each time with the existing data training set is low, and make the distribution of the data training set more uniform, and finally the trained recognition model has better generalization performance.

[0087] Based on the same concept, the present disclosure also provides a data screening device, which can become part or all of an electronic device in the form of software, hardware, or a combination of both. Figure 5 It is a structural block diagram of a data screening device shown according to an exemplary embodiment. Refer to Figure 5 , the data screening device 100 includes:

[0088] The first generation module 110 is used to recognize the data to be recognized using the recognition model and generate the first initial annotation.

[0089] A training module 120, configured to add the data to be recognized and the corresponding first initial annotation to the data training set of the recognition model, and train the recognition model based on the data training set.

[0090] A second generation module 130, configured to generate the perplexity of the data to be screened through the trained recognition model.

[0091] A determination module 140, configured to determine the data to be screened as the target data when the perplexity is greater than a preset perplexity threshold.

[0092] Optionally, the screening device 100 further includes an execution module, and the execution module includes:

[0093] A first generation sub-module, configured to use the target data as the data to be recognized, and repeatedly execute the steps of recognizing the data to be recognized by the recognition model until training the recognition model based on the data training set, so as to generate a trained recognition model.

[0094] A screening sub-module, configured to screen other data to be screened based on the trained recognition model.

[0095] Optionally, the screening sub-module may further be configured to:

[0096] Divide other data to be screened into multiple data subsets with a preset number.

[0097] Take each data subset as the data to be recognized in turn, and repeatedly execute the steps of recognizing the data to be recognized by the recognition model until generating a trained recognition model.

[0098] Obtain the data training set corresponding to the trained recognition model.

[0099] Generate a target data set based on the data in the data training set.

[0100] Optionally, the second generation module 130 includes:

[0101] A second generation sub-module, configured to recognize the data to be screened through the trained recognition model to generate a second initial annotation.

[0102] A first determination sub-module, configured to compare the second initial annotation with the data training set to confirm the fitting degree of the second initial annotation.

[0103] A second determination sub-module, configured to determine the perplexity corresponding to the data to be screened according to the fitting degree.

[0104] Optionally, the first determination sub-module may further be configured to:

[0105] Parse the second initial annotation to determine the corresponding entry data;

[0106] Compare the entry data with the annotation data corresponding to the data training set to generate the word frequency corresponding to the entry data;

[0107] Determine the fitting degree corresponding to the second initial annotation according to the word frequency.

[0108] Optionally, the screening device 100 further includes a determination module, and the determination module is configured to:

[0109] Obtain initial data from the initial database;

[0110] In the case where there is no corresponding annotation information for the initial data, use the initial data as the data to be recognized.

[0111] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment related to the method, and will not be elaborated here.

[0112] Based on the same concept, an embodiment of the present disclosure further provides a computer-readable medium, on which a computer program is stored, and when the program is executed by a processing device, the steps of any of the above data screening methods are implemented.

[0113] Based on the same concept, an embodiment of the present disclosure further provides an electronic device, including:

[0114] A storage device on which at least one computer program is stored;

[0115] At least one processing device, configured to execute the at least one computer program in the storage device to implement the steps of any of the above data screening methods.

[0116] Refer to the following Figure 6 , which shows a schematic structural diagram of an electronic device 600 suitable for implementing an embodiment of the present disclosure. The electronic device in the embodiment of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), vehicle terminals (such as vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The electronic device shown is only an example and should not impose any limitation on the functions and usage scope of the embodiment of the present disclosure.

[0117] As shown in Figure 6As shown, the electronic device 600 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 601, which may perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 602 or the program loaded from the storage device 608 into the random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 are also stored. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. The input / output (I / O) interface 605 is also connected to the bus 604.

[0118] Generally, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 6 the electronic device 600 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.

[0119] Specifically, according to an embodiment of the present disclosure, the process described above with reference to the flowchart may be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program codes for performing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from the network through the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above functions defined in the method of the embodiment of the present disclosure are executed.

[0120] It should be noted that the computer-readable medium described above can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0121] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0122] The above computer-readable medium can be included in the above electronic device; or it can exist separately and not be assembled into the electronic device.

[0123] The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by the electronic device, the electronic device is caused to:

[0124] Use the recognition model to recognize the data to be recognized and generate a first initial annotation;

[0125] Add the data to be recognized and the corresponding first initial annotation to the data training set of the recognition model, and train the recognition model based on the data training set;

[0126] Generate the perplexity of the data to be screened through the trained recognition model;

[0127] When the perplexity is greater than the preset perplexity threshold, determine the data to be screened as the target data.

[0128] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0129] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0130] The modules involved in the embodiments of the present disclosure may be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation to the module itself in some cases.

[0131] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGAs), Application Specific Integrated Circuits (ASICs), Application Specific Standard Products (ASSPs), Systems on Chip (SOCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0132] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media would include electrical connections based on one or more wires, portable computer disks, hard disks, Random Access Memory (RAM), Read Only Memory (ROM), Erasable Programmable Read Only Memory (EPROM or Flash memory), optical fibers, portable compact disk read only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0133] According to one or more embodiments of the present disclosure, Example 1 provides a data screening method, including:

[0134] Using an identification model to identify the data to be screened, and generating a first initial annotation;

[0135] Adding the data to be screened and the corresponding first initial annotation to the data training set corresponding to the identification model, and training the identification model based on the data training set;

[0136] Generating the perplexity of the data to be screened through the trained identification model;

[0137] In the case where the perplexity is greater than a preset perplexity threshold, determining the data to be screened as target data.

[0138] According to one or more embodiments of the present disclosure, Example 2 provides the method of Example 1, and the method further includes:

[0139] Using the target data as the data to be screened, and repeatedly performing the steps of using the identification model to identify the data to be screened and training the identification model based on the data training set until the trained identification model is generated.

[0140] Screening other data to be screened based on the trained identification model.

[0141] According to one or more embodiments of the present disclosure, Example 3 provides the method of Example 2, and screening other data to be screened based on the trained recognition model includes:

[0142] Dividing the other data to be screened into a plurality of data subsets with a preset quantity;

[0143] Taking each data subset as the data to be screened in turn, and repeating the above step of using the recognition model to recognize the data to be screened until the step of generating the trained recognition model;

[0144] Obtaining the data training set corresponding to the trained recognition model;

[0145] Generating a target data set based on the data in the data training set.

[0146] According to one or more embodiments of the present disclosure, Example 4 provides the method of Example 1, and generating the perplexity of the data to be screened through the trained recognition model includes:

[0147] Recognizing the data to be screened through the trained recognition model to generate a second initial annotation;

[0148] Comparing the second initial annotation with the data training set to confirm the fitting degree of the second initial annotation;

[0149] Determining the perplexity corresponding to the data to be screened according to the fitting degree.

[0150] According to one or more embodiments of the present disclosure, Example 5 provides the method of Example 4, and comparing the second initial annotation with the data training set to confirm the fitting degree of the second initial annotation includes:

[0151] Analyzing the second initial annotation to determine the corresponding entry data;

[0152] Comparing the entry data with the annotation data corresponding to the data training set to generate the word frequency corresponding to the entry data;

[0153] Determining the fitting degree corresponding to the second initial annotation according to the word frequency.

[0154] According to one or more embodiments of the present disclosure, Example 6 provides the method of Example 1, and the method includes:

[0155] Obtaining initial data from an initial database;

[0156] In the case that there is no corresponding annotation information for the initial data, using the initial data as the data to be screened.

[0157] According to one or more embodiments of the present disclosure, Example 7 provides a data screening device, which includes:

[0158] A first generation module, configured to identify the data to be screened by using an identification model, and generate a first initial annotation;

[0159] A training module, configured to add the data to be screened and the corresponding first initial annotation to the data training set corresponding to the identification model, and train the identification model based on the data training set;

[0160] A second generation module, configured to generate the perplexity of the data to be screened by using the trained identification model;

[0161] A determination module, configured to determine the data to be screened as target data when the perplexity is greater than a preset perplexity threshold.

[0162] According to one or more embodiments of the present disclosure, Example 8 provides the device of Example 7. The screening device further includes an execution module, and the execution module includes:

[0163] A first generation sub-module, configured to use the target data as the data to be identified, and repeatedly execute the steps of identifying the data to be identified by using the identification model until training the identification model based on the data training set, so as to generate a trained identification model.

[0164] A screening sub-module, configured to screen other data to be screened based on the trained identification model.

[0165] According to one or more embodiments of the present disclosure, Example 9 provides the device of Example 8. The screening sub-module may further be configured to:

[0166] Divide other data to be screened into a plurality of data subsets with a preset number.

[0167] Take each data subset as the data to be identified in turn, and repeatedly execute the steps of identifying the data to be identified by using the identification model until generating a trained identification model.

[0168] Obtain the data training set corresponding to the trained identification model.

[0169] Generate a target data set based on the data in the data training set.

[0170] According to one or more embodiments of the present disclosure, Example 10 provides the device of Example 7. The second generation module includes:

[0171] A second generation sub-module, configured to identify the data to be screened by using the trained identification model, and generate a second initial annotation.

[0172] A first determination sub-module, configured to compare a second initial annotation with a data training set to confirm the fitting degree of the second initial annotation.

[0173] A second determination sub-module, configured to determine the perplexity corresponding to the data to be screened according to the fitting degree.

[0174] According to one or more embodiments of the present disclosure, Example 11 provides the apparatus of Example 10. The first determination sub-module may further be configured to:

[0175] Parse the second initial annotation to determine the corresponding entry data;

[0176] Compare the entry data with the annotation data corresponding to the data training set to generate the word frequency corresponding to the entry data;

[0177] Determine the fitting degree corresponding to the second initial annotation according to the word frequency.

[0178] According to one or more embodiments of the present disclosure, Example 12 provides the apparatus of Example 7. The screening apparatus further includes a determination module, and the determination module is configured to:

[0179] Obtain initial data from an initial database;

[0180] In the case that there is no corresponding annotation information for the initial data, use the initial data as the data to be recognized.

[0181] According to one or more embodiments of the present disclosure, Example 13 provides a computer-readable medium, on which a computer program is stored, and when the program is executed by a processing device, the steps of the method according to any one of Examples 1-6 are implemented.

[0182] According to one or more embodiments of the present disclosure, Example 14 provides an electronic device, including:

[0183] A storage device, on which at least one computer program is stored;

[0184] At least one processing device, configured to execute the at least one computer program in the storage device to implement the steps of the method according to any one of Examples 1-6.

[0185] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principle. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solution formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present disclosure.

[0186] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limitations on the scope of the present disclosure. Certain features that are described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, the various features that are described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.

[0187] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims. Regarding the apparatus in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method, and will not be elaborated here.

Claims

1. A data screening method, characterized in that, Including: Using a recognition model to recognize the data to be recognized, generating a first initial annotation, where the data to be recognized is speech data and the recognition model is a speech recognition model; Adding the data to be recognized and the corresponding first initial annotation to the data training set corresponding to the recognition model, and training the recognition model based on the data training set; Generating the perplexity of the data to be screened through the trained recognition model; the higher the perplexity of the data to be screened indicates the lower the similarity degree of the data to be screened to other data in the data training set; When the perplexity is greater than a preset perplexity threshold, determining the data to be screened as target data; The data to be screened is speech data, and the target data is used to train the recognition model; Among them, generating the perplexity of the data to be screened through the trained recognition model includes: using the trained recognition model to recognize the data to be screened, generating a second initial annotation; comparing the second initial annotation with the data training set to confirm the fitting degree of the second initial annotation; determining the perplexity corresponding to the data to be screened according to the fitting degree; the fitting degree is inversely proportional to the perplexity.

2. The screening method according to claim 1, characterized in that, The method includes: Taking the target data as the data to be recognized, and repeatedly executing the step of using the recognition model to recognize the data to be recognized until the step of training the recognition model based on the data training set to generate the trained recognition model; Screening other data to be screened based on the trained recognition model.

3. The screening method according to claim 2, characterized in that, The screening of other data to be screened based on the trained recognition model includes: Dividing the other data to be screened into a preset number of data subsets; Taking each data subset as the data to be recognized in turn, and repeatedly performing the above steps of using the recognition model to recognize the data to be recognized until the step of generating the trained recognition model; Obtaining the data training set corresponding to the trained recognition model; Generating a target data set based on the data in the data training set.

4. The screening method according to claim 1, characterized in that, The comparing the second initial annotation with the data training set to confirm the fitting degree of the second initial annotation includes: Parsing the second initial annotation to determine the corresponding entry data; Comparing the entry data with the annotation data corresponding to the data training set to generate the word frequency corresponding to the entry data; Determining the fitting degree corresponding to the second initial annotation according to the word frequency.

5. The screening method according to claim 1, characterized in that, The method includes: Obtaining initial data from an initial database; When there is no corresponding annotation information for the initial data, taking the initial data as the data to be recognized.

6. A data screening device, characterized in that, Including: A first generation module, configured to use a recognition model to recognize the data to be recognized, generate a first initial annotation, where the data to be recognized is speech data and the recognition model is a speech recognition model; A training module, configured to add the data to be recognized and the corresponding first initial annotation to the data training set corresponding to the recognition model, and train the recognition model based on the data training set; A second generation module, configured to generate a perplexity of the data to be screened through the trained recognition model; the higher the perplexity of the data to be screened indicates that the data to be screened is less similar to other data in the data training set; A determination module, configured to determine the data to be screened as target data when the perplexity is greater than a preset perplexity threshold; The data to be screened is voice data, and the target data is used to train the recognition model; Wherein, the second generation module is further configured to recognize the data to be screened through the trained recognition model to generate a second initial annotation; compare the second initial annotation with the data training set to confirm the fitting degree of the second initial annotation; according to the fitting degree, determine the perplexity corresponding to the data to be screened; the fitting degree is inversely proportional to the perplexity.

7. The screening device according to claim 6, characterized in that, The apparatus further includes a screening module, configured to: Use the target data as the data to be recognized, and repeatedly execute the step of recognizing the data to be recognized through the recognition model until the step of training the recognition model based on the data training set to generate the trained recognition model; Screen other data to be screened based on the trained recognition model.

8. A computer-readable medium, on which a computer program is stored, characterized in that, When the program is executed by a processing device, it implements the steps of the method according to any one of claims 1-5.

9. An electronic device, characterized in that Including: A storage device, on which at least one computer program is stored; At least one processing device, configured to execute the at least one computer program in the storage device to implement the steps of the method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Title text generation method and device, computer storage medium and electronic equipment

    CN111753533A

  • Dual monolingual cross-entropy-delta filtering of noisy parallel data and use thereof

    US20210026919A1