Marking inspection method and related device, electronic equipment, and storage medium
By extracting the language features of the pronunciation and dividing the molecule, the problems of high cost and long time of manual inspection are solved, and efficient and accurate labeling inspection of the speech recognition system is achieved.
Patent Information
- Application Number
- CN202210482241.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-05
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2042-05-05
AI Technical Summary
In the prior art, the labeling inspection method of the speech recognition system relies on manual inspection, resulting in high cost and time consumption. At the same time, the probability of wrong inspection is increased due to human ear fatigue, which affects the quality of inspection.
By extracting the language features of the pronunciation to be examined, dividing the pronunciation into a subset based on the language features, and evaluating the quality based on the labeling inspection results of the subset, reducing the number of full inspections, and using the feature extraction network to reduce the probability of recognition disorder in different languages.
It reduces the cost and time of labeling inspections, and at the same time improves the quality of inspections, reduces the recognition disorder caused by frequent listening to different languages by the human ear, and improves the inspection efficiency.
Smart Images

Figure CN115050350B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a labeling inspection method and related devices, electronic equipment, and storage media. Background Art
[0002] With breakthroughs in speech recognition technology enabled by deep learning, language recognition technology, as a front-end technology for speech recognition, has been widely used in various industries, including education, entertainment, healthcare, and transportation. However, because language recognition technology is a typical data-driven, supervised learning pattern recognition technology, the quantity and quality of training data directly impacts the system's recognition performance.
[0003] Currently, manual inspection of speech to be inspected is typically performed manually, a process typically divided into two steps. The first involves a full inspection, which involves labeling and checking the entire volume of speech to be inspected. The second involves quality control, which involves randomly selecting a certain percentage of the full inspection data for re-inspection. Manual inspection is not only labor-intensive but also time-consuming. Furthermore, prolonged listening to speech causes auditory fatigue, increasing the probability of false positives and impacting the quality of labeling and inspection. Therefore, improving inspection quality while reducing inspection costs and time has become a pressing issue. Summary of the Invention
[0004] The main technical problem solved by this application is to provide a marking inspection method and related devices, electronic equipment, and storage media, which can improve inspection quality while reducing inspection costs and inspection time.
[0005] In order to solve the above technical problems, the first aspect of the present application provides a labeling and checking method, including: extracting the language features of several speech to be checked respectively; wherein the speech to be checked is labeled with a language category; based on the language features of each speech to be checked, the several speech to be checked are divided into at least one subset; based on the labeling and checking results of some of the speech to be checked in the subset, the labeling quality of the subset is obtained.
[0006] In order to solve the above technical problems, the second aspect of the present application provides a labeling and inspection device, including: a feature extraction module, a set division module and a quality determination module, the feature extraction module is used to respectively extract the language features of several to-be-inspected speech; wherein the to-be-inspected speech is annotated with a language category; the set division module is used to divide the several to-be-inspected speech into at least one subset based on the language features of each to-be-inspected speech; the quality determination module is used to obtain the labeling quality of the subset based on the labeling inspection results of some of the to-be-inspected speech in the subset.
[0007] In order to solve the above technical problems, the third aspect of the present application provides an electronic device, including a memory and a processor coupled to each other, wherein the memory stores program instructions, and the processor is used to execute the program instructions to implement the marking inspection method of the first aspect above.
[0008] In order to solve the above technical problems, the fourth aspect of the present application provides a computer-readable storage medium storing program instructions that can be executed by a processor, and the program instructions are used to implement the marking inspection method of the first aspect.
[0009] The above scheme extracts language features from several speech sounds to be inspected, and the speech sounds to be inspected are annotated with language categories. Based on the language features of each speech sound to be inspected, the speech sounds to be inspected are divided into at least one subset. The annotation quality of the subset is then determined based on the annotation inspection results of the subset of speech sounds to be inspected. Since the speech sounds to be inspected are divided into at least one subset based on the language features of the speech sounds to be inspected, it is possible to ensure that the actual language category of the speech sounds to be inspected in each subset is substantially consistent. On this basis, on the one hand, each time a single subset is inspected, the probability of hearing different languages is greatly reduced, preventing the human ear from misrecognizing different languages due to frequent hearing. Different subsets can even be inspected by different personnel to further reduce the probability of hearing different languages, which helps improve the quality of the annotation inspection. On the other hand, each time a single subset is inspected, only a portion of the speech sounds to be inspected in the subset need to be inspected, rather than the entire subset. This helps reduce the cost and time of annotation inspection. Since the listening time is reduced, it also helps improve the quality of the annotation inspection. Therefore, it can improve the inspection quality while reducing the inspection cost and time. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 This is a flow chart of an embodiment of the marking inspection method of the present application;
[0011] Figure 2 This is a schematic diagram of the process of an embodiment of the marking inspection of the present application;
[0012] Figure 3 is a flow chart of an embodiment of training a feature extraction network;
[0013] Figure 4 1 is a schematic diagram of a process for training a feature extraction network according to an embodiment;
[0014] Figure 5 This is a schematic diagram of the framework of an embodiment of the marking inspection device of the present application;
[0015] Figure 6 This is a schematic diagram of the framework of an embodiment of the electronic device of the present application;
[0016] Figure 7It is a schematic diagram of a framework of an embodiment of a computer-readable storage medium of the present application. DETAILED DESCRIPTION
[0017] The following describes the embodiments of the present application in detail with reference to the accompanying drawings.
[0018] In the following description, for the purpose of explanation rather than limitation, specific details such as specific system structures, interfaces, and technologies are provided to facilitate a thorough understanding of the present application.
[0019] The terms "system" and "network" are often used interchangeably in this document. The term "and / or" is simply a description of an association between related objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " generally indicates that the related objects are in an "or" relationship. Furthermore, "multiple" in this document means two or more than two.
[0020] See also Figure 1 , Figure 1 It is a flowchart of an embodiment of the marking inspection method of the present application.
[0021] Specifically, the following steps may be included:
[0022] Step S11: extracting language features of several speech to be checked respectively.
[0023] In the disclosed embodiment, the speech to be checked can be marked with a language category. It should be noted that the speech to be checked can be marked with a correct language category or an incorrect language category, that is, a number of speech to be checked can contain speech with correct language markings or speech with incorrect language markings, which is not limited here. For example, in a real-life scenario, the speech to be checked, "Hello, how's the weather today?", may be marked with the correct language category "Chinese" or the incorrect language category "English". Other situations can be deduced by analogy, and no further examples will be given here.
[0024] In one implementation scenario, in order to further reduce the probability of the human ear hearing different languages, the language categories marked on the above-mentioned several voices to be checked can all be the same. Specifically, when it is necessary to mark and check N voices to be checked, the N voices to be checked can be divided into at least one set according to the language categories marked on each voice to be checked. In other words, the voices to be checked that are marked with the same language category can be divided into the same set. On this basis, the marking and inspection method in the embodiment of the present disclosure can be used to inspect the voices to be checked in each set respectively. In this way, each time a single set is checked, the probability of the human ear hearing different languages can be greatly reduced. It is even possible to assign different sets to different people to be responsible for them, so as to avoid recognition confusion caused by frequent listening to different languages by the human ear, which helps to greatly improve the quality of marking and inspection.
[0025] In one implementation scenario, in order to improve compatibility with voices from different sources, the voice sources of the above-mentioned several voices to be checked can be different. It should be noted that the voice source represents the collection channel of the voice to be checked, which can specifically include but is not limited to: mobile communication, instant messaging software (such as WeChat, etc.), satellite calls, etc., which are not limited here. Of course, the voice sources of the above-mentioned several voices to be checked can be partially the same or partially different. For ease of understanding, taking three voices to be checked as an example, the voice to be checked "The weather is good today" has a voice source of "mobile communication", the voice to be checked "It will rain tomorrow" has a voice source of "instant messaging software", and the voice to be checked "Help me book a flight for tomorrow" has a voice source of "mobile communication". It can be seen that the voice sources of the first two are different, and the voice sources of the latter two are also different, while the voice sources of the first and third voices are the same. Other situations can be deduced by analogy, and no examples are given here one by one.
[0026] In one implementation scenario, a spectrogram of the speech to be checked can be extracted, and based on the spectrogram, language-related feature information such as frequency distribution and average duration of different phonemes can be extracted to obtain the language characteristics of the speech to be checked. It should be noted that the aforementioned "frequency distribution" and "average duration of different phonemes" are merely artificial features that may be designed in actual applications and do not limit other artificial features. In other words, in actual applications, other artificial features can also be designed based on the spectrogram as part of the language characteristics.
[0027] In one implementation scenario, in order to improve the efficiency and accuracy of language feature extraction, a feature extraction network can be pre-trained so that language features can be extracted from the speech to be inspected based on the feature extraction network. The feature extraction network can be specifically obtained by multi-task joint training based on sample speech. The multi-task can at least include a language category prediction task and a speech source prediction task. The sample speech can be labeled with a sample language category and a sample speech source. In the joint training process, the prediction loss of the speech source prediction task is gradient reversed, so that the feature extraction network can extract as much feature information related to the language category as possible and as little feature information related to the speech source as possible, thereby effectively avoiding the speech to be inspected from being divided into different subsets due to different sources. This can reduce the number of subsets and help improve the efficiency of labeling inspection. It should be noted that the training process of the feature extraction network can refer to the following relevant embodiments and will not be described here.
[0028] In a specific implementation scenario, in the process of feature extraction based on the feature extraction network, the first feature extraction can be performed based on the acoustic features of each speech frame of the speech to be checked to obtain the first features of each speech frame. For example, the acoustic features may include but are not limited to Filter Bank features (i.e., FB features), MFCC (Mel Frequency Cepstral Coefficient) features, etc., which are not limited here. In addition, taking the speech length of the speech to be checked as 8s as an example, when each speech frame is set to 10ms, each speech to be checked can be divided into 800 speech frames. On this basis, the acoustic features can be extracted for each speech frame separately, so that the acoustic features of 800 speech frames can be obtained for each speech to be checked. In order to improve the efficiency of feature extraction, each time the feature is extracted, the M speech to be checked can be used as a batch to extract the language features at the same time, that is, when performing the first feature extraction, the input features can be It is expressed as M*800*d0, where d0 represents the feature dimension of the acoustic feature (such as 128, 256, etc.). Here, when the speech frame is divided in other ways, the same can be applied, and examples are not given here one by one. It should be noted that the first feature extraction is used to extract the deep semantics of the speech frame based on the acoustic feature. In addition, after the first feature extraction, the output feature can be expressed as M*800*d1, where d1 represents the feature dimension of the first feature. On this basis, for each speech frame, a second feature extraction can be performed based on the first feature of the speech frame and its adjacent frames to obtain the second feature of each speech frame. Example Still taking the aforementioned 800 speech frames as an example, taking the i-th speech frame as an example, in the process of extracting its second feature, the second feature of the i-th speech frame can be obtained based on the first feature of the i-th speech frame itself, the first feature of the i-1-th speech frame and the first feature of the i+1-th speech frame, that is, the second feature extraction is used for time series modeling based on the first feature. In addition, after the second feature extraction, the output feature can be expressed as M*800*d2, where d2 represents the feature dimension of the second feature. After obtaining the second feature of each speech frame, feature extraction can be performed based on the second feature of each speech frame. Mapping is performed to obtain the language features of the speech to be processed. It should be noted that feature mapping may include but is not limited to: performing global average pooling or global maximum pooling on the second features of each speech frame belonging to the same speech to be processed. Still taking the aforementioned 800 speech frames as an example, the second features of the 800 speech frames belonging to the same speech to be inspected can be globally pooled, that is, after feature mapping, the output features can be expressed as M*1*d2. The above method first performs the first feature extraction to extract deep semantics, and then performs the second feature extraction to perform time series modeling, which helps to enhance feature expression.
[0029] In a specific implementation scenario, please refer to Figure 2 , Figure 2 This is a process diagram of an embodiment of the present application's marking inspection. Figure 2 As shown, the feature extraction network can include a sequentially connected convolutional neural network, a bidirectional long short-term memory network, and an embedding network, wherein the convolutional neural network is used to perform the first feature extraction, the bidirectional long short-term memory network is used to perform the second feature extraction, and the embedding network is used to perform feature mapping. It should be noted that, as mentioned above, the embedding network can specifically be composed of network layers such as global average pooling or global maximum pooling, which is not limited here. The above method, by sequentially connecting the convolutional neural network, the bidirectional long short-term memory network, and the embedding network, constitutes a feature extraction network, thereby being able to rely on the excellent feature transformation ability of the convolutional neural network to extract deep semantics, and rely on the bidirectional long short-term memory network's ability to model the temporal sequence of related features to perform time series modeling. Connecting the two in series can greatly enhance feature expression. In addition, since the subsequent subset division is only based on the extracted language features, whether the actual language category of the speech to be inspected is included in the language category of the training data for training the feature extraction network does not affect the subsequent subset division and quality measurement, thereby improving the effectiveness of the inspection of various different language data.
[0030] Step S12: Divide the speech to be checked into at least one subset based on the language feature of each speech to be checked.
[0031] Specifically, the plurality of to-be-checked speech can be divided into at least one subset based on the similarity between the language features of the respective checked speech.
[0032] In one implementation scenario, the similarity of the language features between several speech sounds to be checked can be obtained, and a decision threshold can be set. If the similarity is not less than the decision threshold, the two speech sounds to be checked can be divided into the same subset. Conversely, if the similarity is greater than the decision threshold, the two speech sounds to be checked can be divided into different subsets.
[0033] In an implementation scenario, please refer to Figure 2 In order to improve the efficiency and accuracy of set division, clustering can be performed based on the language features of each speech to be checked. For example, the proximity propagation algorithm (AP) can be used to cluster the language features of several speeches to be checked into at least one category, and then the speeches to be checked to which the language features in each category belong are classified into a subset. It should be noted that the specific process of clustering using the AP algorithm can refer to the technical details of the AP algorithm, which will not be described here. In addition, in addition to using the AP algorithm for feature clustering, clustering algorithms including but not limited to DBSCAN, OPTICS, etc. can also be used, which are not limited here.
[0034] Step S13: Based on the annotation check results of part of the speech to be checked in the subset, the annotation quality of the subset is obtained.
[0035] Specifically, in order to reduce the inspection time and labor costs, for each subset, a part of the speech to be inspected can be extracted for annotation inspection first, so that the annotation quality of each subset can be obtained. Then, based on the annotation quality of each subset, it can be further decided whether it is necessary to continue to label and inspect the speech to be inspected that has not been inspected in the subset.
[0036] In one implementation scenario, for each subset, the first number of speech to be checked whose annotation results are correctly labeled as language category can be counted, and the second number of speech to be checked that has been checked in the subset can be counted. Based on this, the annotation quality of the subset can be obtained based on the ratio of the first number to the second number. For ease of description, the first number of speech to be checked that has been correctly labeled can be recorded as K, and the second number of speech to be checked that has been checked in the subset can be recorded as L. The ratio of the two is the current check accuracy rate acc_check of the subset:
[0037] acc_check = K / L……(1)
[0038] Furthermore, for the i-th subset, its current check accuracy can be recorded as acc_check(i). In the above method, the statistical annotation check result is a first number of speech to be checked whose language category is correctly labeled, and a second number of speech to be checked in the subset that has been checked is counted. Based on the ratio of the first number to the second number, the annotation quality of the subset is obtained. Therefore, the annotation quality of the subset can be measured by quantitative statistics, which helps to greatly reduce the complexity of measuring annotation quality.
[0039] In a specific implementation scenario, as mentioned above, the language categories marked by the above-mentioned several voices to be checked can be the same, then at least one subset obtained by dividing the above-mentioned several voices to be checked can all be checked by the same person. Of course, in actual application, each inspection team can be responsible for marking the voices to be checked with different language categories, such as inspection team A is responsible for marking the voices to be checked with Chinese, inspection team B is responsible for marking the voices to be checked with English, and so on, and no more examples are given here. On this basis, for each inspection team, after dividing the voices to be checked for which it is responsible into at least one subset based on the embodiment of the present disclosure, each inspector in the inspection team can be responsible for a different subset. For example, for inspection team A, after dividing the voices to be checked for which it is responsible into Q subsets based on the embodiment of the present disclosure, each inspector in inspection team A can be responsible for a different subset in the Q subsets. Other situations can be deduced by analogy, and no more examples are given here. The above division method can greatly reduce the probability of each inspector hearing different languages, avoid the human ear from frequently listening to different languages and cause recognition confusion, which helps to improve the quality of labeling inspection.
[0040] In a specific implementation scenario, in response to the ratio of the first number to the second number being not less than a preset threshold, it is determined that the annotation quality of the subset is up to standard. It should be noted that the preset threshold can be set according to actual application needs. Specifically, in the case of high requirements for annotation inspection, the preset threshold can be set to be appropriately larger, such as 90%, 95%, etc., or, in the case of relatively loose requirements for annotation inspection, the preset threshold can be set to be appropriately smaller, such as 80%, 85%, etc., which are not limited here. In the above manner, in response to the ratio of the first number to the second number being not less than the preset threshold, it is determined that the annotation quality of the subset is up to standard. The annotation quality of the subset can be determined only by numerical comparison, which is conducive to further reducing the complexity of measuring annotation quality.
[0041] In a specific implementation scenario, in response to the ratio of the first number to the second number being less than a preset threshold, the annotation quality of the subset is determined to be substandard. It should be noted that the preset threshold can be set according to actual application needs. Specifically, in the case of high requirements for annotation inspection, the preset threshold can be set to be appropriately larger, such as 90%, 95%, etc., or, in the case of relatively loose requirements for annotation inspection, the preset threshold can be set to be appropriately smaller, such as 80%, 85%, etc., which are not limited here. In the above manner, in response to the ratio of the first number to the second number being less than the preset threshold, the annotation quality of the subset is determined to be substandard. The annotation quality of the subset can be determined only by numerical comparison, which is conducive to further reducing the complexity of measuring annotation quality.
[0042] In one implementation scenario, in response to the annotation quality of the subset being substandard, a prompt may be given to continue to check the language categories annotated for the unchecked speech to be checked in the subset. It should be noted that, during the continued inspection process, you may choose to inspect all unselected speech to be checked in the subset, or you may choose to inspect some of the unselected speech to be checked in the subset. When the latter is selected, you may re-count the first number of speech to be checked whose annotation results are correct for the language category, and count the second number of speech to be checked in the subset that has been inspected, and based on the ratio of the two, update the current inspection accuracy acc_check of the subset, and re-measure the annotation quality of the subset, and based on the re-measured annotation quality, re-determine whether it is necessary to continue to perform annotation inspection on the speech to be checked in the subset that has not been inspected. In the above manner, in response to the annotation quality of the subset being substandard, a prompt may be given to continue to inspect the language categories annotated for the unchecked speech to be checked in the subset, which helps to improve the accuracy of the measured annotation quality.
[0043] In one implementation scenario, in response to the subset's annotation quality meeting the standard, a prompt is provided to terminate the subset's annotation check. It should be noted that if the subset's annotation quality meets the standard, it can be considered that the subset has passed random inspection. Since random inspections can reflect overall quality to a certain extent, the subset's annotation check can be terminated. The above-mentioned method, in response to the subset's annotation quality meeting the standard, prompting the subset to terminate the annotation check can help speed up the annotation check.
[0044] The above scheme extracts language features from several speech sounds to be inspected, and the speech sounds to be inspected are annotated with language categories. Based on the language features of each speech sound to be inspected, the speech sounds to be inspected are divided into at least one subset. The annotation quality of the subset is then determined based on the annotation inspection results of the subset of speech sounds to be inspected. Since the speech sounds to be inspected are divided into at least one subset based on the language features of the speech sounds to be inspected, it is possible to ensure that the actual language category of the speech sounds to be inspected in each subset is substantially consistent. On this basis, on the one hand, each time a single subset is inspected, the probability of hearing different languages is greatly reduced, preventing the human ear from misrecognizing different languages due to frequent hearing. Different subsets can even be inspected by different personnel to further reduce the probability of hearing different languages, which helps improve the quality of the annotation inspection. On the other hand, each time a single subset is inspected, only a portion of the speech sounds to be inspected in the subset need to be inspected, rather than the entire subset. This helps reduce the cost and time of annotation inspection. Since the listening time is reduced, it also helps improve the quality of the annotation inspection. Therefore, it can improve the inspection quality while reducing the inspection cost and time.
[0045] See also Figure 3 , Figure 3This is a flow chart of an embodiment of training a feature extraction network. Specifically, it may include the following steps:
[0046] Step S31: extracting sample features of the sample speech based on the feature extraction network.
[0047] In the disclosed embodiment, the sample features at least include feature information related to the language. It should be noted that, since the network performance of the feature extraction network for extracting language-related feature information is still relatively poor in the early stages of feature extraction network training, the sample features may also include feature information related to the source of the speech. For example, in the case where the speech source of the sample speech is instant messaging software, the sample features may also include feature information of poor quality such as discontinuous speech due to compression encoding of speech transmission; or, in the case where the speech source of the sample speech is satellite communication, the sample features may also include feature information of large speech delay caused by satellite transmission. Other situations can be deduced by analogy and will not be given one by one here.
[0048] In one implementation scenario, in order to improve the adaptability of the feature extraction network to various speech sources, speech data can be collected from different speech sources in advance. For example, speech data can be collected from speech sources such as mobile communications, instant messaging software, satellite communications, etc. In order to improve the network performance of the feature extraction network, the more language categories involved in the speech data, the better. Similarly, the more speech sources involved in the speech data, the better. In addition, each piece of speech data can contain only one language category. Furthermore, the data volume of speech data of different language categories can be as balanced as possible, and the total effective time meets the training quantity requirements. For example, the effective time of speech data training data of each language category can be no less than 20 hours.
[0049] In a specific implementation scenario, in order to further enrich the amount of sample data and improve the generalization of network training, the speech data can be enhanced by adding white noise, changing speed, disturbing, etc. to obtain speech data corresponding to the speech data, and the original collected speech data and the enhanced speech data can be used together as sample speech.
[0050] In an implementation scenario, please refer to Figure 4 , Figure 4 FIG. 1 is a schematic diagram of a process for training a feature extraction network according to an embodiment of the present invention. Figure 4As shown, a first feature extraction can be performed based on the acoustic features of each speech frame of the sample speech (e.g., first feature extraction is performed through a convolutional neural network) to obtain the first feature of each speech frame. On this basis, for each speech frame, a second feature extraction can be performed based on the first features of the speech frame and its adjacent frames (e.g., second feature extraction is performed through a long short-term memory network) to obtain the second feature of each speech frame, so that feature mapping can be performed based on the second features of each speech (e.g., feature mapping is performed through an embedding network) to obtain the sample features of the sample speech. For details, please refer to the relevant description of extracting language features based on the feature extraction network in the aforementioned disclosed embodiment, which will not be repeated here. Therefore, for both the sample speech and the speech to be checked, sample features and language features can be extracted based on the feature extraction network respectively. For example, the two-stage feature extraction process can be summarized as follows: a first feature extraction is performed based on the acoustic features of each speech frame of the speech to be processed to obtain the first feature of each speech frame, and in the annotation inspection stage, the speech to be processed is the speech to be checked, and in the network training stage, the speech to be processed is the sample speech. On this basis, for each speech frame, a second feature extraction is performed based on the first features of the speech frame and its adjacent frames to obtain the second features of each speech frame. Feature mapping is then performed based on the second features of each speech frame to obtain the speech features of the speech to be processed, and the speech features include at least feature information related to the language. In addition, during the annotation and inspection phase, the speech features are the language features in the aforementioned disclosed embodiment, and during the training phase, the speech features are the sample features in the disclosed embodiment.
[0051] It should be noted that the difference between the above two stages is that in the stage of feature extraction of sample speech based on the feature extraction network, since the feature extraction network has not yet been trained and converged, the extracted sample features contain not only feature information related to the language, but also certain feature information related to the source of the speech. In the stage of feature extraction of the speech to be inspected based on the feature extraction network, since the feature extraction network has been trained and converged, the extracted language features contain very little feature information related to the source of the speech, and contain as much feature information related to the language as possible.
[0052] In one implementation scenario, each training session may include a batch of sample speech to be fed into the feature extraction network for training. The number of samples in each batch may not be limited. For example, each batch may include 16 sample speech, 32 sample speech, etc., which are not limited here. In addition, the sample speech in each batch may be randomly sampled from the full amount of data (i.e., the aforementioned original speech data and enhanced speech data). Furthermore, each batch may include sample speech of at least two different languages. Similarly, each batch may include sample speech from at least two different sources to improve the generalization performance of the feature extraction network.
[0053] Step S32: performing language category prediction based on the sample features to obtain a predicted language category, and performing speech source prediction based on the sample features to obtain a predicted speech source.
[0054] For details, please continue to refer to Figure 4 The sample features can be fed into a language category prediction network to obtain a predicted language category. Similarly, the sample features can be fed into a speech source prediction network to obtain a predicted speech source. It should be noted that the language category prediction network may include, but is not limited to, a fully connected layer, and the speech source prediction network may also include, but is not limited to, a fully connected layer.
[0055] Step S33: Obtain a first loss based on the difference between the predicted language category and the sample language category, obtain a second loss based on the difference between the predicted speech source and the sample speech source, and obtain a network loss based on the first loss and the second loss.
[0056] In the disclosed embodiment, the network loss is positively correlated with the first loss and negatively correlated with the second loss, that is, the greater the first loss and the smaller the second loss, the greater the network loss, or the smaller the first loss and the greater the second loss, the smaller the network loss. Therefore, in the subsequent network parameter optimization process, by minimizing the network loss, the feature extraction network can be forced to extract as much feature information related to the language category as possible and as little feature information related to the voice source as possible, so as to improve the network performance of the feature extraction network, and help to avoid the same language to be inspected from being divided into different subsets due to different voice sources after the feature extraction of the speech to be inspected, which helps to reduce the number of subsets.
[0057] In one implementation scenario, the first loss and the second loss can be calculated using a loss function such as cross entropy. For the specific calculation process, please refer to the technical details of the loss function such as cross entropy, which will not be repeated here.
[0058] In one implementation scenario, to further minimize the intra-class distance between speech features of the same language category and maximize the inter-class distance between features of different language categories, sample features of positive speech can be extracted based on a feature extraction network, and sample features of negative speech can be extracted based on a feature extraction network. The positive speech is annotated with the same sample language category as the sample speech, and the negative speech is annotated with a different sample language category. Based on this, a third loss can be derived based on the first distance between the sample features of the sample speech and the sample features of the positive speech, and the second distance between the sample features of the sample speech and the sample features of the negative speech. The third loss is positively correlated with the first distance and negatively correlated with the second distance. That is, the greater the first distance and the smaller the second distance, the greater the third loss. Conversely, the smaller the first distance and the larger the second distance, the smaller the third loss. On this basis, the network loss can be derived based on the first loss, the second loss, and the third loss. For example, the three losses described above can be summed or weighted. The above method obtains a third loss by measuring the first distance between the sample features of the sample speech and the sample features of the positive speech, and the second distance between the sample features of the sample speech and the sample features of the negative speech. The third loss is positively correlated with the first distance and negatively correlated with the second distance. This can constrain the feature extraction network to minimize the intra-class distance of the same language category and may expand the inter-class distance of different language categories, which helps to improve the network performance of the feature extraction network.
[0059] In a specific implementation scenario, a speech sample labeled with the same language category as the currently trained speech sample can be selected from the previously prepared speech samples as a positive example speech, and a speech sample labeled with a different language category than the currently trained speech sample can be selected from the previously prepared speech samples as a negative example speech. Of course, positive and negative examples can also be selected from other sources, which is not limited here.
[0060] In a specific implementation scenario, for each sample speech, multiple positive speech samples that are labeled with the same sample language category can be obtained. In this case, the sample features of the multiple positive speech samples can be extracted based on the feature extraction network, and the feature distances between the sample features of the sample speech and the sample features of each positive speech sample can be calculated, and the average value of these feature distances can be taken as the first distance. Similarly, for each sample speech, multiple negative speech samples that are labeled with different sample language categories can also be obtained. In this case, the sample features of the multiple negative speech samples can be extracted based on the feature extraction network, and the feature distances between the sample features of the sample speech and the sample features of each negative speech sample can be calculated, and the average value of these feature distances can be taken as the second distance. It should be noted that the feature distance can be obtained based on cosine similarity. For the convenience of description, the sample speech can be denoted as x n , and the positive or negative example of the sample speech is recorded as x k , the mathematical function of the feature extraction network can be recorded as f(·), and the sample feature of the sample speech can be recorded as f(x n ), the sample features of positive or negative speech can be recorded as f(x k ), then the characteristic distance d(f(x n ),f(x k )) can be expressed as:
[0061]
[0062] In a specific implementation scenario, for the convenience of description, the nth sample speech belonging to the i-th sample language category in the current batch can be recorded as And the sample speech The sample language category marked is And the sample speech The source of the labeled sample speech is On this basis, the network loss Loss can be expressed as:
[0063]
[0064] In the above formula (3), CELoss represents the first loss, DomainLoss represents the second loss, TripletLoss represents the third loss, σ represents the normalization function (e.g., softmax), g1 represents the mathematical function of the language category prediction network, g2 represents the mathematical function of the speech source prediction network, d + Indicates the first distance, d _represents the second distance, α represents the margin parameter, which is used to control the degree of dispersion between positive and negative examples (for example, it can be 0.2). N represents the batch size.
[0065] Step S34: Based on the network loss, adjust the network parameters of the feature extraction network.
[0066] Specifically, after calculating the network loss, the network parameters of the feature extraction network can be adjusted based on the network loss using optimization methods such as gradient descent. The specific adjustment process of the network parameters can refer to the technical details of optimization methods such as gradient descent, which will not be repeated here. In addition, during the training process of the feature extraction network, the learning rate can be set to 0.1, which is not limited here. Through the above steps, the feature extraction network can be iteratively trained until the training converges, and the network parameters of the feature extraction network with the lowest network loss can be fixed to obtain the feature extraction network used in the subsequent annotation inspection process.
[0067] The above scheme extracts sample features of the sample speech based on the feature extraction network, and the sample features include at least feature information related to the language. On this basis, the language category is predicted based on the sample features to obtain the predicted language category, and the speech source is predicted based on the sample features to obtain the predicted speech source. Thus, based on the difference between the predicted language category and the sample language category, a first loss is obtained, and based on the difference between the predicted speech source and the sample speech source, a second loss is obtained. Based on the first loss and the second loss, a network loss is obtained, and the network loss is positively correlated with the first loss and negatively correlated with the second loss. Based on the network loss, the network parameters of the feature extraction network are adjusted. Therefore, in the subsequent network parameter optimization process, by minimizing the network loss, the feature extraction network can be forced to extract as much feature information related to the language category as possible and as little feature information related to the speech source as possible, thereby improving the network performance of the feature extraction network. This also helps to avoid the division of the same language to be inspected into different subsets due to different speech sources after the feature extraction of the speech to be inspected, thereby helping to reduce the number of subsets.
[0068] See also Figure 5 , Figure 5 Schematic diagram of the framework of an embodiment of a labeling and checking device 50 of the present application. Labeling and checking device 50 includes a feature extraction module 51, a set partitioning module 52, and a quality determination module 53. Feature extraction module 51 is configured to extract language features of a plurality of speech items to be checked, wherein the speech items to be checked are labeled with language categories. Set partitioning module 52 is configured to partition the plurality of speech items to be checked into at least one subset based on the language features of each speech item to be checked. Quality determination module 53 is configured to determine the labeling quality of the subset based on the labeling and checking results of a portion of the speech items to be checked in the subset.
[0069] The above scheme extracts language features from several speech sounds to be inspected, and the speech sounds to be inspected are annotated with language categories. Based on the language features of each speech sound to be inspected, the speech sounds to be inspected are divided into at least one subset. The annotation quality of the subset is then determined based on the annotation inspection results of the subset of speech sounds to be inspected. Since the speech sounds to be inspected are divided into at least one subset based on the language features of the speech sounds to be inspected, it is possible to ensure that the actual language category of the speech sounds to be inspected in each subset is substantially consistent. On this basis, on the one hand, each time a single subset is inspected, the probability of hearing different languages is greatly reduced, preventing the human ear from misrecognizing different languages due to frequent hearing. Different subsets can even be inspected by different personnel to further reduce the probability of hearing different languages, which helps improve the quality of the annotation inspection. On the other hand, each time a single subset is inspected, only a portion of the speech sounds to be inspected in the subset need to be inspected, rather than the entire subset. This helps reduce the cost and time of annotation inspection. Since the listening time is reduced, it also helps improve the quality of the annotation inspection. Therefore, it can improve the inspection quality while reducing the inspection cost and time.
[0070] In some disclosed embodiments, the quality determination module 53 includes a quantity statistics submodule, which is used to count the first number of speech to be checked whose labeling check results show that the language category is correctly labeled, and to count the second number of speech to be checked that has been checked in the subset; the quality determination module 53 includes a quality measurement submodule, which is used to obtain the labeling quality of the subset based on the ratio of the first number to the second number.
[0071] Therefore, the statistical annotation check result is the first number of speech to be checked whose language category is correctly labeled, and the second number of speech to be checked in the subset that has been checked is counted, and based on the ratio of the first number to the second number, the annotation quality of the subset is obtained. Therefore, the annotation quality of the subset can be measured by quantitative statistics, which is beneficial to greatly reduce the complexity of measuring the annotation quality.
[0072] In some disclosed embodiments, the quality metric submodule includes a first metric unit for determining that the annotation quality of the subset is up to standard in response to a ratio of the first number to the second number being not less than a preset threshold; the quality metric submodule includes a second metric unit for determining that the annotation quality of the subset is not up to standard in response to a ratio of the first number to the second number being less than a preset threshold.
[0073] Therefore, in response to the ratio of the first number to the second number being not less than a preset threshold, the annotation quality of the subset is determined to be up to standard, and / or, in response to the ratio of the first number to the second number being less than a preset threshold, the annotation quality of the subset is determined to be substandard. The annotation quality of the subset can be determined only by numerical comparison, which is conducive to further reducing the complexity of measuring annotation quality.
[0074] In some disclosed embodiments, the annotation checking device 50 includes a first prompt module for prompting to continue checking the language category marked by the unchecked speech to be checked in the subset in response to the annotation quality of the subset being unsatisfactory; the annotation checking device 50 includes a second prompt module for prompting to end the annotation check of the subset in response to the annotation quality of the subset being up to standard.
[0075] Therefore, in response to the annotation quality of the subset being unsatisfactory, a prompt is given to continue checking the language categories annotated for the unchecked speech in the subset, which helps to improve the accuracy of measuring the annotation quality; and in response to the annotation quality of the subset being up to standard, a prompt is given to end the annotation check for the subset, which can help to speed up the annotation check.
[0076] In some disclosed embodiments, language features are extracted from the speech to be inspected based on a feature extraction network, and the feature extraction network is obtained by multi-task joint training based on sample speech. The multi-tasks include at least a language category prediction task and a speech source prediction task. The sample speech is annotated with a sample language category and a sample speech source, and during the joint training process, the prediction loss of the speech source prediction task is gradient reversed.
[0077] Therefore, the feature extraction network can extract as much feature information related to the language category as possible and as little feature information related to the speech source as possible, thereby effectively avoiding the subsequent division of the speech to be inspected into different subsets due to different sources, thereby reducing the number of subsets and helping to improve the efficiency of labeling and inspection.
[0078] In some disclosed embodiments, the labeling and checking device 50 includes a sample extraction module for extracting sample features of the sample speech based on a feature extraction network; wherein the sample features include at least feature information related to the language; the labeling and checking device 50 includes a sample prediction module for predicting the language category based on the sample features to obtain the predicted language category, and predicting the speech source based on the sample features to obtain the predicted speech source; the labeling and checking device 50 includes a first measurement module for obtaining a first loss based on the difference between the predicted language category and the sample language category, the labeling and checking device 50 includes a second measurement module for obtaining a second loss based on the difference between the predicted speech source and the sample speech source, the labeling and checking device 50 includes a loss measurement module for obtaining a network loss based on the first loss and the second loss; wherein the network loss is positively correlated with the first loss and negatively correlated with the second loss; the labeling and checking device 50 includes a parameter adjustment module for adjusting the network parameters of the feature extraction network based on the network loss.
[0079] Therefore, in the subsequent network parameter optimization process, by minimizing the network loss, the feature extraction network can be forced to extract as much feature information related to the language category as possible and as little feature information related to the speech source as possible, so as to improve the network performance of the feature extraction network. This will also help to avoid the same language to be inspected being divided into different subsets due to different speech sources after the feature extraction of the speech to be inspected, which will help to reduce the number of subsets.
[0080] In some disclosed embodiments, the annotation and inspection device 50 includes a control extraction module for extracting sample features of the positive example speech based on a feature extraction network, and extracting sample features of the negative example speech based on a feature extraction network; wherein the positive example speech is annotated with a sample language category that is the same as the sample speech, and the negative example speech is annotated with a sample language category that is different from the sample speech; the annotation and inspection device 50 includes a third measurement module for obtaining a third loss based on a first distance between the sample features of the sample speech and the sample features of the positive example speech, and a second distance between the sample features of the sample speech and the sample features of the negative example speech; wherein the third loss is positively correlated with the first distance and negatively correlated with the second distance; the loss measurement module is specifically used to obtain the network loss based on the first loss, the second loss and the third loss.
[0081] Therefore, by measuring the first distance between the sample features of the sample speech and the sample features of the positive speech, and the second distance between the sample features of the sample speech and the sample features of the negative speech, a third loss is obtained, and the third loss is positively correlated with the first distance and negatively correlated with the second distance, thereby constraining the feature extraction network to narrow the intra-class distance of the same language category as much as possible, and possibly expand the inter-class distance of different language categories, which helps to improve the network performance of the feature extraction network.
[0082] In some disclosed embodiments, the feature extraction module 51 includes a first extraction submodule for performing first feature extraction based on the acoustic features of each speech frame of the speech to be processed to obtain the first feature of each speech frame; wherein, in the annotation inspection stage, the speech to be processed is the speech to be inspected, and in the network training stage, the speech to be processed is the sample speech; the feature extraction module 51 includes a second extraction submodule for performing second feature extraction on each speech frame based on the first features of the speech frame and its adjacent frames to obtain the second feature of each speech frame; the feature extraction module 51 includes a feature mapping submodule for performing feature mapping based on the second features of each speech frame to obtain the speech features of the speech to be processed; wherein, the speech features at least include feature information related to the language.
[0083] Therefore, performing the first feature extraction to extract deep semantics and then performing the second feature extraction to perform temporal modeling can help enhance feature expression.
[0084] In some disclosed embodiments, the feature extraction network includes a sequentially connected convolutional neural network, a bidirectional long short-term memory network, and an embedding network, and the convolutional neural network is used to perform a first feature extraction, the bidirectional long short-term memory network is used to perform a second feature extraction, and the embedding network is used to perform feature mapping.
[0085] Therefore, by sequentially connecting the convolutional neural network, the bidirectional long short-term memory network and the embedding network, a feature extraction network is formed, which can rely on the excellent feature transformation ability of the convolutional neural network to extract deep semantics, and rely on the bidirectional long short-term memory network's ability to model the time sequence of related features to perform time series modeling. Connecting the two in series can greatly enhance feature expression.
[0086] See also Figure 6 , Figure 6 This is a schematic diagram of an embodiment of an electronic device 60 of the present application. Electronic device 60 includes a memory 61 and a processor 62 coupled to each other. Memory 61 stores program instructions, and processor 62 is configured to execute the program instructions to implement the steps of any of the aforementioned label inspection method embodiments. Specifically, electronic device 60 may include, but is not limited to, desktop computers, laptop computers, servers, mobile phones, tablet computers, and the like.
[0087] Specifically, the processor 62 is used to control itself and the memory 61 to implement the steps in any of the above-mentioned marking inspection method embodiments. The processor 62 can also be called a CPU (Central Processing Unit). The processor 62 may be an integrated circuit chip with signal processing capabilities. The processor 62 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. In addition, the processor 62 can be implemented by an integrated circuit chip.
[0088] The above scheme divides a number of speech to be checked into at least one subset based on the language characteristics of the speech to be checked. Therefore, it can ensure that the actual language category of the speech to be checked in each subset is basically consistent as much as possible. On this basis, on the one hand, each time a single subset is checked, the probability of listening to different languages can be greatly reduced, avoiding the human ear from misrecognizing due to frequent listening to different languages. Different subsets can even be handed over to different people for inspection to further reduce the probability of the human ear listening to different languages, which helps to improve the quality of annotation inspection. On the other hand, each time a single subset is checked, only part of the speech to be checked in the subset needs to be checked, without the need for full inspection, which helps to reduce the cost and time of annotation inspection. Moreover, due to the reduction in listening time, it also helps to improve the quality of annotation inspection. Therefore, it is possible to improve the inspection quality while reducing the inspection cost and time.
[0089] See also Figure 7 , Figure 7 Schematic diagram of a computer-readable storage medium 70 according to an embodiment of the present invention. The computer-readable storage medium 70 stores program instructions 71 that can be executed by a processor, and the program instructions 71 are used to implement the steps of any of the above-mentioned marking inspection method embodiments.
[0090] The above scheme divides a number of speech to be checked into at least one subset based on the language characteristics of the speech to be checked. Therefore, it can ensure that the actual language category of the speech to be checked in each subset is basically consistent as much as possible. On this basis, on the one hand, each time a single subset is checked, the probability of listening to different languages can be greatly reduced, avoiding the human ear from misrecognizing due to frequent listening to different languages. Different subsets can even be handed over to different people for inspection to further reduce the probability of the human ear listening to different languages, which helps to improve the quality of annotation inspection. On the other hand, each time a single subset is checked, only part of the speech to be checked in the subset needs to be checked, without the need for full inspection, which helps to reduce the cost and time of annotation inspection. Moreover, due to the reduction in listening time, it also helps to improve the quality of annotation inspection. Therefore, it is possible to improve the inspection quality while reducing the inspection cost and time.
[0091] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.
[0092] The above description of the various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced with each other and will not be repeated herein for the sake of brevity.
[0093] In the several embodiments provided in this application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation methods described above are only schematic. For example, the division of modules or units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical or other forms.
[0094] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0095] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0096] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of each embodiment method of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
Claims
1. A marking inspection method, characterized in that: include: Extracting language features of a plurality of speech to be checked respectively; wherein the speech to be checked is annotated with a language category; Dividing the plurality of speech to be checked into at least one subset based on the language feature of each of the speech to be checked; Obtaining the annotation quality of the subset based on the annotation inspection results of the portion of the to-be-inspected speech in the subset; In which, the language feature is extracted from the speech to be inspected based on a feature extraction network, and the feature extraction network is obtained by multi-task joint training based on sample speech. The multi-tasks include at least a language category prediction task and a speech source prediction task. The sample speech is annotated with a sample language category and a sample speech source, and during the joint training process, the prediction loss of the speech source prediction task is gradient reversed.
2. The method according to claim 1, characterized in that The obtaining of the annotation quality of the subset based on the annotation inspection results of the portion of the to-be-inspected speech in the subset includes: Counting a first number of the speech to be checked whose marking check result is that the language category is correctly marked, and counting a second number of the speech to be checked in the subset that has been checked; The labeling quality of the subset is obtained based on a ratio of the first number to the second number.
3. The method according to claim 2, characterized in that The obtaining, based on the ratio of the first quantity to the second quantity, the annotation quality of the subset includes: In response to a ratio of the first number to the second number being not less than a preset threshold, determining that the annotation quality of the subset meets the standard; And / or, in response to a ratio of the first number to the second number being less than a preset threshold, determining that the annotation quality of the subset is substandard.
4. The method according to claim 1, wherein After obtaining the annotation quality of the subset based on the annotation inspection results of the portion of the to-be-inspected speech in the subset, the method further includes: In response to the annotation quality of the subset being unsatisfactory, prompting to continue to inspect the language categories of the uninspected speech to be inspected in the subset; And / or, in response to the annotation quality of the subset being up to standard, a prompt is provided to end the annotation check on the subset.
5. The method according to claim 1, wherein The training steps of the feature extraction network include: Extracting sample features of the sample speech based on the feature extraction network; wherein the sample features at least include feature information related to the language; Performing language category prediction based on the sample features to obtain a predicted language category, and performing speech source prediction based on the sample features to obtain a predicted speech source; A first loss is obtained based on a difference between the predicted language category and the sample language category, a second loss is obtained based on a difference between the predicted speech source and the sample speech source, and a network loss is obtained based on the first loss and the second loss; wherein the network loss is positively correlated with the first loss and negatively correlated with the second loss; Based on the network loss, network parameters of the feature extraction network are adjusted.
6. The method according to claim 5, characterized in that Before obtaining the network loss based on the first loss and the second loss, the method further includes: Extracting sample features of positive speech based on the feature extraction network, and extracting sample features of negative speech based on the feature extraction network; wherein the positive speech is annotated with the same sample language category as the sample speech, and the negative speech is annotated with a sample language category different from the sample speech; Obtaining a third loss based on a first distance between the sample feature of the sample speech and the sample feature of the positive example speech, and a second distance between the sample feature of the sample speech and the sample feature of the negative example speech; wherein the third loss is positively correlated with the first distance and negatively correlated with the second distance; The obtaining of a network loss based on the first loss and the second loss includes: The network loss is obtained based on the first loss, the second loss and the third loss.
7. The method according to claim 1, characterized in that The feature extraction step of the feature extraction network includes: Extracting a first feature based on the acoustic features of each speech frame of the speech to be processed to obtain a first feature of each speech frame; wherein, in the annotation inspection stage, the speech to be processed is the speech to be inspected, and in the network training stage, the speech to be processed is the sample speech; For each speech frame, extract a second feature based on the first feature of the speech frame and its adjacent frames to obtain the second feature of each speech frame; Feature mapping is performed based on the second feature of each speech frame to obtain speech features of the speech to be processed; wherein the speech features at least include feature information related to the language.
8. The method according to claim 7, characterized in that The feature extraction network includes a sequentially connected convolutional neural network, a bidirectional long short-term memory network and an embedding network, and the convolutional neural network is used to perform the first feature extraction, the bidirectional long short-term memory network is used to perform the second feature extraction, and the embedding network is used to perform the feature mapping.
9. A marking inspection device, characterized in that: include: A feature extraction module is used to extract language features of a plurality of speech to be checked, wherein the speech to be checked is annotated with a language category; A set division module, configured to divide the plurality of speech to be checked into at least one subset based on the language feature of each of the speech to be checked; a quality determination module, configured to obtain the annotation quality of the subset based on the annotation inspection results of the portion of the speech to be inspected in the subset; In which, the language feature is extracted from the speech to be inspected based on a feature extraction network, and the feature extraction network is obtained by multi-task joint training based on sample speech. The multi-tasks include at least a language category prediction task and a speech source prediction task. The sample speech is annotated with a sample language category and a sample speech source, and during the joint training process, the prediction loss of the speech source prediction task is gradient reversed.
10. An electronic device, characterized in that: The method comprises a memory and a processor coupled to each other, wherein the memory stores program instructions, and the processor is used to execute the program instructions to implement the marking inspection method according to any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that Program instructions that can be executed by a processor are stored, and the program instructions are used to implement the marking inspection method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Data annotation accuracy verification method and device, electronic equipment and storage medium
CN111354340A
Voice data automatic annotation quality evaluation method
CN112435651A