A Data Quality Inspection Method for Minor Languages ​​Based on Multidimensional Confidence

By employing a multi-dimensional confidence-based data quality inspection method for minority languages, combined with text, audio, and phonetic consistency inspection modules, the problem of low efficiency and difficulty in ensuring the quality of minority language data quality inspection is solved. This achieves efficient and refined quality inspection results and is suitable for data quality inspection of large-scale databases.

CN115906003BActive Publication Date: 2026-04-03UNIV OF SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-08
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

In existing speech recognition systems, the quality inspection efficiency for minority language data is low and the quality is difficult to guarantee. In particular, for rare languages, the quality inspection time span is large, and the randomness and subjectivity of manual quality inspection are strong, making it difficult to meet the quality inspection needs of large-scale databases.

Method used

A multi-dimensional confidence-based method for quality inspection of minority language data is adopted. By constructing formatted data, and using pre-trained text, audio and pronunciation consistency quality inspection modules, confidence scores of text, audio and pronunciation consistency are calculated respectively, and unqualified data are automatically filtered out. Combined with manual quality inspection, the quality inspection effect and efficiency are ensured.

Benefits of technology

It achieves refined quality inspection from three dimensions: text, audio, and phonetic consistency, which improves inspection efficiency and ensures quality. Especially in the case of rare languages, it can efficiently complete the initial screening of data and reduce the time and cost of manual quality inspection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115906003B_ABST
    Figure CN115906003B_ABST
Patent Text Reader

Abstract

This invention discloses a method for quality inspection of minority language data based on multidimensional confidence, comprising: Step S1, constructing formatted data: constructing audio and corresponding labeled text into audio-text pairs; Step S2, labeled text attribute quality inspection: extracting labeled text, using a pre-trained text attribute confidence model to obtain a text confidence score; if the score is greater than a preset threshold, the quality inspection is qualified; otherwise, manual quality inspection is required; Step S3, audio attribute quality inspection: extracting audio, using a pre-trained audio attribute confidence model to obtain an audio confidence score; if the score is greater than a preset threshold, the quality inspection is qualified; otherwise, manual quality inspection is required; Step S4, pronunciation consistency quality inspection: using a pre-trained speech recognition confidence model to obtain a pronunciation consistency confidence score for the audio-text pairs; if the score is greater than a preset threshold, the quality inspection is qualified; otherwise, manual quality inspection is required. This method can improve the quality inspection quality and efficiency of finished databases or large-scale labeled data, saving quality inspectors' time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data quality inspection for speech recognition systems, and more particularly to a method for quality inspection of data in less commonly spoken languages. Background Technology

[0002] With the deepening of intelligent information technology, artificial intelligence has been applied and popularized in many aspects of life. Correspondingly, the algorithms and model structures of artificial intelligence are becoming increasingly complex, relying on more training data for iterative updates. Although unsupervised and weakly supervised methods have made some breakthroughs in recent years, supervised data remains one of the most effective means of improving model performance, based on data effectiveness.

[0003] With the increasing demand for data, supervised data can be purchased and produced through procurement or self-construction. However, due to the large volume of data, it is often difficult to effectively control the quality of supervised data. Therefore, supervised data quality control is becoming an increasingly important and indispensable step.

[0004] In existing speech recognition systems, the databases used mainly consist of audio and its corresponding annotated text. For quality inspection tasks involving ultra-large-scale databases, the current approach primarily relies on data sampling, which involves manually inspecting a certain percentage of data from the database to be inspected using stratified sampling or random sampling methods.

[0005] The existing data quality inspection methods mainly rely on manual sampling and quality control, which has the following main drawbacks:

[0006] 1) Manual sampling is highly random and cannot guarantee the effectiveness of the overall data.

[0007] 2) Some subjective quality inspection indicators are not confirmed in written form, which may lead to inconsistent quality of manual data inspection.

[0008] 3) When dealing with less commonly spoken languages, especially rare ones, the difficulty of manual quality inspection increases dramatically. The overall quality inspection process is lengthy and inefficient.

[0009] In view of this, the present invention is hereby proposed. Summary of the Invention

[0010] The purpose of this invention is to provide a multidimensional confidence-based method for quality inspection of minority language data, which can improve the quality and efficiency of quality inspection of finished databases or large-scale labeled data, save the time of quality inspectors, and improve the effectiveness of data use, thereby achieving the effect of cost reduction and efficiency improvement in the data quality inspection process.

[0011] The objective of this invention is achieved through the following technical solution:

[0012] A method for quality inspection of minority language data based on multidimensional confidence includes:

[0013] Step S1, construct formatted data:

[0014] The audio data to be inspected is matched with the corresponding labeled text data to construct a formatted audio-text pair data.

[0015] Step S2, quality inspection of labeled text attributes:

[0016] The labeled text data is extracted from the audio-text pair data constructed in step S1 and sent to the labeled text attribute detection module. The labeled text attribute detection module uses a pre-trained text attribute confidence model to calculate the text confidence score of the labeled text attributes of the labeled text data. If the text confidence score is greater than a preset text qualification threshold, the labeled text data is determined to be qualified. Audio attribute quality inspection is then performed on the audio-text pair data corresponding to the qualified labeled text data. If the text confidence score is less than the preset text qualification threshold, the labeled text data is determined to be unqualified. Manual quality inspection is then performed on the unqualified labeled text data. If the manual quality inspection is qualified, audio attribute quality inspection is then performed on the audio-text pair data corresponding to the qualified labeled text data. If the manual quality inspection is unqualified, the audio-text pair data corresponding to the unqualified labeled text data is determined to be unqualified data.

[0017] Step S3, Audio Attribute Quality Inspection:

[0018] The audio data from the audio-text pairs that passed the quality inspection in step S2 is extracted and sent to the audio attribute quality inspection module. The audio attribute quality inspection module uses a pre-trained audio attribute confidence model to calculate the audio confidence score of the audio attributes of the audio data. If the audio confidence score is greater than a preset audio pass threshold, the audio data is determined to pass the quality inspection, and the audio-text pairs corresponding to the qualified audio data are subjected to pronunciation consistency quality inspection. If the audio confidence score is less than the preset audio pass threshold, the audio data is determined to fail the quality inspection, and the audio data that fails the quality inspection is subjected to manual quality inspection. If the manual quality inspection is qualified, the audio-text pairs corresponding to the qualified audio data are subjected to pronunciation consistency quality inspection. If the manual quality inspection is unqualified, the audio-text pairs corresponding to the unqualified audio data are determined to be unqualified data.

[0019] Step S4, Phonetic Consistency Quality Inspection:

[0020] The audio-text pairs that pass the quality inspection in step S3 are sent to the pronunciation consistency quality inspection module. The pronunciation consistency quality inspection module uses a pre-trained speech recognition confidence model to calculate the pronunciation consistency confidence score of the audio and text consistency alignment relationship of the audio-text pairs. If the pronunciation consistency confidence score is greater than the preset pronunciation consistency pass threshold, the audio-text pairs are determined to pass the quality inspection. If the pronunciation consistency confidence score is less than the preset pronunciation consistency pass threshold, the audio-text pairs are determined to fail the quality inspection. The audio and text consistency alignment relationship of the audio-text pairs that fail the quality inspection is then manually inspected. If the manual inspection passes, the audio-text pairs are determined to pass the quality inspection. If the manual inspection fails, the audio-text pairs that fail the quality inspection are determined to be unqualified data.

[0021] Compared with existing technologies, the multidimensional confidence-based method for quality inspection of minority language data provided by this invention has the following advantages:

[0022] 1) We conducted a meticulous overall data quality inspection across three dimensions: text, audio, and phonetic consistency. For each dimension, the attribute quality inspection was conducted using different dimensions and scales to ensure the quality inspection results met the requirements of actual database construction.

[0023] 2) Quality inspections in different dimensions all adopt model recognition methods to replace manual data quality inspection, which greatly improves the efficiency of quality inspection while ensuring the quality inspection effect.

[0024] 3) Data quality inspection combines model recognition with manual processing. In the context of scarce labeled data, it can effectively complete the initial screening of data while ensuring the quality inspection effect and efficiency. Attached Figure Description

[0025] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0026] Figure 1 A flowchart of a method for quality inspection of minority language data based on multidimensional confidence, provided in an embodiment of the present invention.

[0027] Figure 2 This is a schematic diagram of the speech recognition confidence model of the multidimensional confidence-based method for quality inspection of minority language data provided in an embodiment of the present invention. Detailed Implementation

[0028] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the specific content of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments, which do not constitute a limitation of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of the present invention.

[0029] First, the following explanations are provided for the terms that may be used in this article:

[0030] The term "and / or" means that either or both can be achieved simultaneously. For example, X and / or Y means that it includes both "X" or "Y" as well as the three cases of "X and Y".

[0031] The terms "comprising," "including," "containing," "having," or other similar semantic descriptions should be interpreted as non-exclusive inclusion. For example, including a technical feature element (such as raw material, component, ingredient, carrier, dosage form, material, size, part, component, mechanism, device, step, process, method, reaction conditions, processing conditions, parameter, algorithm, signal, data, product or article of manufacture, etc.) should be interpreted as including not only the expressly listed technical feature element, but also other technical feature elements that are not expressly listed and are well-known in the art.

[0032] The term "composed of" excludes any technical features not expressly listed. When used in a claim, it closes the claim to exclude all technical features other than those expressly listed, except for associated conventional impurities. If the term appears only in a clause of a claim, it limits the claim to the elements expressly listed in that clause; elements recited in other clauses are not excluded from the overall claim.

[0033] Unless otherwise explicitly specified or limited, the terms "installation," "connection," "linking," and "fixing," etc., should be interpreted broadly. For example, they can refer to fixed connections, detachable connections, or integral connections; they can refer to mechanical connections or electrical connections; they can refer to direct connections or indirect connections through an intermediate medium; and they can refer to the internal connection between two components. Those skilled in the art can understand the specific meaning of the above terms in this document according to the specific circumstances.

[0034] When concentration, temperature, pressure, size, or other parameters are expressed as numerical ranges, such ranges should be understood to specifically disclose all ranges formed by any pairing of upper limits, lower limits, or preferred values ​​within that range, regardless of whether the range is explicitly stated; for example, if the numerical range "2 to 8" is stated, then that range should be interpreted to include ranges such as "2 to 7", "2 to 6", "5 to 7", "3 to 4 and 6 to 7", "3 to 5 and 7", "2 and 5 to 7", etc. Unless otherwise stated, the numerical ranges described herein include both their endpoints and all integers and fractions within that range.

[0035] The terms “center,” “longitudinal,” “lateral,” “length,” “width,” “thickness,” “upper,” “lower,” “front,” “back,” “left,” “right,” “vertical,” “horizontal,” “top,” “bottom,” “inner,” “outer,” “clockwise,” and “counterclockwise” indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience and simplification of description and do not imply that the device or component referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this document.

[0036] The following is a detailed description of the multidimensional confidence-based method for quality inspection of minority language data provided by this invention. Contents not described in detail in the embodiments of this invention are prior art known to those skilled in the art. Where specific conditions are not specified in the embodiments of this invention, they should be performed according to conventional conditions in the art or conditions recommended by the manufacturer. Reagents or instruments used in the embodiments of this invention, unless otherwise specified by the manufacturer, are all commercially available conventional products.

[0037] like Figure 1 As shown, this embodiment of the invention provides a method for quality inspection of minority language data based on multidimensional confidence, including the following steps:

[0038] Step S1, construct formatted data:

[0039] The audio data to be inspected is matched with the corresponding labeled text data to construct a formatted audio-text pair data.

[0040] Step S2, quality inspection of labeled text attributes:

[0041] The labeled text data is extracted from the audio-text pair data constructed in step S1 and sent to the labeled text attribute detection module. The labeled text attribute detection module uses a pre-trained text attribute confidence model to calculate the text confidence score of the labeled text attributes of the labeled text data. The text confidence score is used to determine whether the labeled text meets the database construction requirements. Specifically, if the text confidence score is greater than a pre-set text qualification threshold, the labeled text data is determined to be qualified. Audio attribute quality inspection is then performed on the audio-text pair data corresponding to the qualified labeled text data. If the text confidence score is less than the pre-set text qualification threshold, the labeled text data is determined to be unqualified. Manual quality inspection is then performed on the unqualified labeled text data. If the manual quality inspection is qualified, audio attribute quality inspection is then performed on the audio-text pair data corresponding to the qualified labeled text data. If the manual quality inspection is unqualified, the audio-text pair data corresponding to the unqualified labeled text data is determined to be unqualified data.

[0042] Step S3, Audio Attribute Quality Inspection:

[0043] The audio data from the audio-text pairs that passed the quality inspection in step S2 is extracted and sent to the audio attribute quality inspection module. The audio attribute quality inspection module uses a pre-trained audio attribute confidence model to calculate the audio confidence score of the audio attributes of the audio data. The audio confidence score is used to determine whether the audio meets the database construction requirements. Specifically, if the audio confidence score is greater than a pre-set audio qualification threshold, the audio data is determined to be qualified, and the audio-text pairs corresponding to the qualified audio data are subjected to pronunciation consistency quality inspection. If the audio confidence score is less than the pre-set audio qualification threshold, the audio data is determined to be unqualified, and the unqualified audio data is subjected to manual quality inspection. If the manual quality inspection is qualified, the audio-text pairs corresponding to the qualified audio data are subjected to pronunciation consistency quality inspection. If the manual quality inspection is unqualified, the audio-text pairs corresponding to the unqualified audio data are determined to be unqualified data.

[0044] Step S4, Phonetic Consistency Quality Inspection:

[0045] The audio-text pairs that pass the quality inspection in step S3 are sent to the pronunciation consistency quality inspection module. The pronunciation consistency quality inspection module uses a pre-trained speech recognition confidence model to calculate the pronunciation consistency confidence score of the audio and text consistency alignment relationship of the audio-text pairs. Based on the pronunciation consistency confidence score, it is determined whether the audio + standard text pair meets the database construction requirements. Specifically, if the pronunciation consistency confidence score is greater than the preset pronunciation consistency pass threshold, the audio-text pairs are determined to pass the quality inspection. If the pronunciation consistency confidence score is less than the preset pronunciation consistency pass threshold, the audio-text pairs are determined to fail the quality inspection. The audio and text consistency alignment relationship of the audio-text pairs that fail the quality inspection is then manually inspected. If the manual inspection passes, the audio-text pairs are determined to pass the quality inspection. If the manual inspection fails, the audio-text pairs that fail the quality inspection are determined to be unqualified data.

[0046] As can be seen, the thresholds involved in each step of the above method can be determined according to the actual quality inspection requirements, such as based on historical quality inspection results or based on the knowledge of the quality inspector.

[0047] As can be seen, in each step of the above method, for data that is determined to be unqualified in quality inspection, which is data that does not meet the requirements for establishing a database, the data can be returned to the data supplier or deleted.

[0048] In step S1 of the above method, the audio data to be inspected and the corresponding labeled text data are matched in the following manner to construct formatted audio-text pairs, including:

[0049] If there are N samples to be inspected, the audio data is labeled as follows:

[0050] ;

[0051] The labeled text data corresponding to each audio data point are marked as follows:

[0052] ;

[0053] Then, each audio data to be inspected is matched one-to-one with its corresponding labeled text data to obtain N formatted audio-text pairs. Each audio-text pair is represented by [A - T]. All N audio-text pairs are represented as follows:

[0054] [ ], [ ], [ ]…[ ].

[0055] In step S2 of the above method, the confidence score of the labeled text attributes of the labeled text data is calculated using a pre-trained text attribute confidence model in the following manner:

[0056] The text confidence features of the labeled text data are extracted and used as input to a pre-trained text attribute confidence model. The confidence score corresponding to the input text confidence features is calculated through the text attribute confidence model.

[0057] The text attribute confidence model described above employs either a transformer neural network model or a conformer neural network model. The training and application methods for this text attribute confidence model are essentially the same as existing transformer neural network models or conformer neural network models, and will not be detailed here.

[0058] In the above method, the extracted text confidence features of the labeled text data are constructed by concatenating the domain information features, lexical information features, sentence structure features, and sensitive word features of the labeled text; among them,

[0059] Domain information features of the labeled text The result is calculated using the following formula:

[0060] , ;

[0061] Where N is the total number of fields involved in all the annotated texts of the quality inspection, and the total number of fields is determined manually by language experts; It is the probability that the labeled text belongs to the i-th domain, as predicted by a pre-trained domain classification model. The domain classification model is a classification model built based on the text domain training set data and using an LSTM or BERT neural network. It is the difficulty coefficient of the annotation accuracy obtained by statistically analyzing the annotated text of the i-th domain on the training set, and is represented by the proportion of the amount of text data collected in the domain to the total amount of the entire text training set; These are the hidden layer vectors before the Softmax function of the domain classification model, representing the information features of different domains;

[0062] Lexical information features of the annotated text The result is calculated using the following formula:

[0063] ;

[0064] Where N represents the number of words in the sentence of the annotated text; The highest word frequency of the vocabulary in the training set of the corresponding language text of the labeled text; This represents the word frequency of the i-th word in the annotated text sentence; This indicates the total number of words in the list of localized vocabulary. This represents the sum of word frequencies in the annotated text that match the localized vocabulary; and These are hyperparameters, whose specific values ​​are determined through parameter adjustment during the training of the text attribute confidence model;

[0065] The sentence structure features of the annotated text The result is calculated using the following formula:

[0066] ;

[0067] in, and These represent the depths of the syntactic parsing tree and the semantic dependency tree of the annotated text, respectively. and These represent the breadth of the syntactic parsing tree and the semantic dependency tree of the annotated text, respectively; and These represent the number of different node labels in the syntactic parsing tree and semantic dependency tree of the annotated text, respectively; and These are hyperparameters, whose specific values ​​are determined through parameter adjustment during the training of the text attribute confidence model;

[0068] The sensitive word features of the annotated text The result is calculated using the following formula:

[0069] ;

[0070] Where f is the hidden layer vector feature before the output layer of the pre-trained sensitive word classification model, representing the probability prediction result of sensitive words in the labeled text; the sensitive word classification model is obtained by training an NLP model based on the collected sensitive word text training data. The NLP model can adopt commonly used model structures such as BERT, which will not be elaborated further. The embedding representation of each word in the annotated text (also known as word embedding representation); Part-of-speech features of each word in the labeled text predicted for Named Entity Recognition (NER).

[0071] In the above method, a transfer-based parsing algorithm is used to construct the syntactic parsing tree and semantic dependency tree of the annotated text sentences; it should be understood that other parsing algorithms can also be used to construct the syntactic parsing tree and semantic dependency tree of the annotated text sentences, which does not constitute a limitation of the present invention;

[0072] Sentence structure features of annotated text that can determine sentence repetition rate The result is calculated using the following formula:

[0073] ;

[0074] Where N is the total number of clusters after clustering the sentence structure features of the labeled text in the data to be inspected; It is a characteristic representation of the cluster centers; It is the sum of the number of sentences in the current cluster whose distance meets the preset similarity threshold; This represents the total number of sentences in the current cluster.

[0075] In step S3 of the above method, the confidence score of the audio attributes of the audio data is calculated using a pre-trained audio attribute confidence model in the following manner:

[0076] Audio confidence features of audio data are extracted and used as input to a pre-trained audio attribute confidence model. The audio confidence score corresponding to the input audio confidence features is then calculated using the audio attribute confidence model.

[0077] The audio confidence features of the extracted audio data include:

[0078] Accent characteristics and quality characteristics of audio; among them,

[0079] The accent features of the audio The result is calculated using the following formula:

[0080] ;

[0081] Where N is the total number of accents contained in the audio data collected in the database; It is the probability of the i-th accent predicted by a pre-trained acoustic classification model that classifies different accents; the acoustic classification model is obtained by training a neural network model (NN) based on collected and labeled training data of different accents. The NN model can be a common structure such as CNN, transformer or conformer, which will not be elaborated here. It is the difficulty coefficient of each accent obtained statistically from the training set of the acoustic classification model that classifies different accents, and is represented by the proportion of the amount of audio data collected in the database to the total amount of the entire audio training set.

[0082] The quality characteristics of the audio The result is calculated using the following formula:

[0083] ;

[0084] in, This represents the hidden layer vector features before the output layer of the audio quality classification model; the audio quality classification model can be used for feature modeling of audio quality; the structure of this audio quality classification model can reuse CNNs, transformers, or converters related to ASR recognition, which will not be elaborated here. Audio data with substandard quality can be pre-constructed for training the speech quality classification model. Substandard audio data includes audio data with truncation, speaker popping, excessive background noise, stuttering, etc.; x represents the input features of the audio data, which can be any one of filter bank features, Mel-frequency cepstral coefficient features (MFCC features), or perceptual linear prediction features (PLP features).

[0085] In step S4 of the above method, the alignment confidence score of the audio-text consistency alignment relationship of the data is calculated using a pre-trained speech recognition confidence model in the following manner:

[0086] The audio features of the audio text pairs that passed the quality inspection in step S3 are extracted and used as input to a pre-trained speech recognition confidence model. The alignment confidence score of the audio and text consistency alignment relationship of the audio text pairs is calculated by the speech recognition confidence model. Preferably, the extracted audio features are conventional features such as filterbank and MFCC.

[0087] In step S4 above, the extracted audio features and corresponding labeled text are used as input, and the corresponding speech recognition information is extracted through an existing speech recognition baseline system. The speech recognition information includes: the acoustic model score of each word. Language model score Posterior probability score and the duration of each word ;

[0088] The acoustic model score of each word obtained Language model score Posterior probability score and the duration of each word As input to the speech recognition confidence model, the correctness of each word's labeling is output. The output of each correctly labeled word is then processed by the softmax function to obtain the corresponding confidence score for phonetic consistency. It is known that the meaning of each word differs across languages; for example, in Chinese it can be a Chinese character, while in English it is a word. Therefore, each word mentioned above refers to a possible modeling unit corresponding to a specific language.

[0089] In the aforementioned existing speech recognition baseline systems, the acoustic and language models used to extract audio features for speech recognition information can all be existing models of the same type.

[0090] In step S4 above, the construction method of the speech recognition confidence model is the same as that of the existing mainstream speech recognition technology solutions. The model can be trained by frame-level label prediction, such as using TDNN, CNN and other model structures, which will not be described in detail here.

[0091] To more clearly demonstrate the technical solution and its effects provided by the present invention, the following describes in detail the method for quality inspection of minority language data based on multidimensional confidence provided by the present invention, using specific embodiments.

[0092] Example 1

[0093] This invention provides a method for quality inspection of minority language data based on multidimensional confidence, the process of which is as follows: Figure 1 As shown, it includes the following steps:

[0094] Step S1: Construct formatted data:

[0095] The audio data requiring quality inspection is matched with its corresponding labeled text data to form formatted audio-text pairs, facilitating subsequent feature processing and quality inspection result judgment. For example, the database to be inspected contains N data samples, of which N are audio data. Can be marked as:

[0096] ;

[0097] in, These are the audio data respectively;

[0098] N audio data points correspond to N labeled text data points Can be marked as:

[0099] ;

[0100] in, These are the labeled text data respectively;

[0101] Ultimately, we will obtain N formatted audio text pairs represented as [A - T], which expands to:

[0102] [A - T] =[ ], [ ], [ ]…[ ].

[0103] Step S2: Text attribute quality inspection:

[0104] The audio text pairs to be inspected are constructed, and all the labeled text is extracted from the data. The labeled text attribute detection module is then sent to the labeled text attribute detection module. The labeled text attribute detection module uses a pre-trained text attribute confidence model to calculate the text confidence score of the labeled text, which is used to determine whether the labeled text meets the database construction requirements.

[0105] If the text confidence score is greater than the preset text acceptance threshold, the quality inspection is considered passed. The audio text pairs contained in the accepted text will proceed to step S3 for further audio attribute quality inspection. If the confidence score is less than the preset text acceptance threshold, the quality inspection is considered failed. The unacceptable texts are then filtered out and manually inspected by quality inspectors. The audio text pairs associated with the accepted texts will continue to step three to complete the following quality inspection process. However, the audio text pairs associated with the unacceptable texts will be marked as non-compliant or bad data and ultimately returned to the data provider or deleted.

[0106] Generally, text in a database should accurately cover a wide range of fields, have fluent sentence structures, and diverse expressions. It should also reflect the spoken language habits of native speakers during communication. Furthermore, it should comply with local laws and regulations and religious beliefs, and should not contain sensitive information such as terrorism, violence, pornography, or racial discrimination. Therefore, the text attribute confidence score is used to characterize whether the audio text meets the above conditions and complies with the requirements for database construction. To better calculate the text attribute confidence score, this invention's method is based on a text attribute confidence model to calculate the confidence score. Specifically, the text attribute confidence features of the text are extracted and used as input to the text attribute confidence model, which then calculates the final output text confidence score.

[0107] The method for constructing the text attribute confidence model is as follows:

[0108] Text attribute confidence models can employ neural network models such as transformers or conformers. A large amount of language-related text is pre-collected, covering numerous domains. Text confidence features for each sentence are extracted, and human experts annotate the corresponding confidence scores to construct a training set. The trained text attribute confidence model is then obtained using this training set. The specific training method is the same as existing transformer or conformer neural network model schemes, and will not be detailed here.

[0109] The text confidence features include: domain information features, lexical information features, sentence structure features, and sensitive word features; the extraction process for each feature is as follows:

[0110] 21) Domain information features of text:

[0111] The text in the database should accurately record key information from different fields, including various professional terms. Text in certain fields may require prior knowledge from industry experts for accurate annotation, placing higher demands on quality inspectors. This invention's method allows experts to determine the specific professional fields and their number, and by statistically analyzing the collected field text data, its proportion to the total text volume is used as the difficulty coefficient for text annotation accuracy in each field. Simultaneously, a field classification model is trained based on existing data across multiple fields to predict the probability of each field in the text data to be inspected, ultimately obtaining the field information features of the text. Specifically:

[0112] , ;

[0113] Where N is the total number of domains, It is the probability of the i-th domain predicted by the domain classification model. It is the difficulty coefficient of the i-th domain obtained statistically from the training set. These are the hidden layer vectors before the Softmax domain classification model, representing the information features of different domains;

[0114] 22) Lexical information features of text:

[0115] Text in a database generally has requirements regarding lexical richness. For example, in the realm of common spoken language, if text expansion is done using a method of "fixed sentence patterns + entity slot values," the lexical richness will be significantly reduced. On the other hand, for different languages, the text needs to express the authentic language habits and writing styles of native speakers. For example, we don't want Swedish texts to be filled with discussions about housing prices and food in Beijing's Chaoyang District. Therefore, from a lexical richness perspective, for different domain corpora, the frequency of different words appearing in the total text set can be calculated as a reference for lexical richness. Simultaneously, language experts can identify and construct localized keyword databases for different languages. For example, Korean includes 청와대 (Blue House), and Russian includes Московский Кремль (Kremlin Palace). The number of keyword hits in the database can be statistically analyzed to measure the degree of localization. Specifically:

[0116] ;

[0117] Where N represents the number of words in the input text sentence. This represents the highest word frequency of the vocabulary in the corresponding language text training set. This represents the frequency of the i-th word in the sentence. This represents the total number of words in the localized vocabulary list. This represents the total frequency of words in the input text that match words from the localized vocabulary. Additionally, and These are hyperparameters, whose specific values ​​are determined through parameter adjustments during the training of the confidence model.

[0118] 23) Sentence structure features of the text:

[0119] The text in the database needs to retain complete semantic information and possess a certain sentence structure, such as subject, verb, and object. However, different languages ​​have vastly different sentence structures and word formats, requiring the text to accurately express these elements. By analyzing the sentence structure components, a syntactic parsing tree and a semantic dependency tree can be obtained. Furthermore, by parsing these trees, feature values ​​of the text's sentence structure can be calculated. Current methods for sentence grammar analysis include graph-based parsing algorithms and transition-based parsing algorithms, all of which have achieved high performance. Without loss of generality, this invention employs a transition-based parsing algorithm to construct the syntactic parsing tree and semantic dependency tree. Specific syntactic parsing algorithms are not detailed here, and the sentence structure features are calculated based on this. Specifically, this invention uses the following method for feature value calculation:

[0120] ;

[0121] in and These represent the depths of the syntax tree and the semantic tree, respectively. and These represent the breadth of the syntax tree and the semantic tree, respectively. and These represent the number of distinct node labels in the syntax tree and syntactic tree, respectively. The depth, breadth, and number of distinct node labels in the syntax tree and semantic tree represent the complexity of the tree, and thus the complexity of the sentence meaning and syntax. Additionally, and These are hyperparameters, whose specific values ​​are learned during the training of the confidence model.

[0122] Furthermore, sentence structure features can be used to determine sentence repetition rate. Database text typically has requirements regarding sentence fluency and repetition rate. Generally, it's desirable for the database to cover a wider range of sentence types, while the repetition rate between different sentence types should not be particularly high. For example, the following text has relatively low sentence structure coverage:

[0123] Xiaoming and Xiaohua are good friends, and they made a promise to go to the park tomorrow.

[0124] Xiao Huang and Xiao Hua are good friends, and they made a promise to go to the park tomorrow.

[0125] Xiao Huang and Xiao Hua are good friends, and they made a promise to go to the park the day after tomorrow.

[0126] Xiao Huang and Xiao Hua are good friends, and they made a promise to go to the amusement park tomorrow.

[0127] Based on this, the sentence structure features of text in the database can be clustered, and the distance between different texts in each class can be calculated. This distance can be measured using Levenshtein distance, Consin distance, etc. Texts with the same or similar distances, if their statistical frequency is high, may have problematic sentence repetition. Therefore, specifically, the sentence structure features of the text can be further rewritten as follows:

[0128] ;

[0129] Where N is the total number of clusters after sentence pattern feature clustering. It is a feature representation of the cluster centers. It is the sum of sentence counts within the current cluster that are close in distance (within a defined range, e.g., L≤5). This represents the total number of sentences in the current cluster.

[0130] 24) Sensitive word features of the text:

[0131] The text in the database must comply with local laws, regulations, and religious beliefs in the respective language, and cannot contain any sensitive information. Therefore, quality control of sensitive words in the text is necessary. Different languages ​​have different grammatical structures, and the variations in words and parts of speech are also significant. For example, Russian nouns have six declensions, and the same word may have different spellings in different contexts. The method of this invention allows language experts to manually determine the list of sensitive keywords for different languages. By collecting text data containing sensitive words, a sensitive word classification model is trained to predict the probability of sensitive words in the text to be quality controlled. Simultaneously, existing named entity recognition algorithms can be used to determine the lexical attributes in the text, such as nouns and verbs, and these attributes are fused with the input features of the aforementioned sensitive word classification model to obtain the final sensitive word feature representation.

[0132] ;

[0133] Where f is the hidden layer vector feature before the output layer of the sensitive word classification model, and its input is the embedding representation of each word in the text. Part-of-speech features predicted by NER .

[0134] Step S3: Audio Attribute Quality Inspection:

[0135] After the text quality check is passed, audio attribute quality check is performed. The audio data is extracted from the constructed audio text pair data and sent to the audio attribute quality check module. Using the audio attribute confidence model, the audio confidence score of the audio data is calculated to determine whether the audio data meets the database construction requirements.

[0136] If the audio confidence score is greater than the preset audio pass threshold, the quality inspection is deemed passed. Audio data that passes the quality inspection will have its audio-text pairs transferred to step S4 for further quality inspection. If the audio confidence score is less than the preset audio pass threshold, the quality inspection is deemed failed, and the failed audio data is filtered out for manual quality inspection by a quality inspector. Audio data that passes the manual quality inspection will have its associated audio-text pairs continue to step S4 to complete the following pronunciation consistency quality inspection process. However, audio data that fails the manual quality inspection will have its associated audio-text pairs marked as non-compliant or bad data, and will ultimately be returned to the data provider or deleted.

[0137] Generally, audio data in a database should guarantee recording quality, with signal-to-noise ratio and energy amplitude meeting database construction requirements. Furthermore, purchased voice data should not contain truncated segments, leading to semantic incompleteness. In certain scenarios, additional requirements may apply to audio data, such as accent coverage and background noise types. To better calculate audio attribute confidence, this invention's method is based on an audio attribute confidence model to calculate confidence scores. Specifically, the audio attribute confidence features of the audio data are extracted and used as input to the audio attribute confidence model. The final output audio confidence score is then calculated using the audio attribute confidence model.

[0138] The method for constructing the audio attribute confidence model is as follows:

[0139] The audio attribute confidence model can employ an Encode-Deocder neural network model. This involves pre-collecting a large amount of language-related audio data, covering numerous speakers and native speakers from different regions, and extracting the audio confidence features for each sentence. Simultaneously, the confidence scores are manually labeled to construct a training set. The trained audio attribute confidence model is then obtained using this training set. The specific training method is the same as existing Encode-Deocder neural network model schemes and will not be detailed here.

[0140] The audio confidence features include: audio accent features and audio quality features. The extraction process for each feature is as follows:

[0141] 31) Accent characteristics of audio:

[0142] The audio data in the database, as required, needs to be differentiated by accent, especially for languages ​​with significant regional accents, such as Portuguese. The pronunciation of Portuguese varies considerably across different regions of Europe, South America, and Africa, necessitating the differentiation of audio data. Furthermore, in this invention's method, language experts can label the accents of different languages ​​based on their major regional pronunciations. The proportion of each accent to the total audio volume is then used as a difficulty coefficient for each accent. Simultaneously, acoustic classification models for different accents are collected and trained to predict the probability of each domain in the audio data to be inspected, ultimately yielding the accent information features of the audio, as detailed below:

[0143] ;

[0144] Where N is the total number of accents, It is the probability of the i-th accent predicted by the accent classification model. It is the difficulty coefficient of the i-th accent obtained statistically from the training set.

[0145] 32) Audio quality characteristics:

[0146] The audio data in the database needs to be untrunculated, and the relevant signal-to-noise ratio, energy amplitude, and background noise must meet the database construction requirements. The method of this invention can perform feature modeling of audio quality by training a speech quality classification model. Audio data that does not meet the quality standards can be pre-constructed, such as audio data with truncation, microphone popping, excessive background noise, stuttering, etc., for training the speech quality analysis model. The hidden layer vector before the output layer is then used as the audio quality feature. Specifically:

[0147] ;

[0148] in, It expresses the hidden layer vector features before the output layer, and x represents the input features of the audio data, which can be any of the filterbank, MFCC, PLP features, etc.

[0149] Step S4: Phonetic consistency quality inspection:

[0150] After the audio quality inspection is passed, a final phonetic consistency quality inspection is performed. The constructed audio-text pair data is sent to the audio attribute quality inspection module. Using the speech recognition confidence model, the phonetic consistency confidence score of the audio and text consistency alignment relationship is calculated to determine whether the audio data meets the database construction requirements.

[0151] If the confidence score of phonetic consistency is greater than a pre-set qualified threshold for phonetic consistency, it is determined that the quality inspection is qualified. The audio data that passes the quality inspection and its corresponding annotation text will be provided to the data user as the final qualified data. If the confidence score of phonetic consistency is less than the pre-set qualified threshold for phonetic consistency, it is determined that the quality inspection is unqualified, and the unqualified audio plus text is screened out for manual quality inspection by a quality inspector. The audio plus text that passes the manual quality inspection still belongs to qualified data and can continue to be provided to the data user. However, the audio plus text that fails the manual quality inspection will be marked as non-compliant or bad data and will ultimately be sent back to the data provider or deleted;

[0152] Generally speaking, phonetic consistency is an important part of supervised data: existing supervised neural network models all need to let the model learn the mapping relationship between the input audio and the corresponding annotation text. To better quality-check phonetic consistency, the method of the present invention calculates the confidence score of phonetic consistency based on a speech recognition confidence model. Specifically, when calculating, it is necessary to extract the audio features of the audio data, such as conventional filterbank, MFCC, etc., as the input of the speech recognition confidence model, and calculate the confidence score of phonetic consistency through the output of the speech recognition confidence model.

[0153] The speech recognition confidence model described above adopts model structures such as TDNN and CNN. Its construction method is the same as that of existing mainstream speech recognition model solutions, and frame-level label prediction is used for model training, which will not be elaborated here.

[0154] The process of extracting the confidence score of phonetic consistency is as follows:

[0155] Speech recognition usually uses Viterbi decoding to obtain the final speech recognition text. Therefore, the decoding path of the speech recognition model can be constrained and defined as a text path that only contains the annotation. Then, the acoustic model score of each word can be obtained in this process , the language model score , the posterior probability score , and the duration of each word and other speech recognition information. The definitions and score calculation methods of the above-mentioned modules are basically the same as those of the traditional scheme and will not be elaborated here. Through an existing speech recognition baseline system, the input speech signal and the corresponding recognition annotation text content are used to extract various score features by the acoustic model, language model, etc., including: , , and are used as the input of the recognition confidence model, and whether each word is correctly annotated is used as the output or target of the speech recognition confidence model to update and train the model parameters, and a speech recognition confidence model with a structure as Figure 2 is obtained, as shown in the following formula:

[0156] ;

[0157] in, These are the weight coefficients for the corresponding acoustic model score, language model score, posterior probability score, and duration of each word, which can be iterated and updated during training until the optimal parameters are found.

[0158] Finally, the output of the speech recognition confidence model is processed by the softmax function to obtain a probability between [0…1]. The specific value of the probability is the confidence score of the corresponding pronunciation consistency.

[0159] In summary, the quality inspection method of the present invention has at least the following advantages:

[0160] 1) We conducted a meticulous overall data quality inspection across three dimensions: text, audio, and phonetic consistency. For each dimension's attribute quality inspection, we analyzed and considered different dimensions and scales to ensure the quality inspection results met the requirements of actual database construction.

[0161] 2) The quality inspection of different dimensions adopts the model solution to replace manual data quality inspection, which greatly improves the quality inspection efficiency while ensuring the quality inspection effect.

[0162] 3) Data quality inspection combines model and manual methods, which can effectively complete the initial screening of data when labeled data is scarce, while ensuring the quality inspection effect and efficiency.

[0163] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0164] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims. The information disclosed in the background section is intended only to enhance the understanding of the overall background technology of the present invention and should not be construed as an admission or implication in any way that such information constitutes prior art known to those skilled in the art.

Claims

1. A method for quality inspection of minority language data based on multidimensional confidence, characterized in that, include: Step S1, construct formatted data: The audio data to be inspected is matched with the corresponding labeled text data to construct a formatted audio-text pair data. Step S2, quality inspection of labeled text attributes: The labeled text data is extracted from the audio-text pair data constructed in step S1 and sent to the labeled text attribute detection module. The labeled text attribute detection module uses a pre-trained text attribute confidence model to calculate the text confidence score of the labeled text attributes of the labeled text data. If the text confidence score is greater than a preset text qualification threshold, the labeled text data is determined to be qualified. Audio attribute quality inspection is then performed on the audio-text pair data corresponding to the qualified labeled text data. If the text confidence score is less than the preset text qualification threshold, the labeled text data is determined to be unqualified. Manual quality inspection is then performed on the unqualified labeled text data. If the manual quality inspection is qualified, audio attribute quality inspection is then performed on the audio-text pair data corresponding to the qualified labeled text data. If the manual quality inspection is unqualified, the audio-text pair data corresponding to the unqualified labeled text data is determined to be unqualified data. In step S2, the confidence score of the labeled text attributes of the labeled text data is calculated using a pre-trained text attribute confidence model in the following manner: The text confidence features of the labeled text data are extracted and used as input to a pre-trained text attribute confidence model. The confidence score corresponding to the input text confidence features is calculated through the text attribute confidence model. The extracted text confidence features of the annotated text data are constructed by concatenating the domain information features, lexical information features, sentence structure features, and sensitive word features of the annotated text; among them, Domain information features of the labeled text The result is calculated using the following formula: , ; Where N is the total number of fields involved in all the annotated texts of the quality inspection, and the total number of fields is determined manually by language experts; It is the probability that the labeled text belongs to the i-th domain, as predicted by a pre-trained domain classification model. The domain classification model is a classification model built based on the text domain training set data and using an LSTM or BERT neural network. It is the difficulty coefficient of the annotation accuracy obtained by statistically analyzing the annotated text of the i-th domain on the training set, and is represented by the proportion of the amount of text data collected in the domain to the total amount of the entire text training set; These are the hidden layer vectors before the Softmax function of the domain classification model, representing the information features of different domains; Lexical information features of the annotated text The result is calculated using the following formula: ; Where N represents the number of words in the sentence of the annotated text; The highest word frequency of the vocabulary in the training set of the corresponding language text of the labeled text; This represents the word frequency of the i-th word in the annotated text sentence; This indicates the total number of words in the list of localized vocabulary. This represents the sum of word frequencies in the annotated text that match the localized vocabulary; and These are hyperparameters, whose specific values ​​are determined through parameter adjustment during the training of the text attribute confidence model; The sentence structure features of the annotated text The result is calculated using the following formula: ; in, and These represent the depths of the syntactic parsing tree and the semantic dependency tree of the sentences in the annotated text, respectively. and These represent the breadth of the syntactic parsing tree and the semantic dependency tree of the annotated text, respectively; and These represent the number of different node labels in the syntactic parsing tree and semantic dependency tree of the annotated text, respectively; and These are hyperparameters, whose specific values ​​are determined through parameter adjustment during the training of the text attribute confidence model; The sensitive word features of the annotated text The result is calculated using the following formula: ; Where f is the hidden layer vector feature before the output layer of the pre-trained sensitive word classification model, representing the probability prediction result of sensitive words in the labeled text; the sensitive word classification model is obtained by training an NLP model based on the collected sensitive word text training data, and the NLP model adopts the commonly used BERT model structure; This refers to the embedding of each word in the annotated text. Part-of-speech features for each word in the labeled text for named entity recognition; A transition-based parsing algorithm is used to construct syntactic parsing trees and semantic dependency trees for sentences in annotated text. Sentence structure features of annotated text that can determine sentence repetition rate The result is calculated using the following formula: ; Where N is the total number of clusters after clustering the sentence structure features of the labeled text in the data to be inspected; It is a characteristic representation of the cluster centers; It is the sum of the number of sentences in the current cluster whose distance meets the preset similarity threshold; This represents the total number of sentences in the current cluster; Step S3, Audio Attribute Quality Inspection: The audio data from the audio-text pairs that passed the quality inspection in step S2 is extracted and sent to the audio attribute quality inspection module. The audio attribute quality inspection module uses a pre-trained audio attribute confidence model to calculate the audio confidence score of the audio attributes of the audio data. If the audio confidence score is greater than a preset audio pass threshold, the audio data is determined to pass the quality inspection, and the audio-text pairs corresponding to the qualified audio data are subjected to pronunciation consistency quality inspection. If the audio confidence score is less than the preset audio pass threshold, the audio data is determined to fail the quality inspection, and the audio data that fails the quality inspection is subjected to manual quality inspection. If the manual quality inspection is qualified, the audio-text pairs corresponding to the qualified audio data are subjected to pronunciation consistency quality inspection. If the manual quality inspection is unqualified, the audio-text pairs corresponding to the unqualified audio data are determined to be unqualified data. In step S3, the confidence score of the audio attributes of the audio data is calculated using a pre-trained audio attribute confidence model in the following manner: The audio confidence features of the audio data are extracted and used as input to a pre-trained audio attribute confidence model. The audio confidence score corresponding to the input audio confidence features is then calculated using the audio attribute confidence model. The audio confidence features of the extracted audio data include: Accent characteristics and quality characteristics of audio; among them, The accent features of the audio The result is calculated using the following formula: ; Where N is the total number of accents contained in the audio data collected in the database; It is the probability of the i-th accent predicted by a pre-trained acoustic classification model that classifies different accents. The acoustic classification model is obtained by training an NN neural network model based on collected and labeled training data of different accents. The NN neural network model can be any one of the CNN model, transformer model, or conformer model. It is the difficulty coefficient of each accent obtained statistically from the training set of the acoustic classification model that classifies different accents, and is represented by the proportion of the amount of audio data collected in the database to the total amount of the entire audio training set. The quality characteristics of the audio The result is calculated using the following formula: ; in, The vector features represent the hidden layer features before the output layer of the audio quality classification model. The audio quality classification model is used for feature modeling of audio quality. This model reuses any one of the CNN, transformer, or conformer models related to ASR recognition. The model is trained using pre-constructed audio data with substandard quality, including audio data with truncation, microphone popping, excessive background noise, stuttering, or stammering. x represents the input features of the audio data, which can be any one of filter bank features, Mel-frequency cepstral coefficient features, or perceptual linear prediction features. Step S4, Phonetic Consistency Quality Inspection: The audio-text pairs that pass the quality inspection in step S3 are sent to the pronunciation consistency quality inspection module. The pronunciation consistency quality inspection module uses a pre-trained speech recognition confidence model to calculate the pronunciation consistency confidence score of the audio and text consistency alignment relationship of the audio-text pairs. If the pronunciation consistency confidence score is greater than the preset pronunciation consistency pass threshold, the audio-text pairs are determined to pass the quality inspection. If the pronunciation consistency confidence score is less than the preset pronunciation consistency pass threshold, the audio-text pairs are determined to fail the quality inspection. The audio and text consistency alignment relationship of the audio-text pairs that fail the quality inspection is then manually inspected. If the manual inspection is successful, the audio-text pairs are determined to pass the quality inspection. If the manual inspection is unsuccessful, the audio-text pairs are determined to be unsuccessful data. In step S4, the alignment confidence score of the audio-text consistency alignment relationship of the data is calculated using a pre-trained speech recognition confidence model in the following manner: Extract the audio features of the audio data of the audio text pairs that have passed the quality inspection in step S3, and use them as input to the pre-trained speech recognition confidence model. The alignment confidence score of the audio and text consistency alignment relationship of the audio text pairs is calculated by the speech recognition confidence model. The extracted audio features and corresponding labeled text are used as input, and the corresponding speech recognition information is extracted through an existing speech recognition baseline system. The speech recognition information includes: the acoustic model score of each word. Language model score Posterior probability score and the duration of each word ; The acoustic model score of each word obtained Language model score Posterior probability score and the duration of each word As input to the speech recognition confidence model, whether each word is correctly labeled is output. The output of whether each word is correctly labeled is processed by the softmax function to obtain the corresponding confidence score of pronunciation consistency.

2. The method for quality inspection of minority language data based on multidimensional confidence as described in claim 1, characterized in that, In step S1, the audio data to be inspected and the corresponding labeled text data are matched in the following manner to construct formatted audio-text pairs, including: If there are N samples to be inspected, the audio data is labeled as follows: ; The labeled text data corresponding to each audio data point are marked as follows: ; Then, each audio data to be inspected is matched one-to-one with the corresponding labeled text data to obtain N formatted audio-text pairs. Each audio-text pair is represented by [A - T]. All N audio-text pairs are represented as follows: [ ], [ ], [ ]…[ ]。 3. The method for quality inspection of minority language data based on multidimensional confidence as described in claim 1, characterized in that, The text attribute confidence model uses either a transformer neural network model or a conformer neural network model.

Citation Information

Patent Citations

  • Call speech quality inspection method and device, computer equipment and storage medium

    CN109151218A

  • Quality test method, apparatus and device for customer service recording, and storage medium

    WO2021169423A1