Data augmentation processing method and apparatus
Patent Information
- Application Number
- CN202211032174.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-26
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2042-08-26
AI Technical Summary
[0005]本申请提供一种数据扩充处理方法及装置,以解决如何自动扩充词库中的标注数据的技术问题
[0022]第五方面,本申请还提供一种计算机程序产品,包括计算机程序,该计算机程序被处理器执行时实现第一方面所提供的任意一种可能的数据扩充处理方法。
Smart Images

Figure CN116150313B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of natural language processing, and in particular to a data augmentation processing method and apparatus. Background Technology
[0002] As internet information technology is increasingly applied in the financial sector and companies continue to strengthen their innovation efforts, market competition is becoming increasingly fierce. Managing and controlling the quality of customer service systems has become an important part of the daily work of business managers, and intelligent voice quality inspection is a major component of this.
[0003] Currently, model-based intelligent speech quality inspection methods are gaining popularity due to their high accuracy and ability to fully understand semantics. However, these methods primarily utilize supervised learning to build models for target word detection, relying heavily on large amounts of labeled data in databases. However, in existing technologies, this labeled data requires manual annotation, resulting in high costs and low efficiency.
[0004] This makes how to automatically expand the labeled data in the lexicon a technical problem that urgently needs to be solved. Summary of the Invention
[0005] This application provides a data augmentation processing method and apparatus to solve the technical problem of how to automatically augment the labeled data in a thesaurus.
[0006] In a first aspect, this application provides a data augmentation processing method, comprising:
[0007] Acquire audio files and perform speech recognition on the audio files to obtain speech recognition results;
[0008] The speech recognition results are filtered using a target vocabulary to obtain the filtered results. Multiple feature vectors are determined by extracting features from the filtered results and the speech recognition results in multiple dimensions through various expression methods. Each feature vector contains feature information in at least one dimension.
[0009] The augmented dataset is determined from the sentences corresponding to the speech recognition results and / or filtering results based on multiple feature vectors, multiple weight values, and preset similarity thresholds, and then added to the expanded database. The weight values correspond to the feature vectors. The expanded database is used for quality inspection of speech service data.
[0010] This application provides a data augmentation processing method that extracts features from speech recognition results and filtering results across multiple dimensions using various expression methods. Each expression method yields a feature vector set, or the speech recognition result and filtering result each correspond to a feature vector set. All feature vectors in these sets reflect features identical or similar to the target word. Each feature vector can be a multi-dimensional vector, where one-dimensional feature information may correspond to a multi-dimensional vector, or a multi-dimensional vector may correspond to feature information from multiple dimensions. In summary, the multiple feature vectors obtained through the above method can uncover more words or phrases with features identical or similar to the target word from more dimensions or perspectives. These words or phrases are combined into an augmented dataset and added to the augmented database. This enables automatic multi-dimensional augmentation of the labeled data in the lexicon, reducing the cost of manual data collection and creation, and improving the efficiency and richness of the labeled data in the lexicon. Furthermore, the augmented database can include the target lexicon and the augmented dataset, or it can include only the augmented dataset. In daily business management, the extended database can be called, or the extended database and the target dictionary can be used together to perform quality inspection on the recordings of voice services provided by business personnel, which can improve the efficiency and quality of business management and reduce management costs.
[0011] Secondly, this application provides a data augmentation processing apparatus, comprising:
[0012] The acquisition module is used to acquire audio files;
[0013] Processing module, used for:
[0014] Perform speech recognition on the audio file and use the recognition result as the speech recognition outcome;
[0015] The speech recognition results are filtered using a target vocabulary, and multiple features are extracted from the filtered results and the speech recognition results using various expression methods to determine multiple feature vectors. Each feature vector contains feature information of at least one dimension.
[0016] The augmented dataset is determined from the sentences corresponding to the speech recognition results based on multiple feature vectors, multiple weight values, and a preset similarity threshold. The weight values correspond to the feature vectors.
[0017] The augmented dataset is added to the augmented database, and the augmented database is used to perform quality control and supervision of voice service data.
[0018] Thirdly, this application provides an electronic device, comprising:
[0019] Memory, used to store program instructions;
[0020] The processor is configured to call and execute program instructions in the memory to perform any of the possible data expansion processing methods provided in the first aspect.
[0021] Fourthly, this application provides a storage medium in which a computer program is stored, the computer program being used to execute any of the possible data augmentation processing methods provided in the first aspect.
[0022] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, implements any of the possible data augmentation processing methods provided in the first aspect.
[0023] This application provides a data augmentation processing method and apparatus, which extracts features from speech recognition results and their filtering results from multiple dimensions through various expression methods, thereby improving the accuracy of feature information extraction for target words. Furthermore, by first expanding the basic target vocabulary and then repeatedly filtering the data to be augmented, the generalization ability of the model in this application is improved from another perspective. Attached Figure Description
[0024] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0025] Figure 1 This is a schematic diagram illustrating an application scenario of a data augmentation processing method provided in an embodiment of this application.
[0026] Figure 2 A flowchart illustrating a data augmentation processing method provided in this application;
[0027] Figure 3 A flowchart illustrating another data augmentation processing method provided for the implementation of this application;
[0028] Figure 4 A schematic diagram illustrating the process of identifying and filtering sentiment tendencies provided in an embodiment of this application;
[0029] Figure 5 A schematic diagram illustrating the first feature vector generation method provided in this application embodiment;
[0030] Figure 6 A schematic diagram illustrating the second feature vector generation method provided in this application embodiment;
[0031] Figure 7 This is a schematic diagram of the structure of a data augmentation processing device provided in an embodiment of this application;
[0032] Figure 8 This is a schematic diagram of the structure of an electronic device provided in this application.
[0033] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0034] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort, including but not limited to combinations of multiple embodiments, are within the scope of protection of this application.
[0035] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented, for example, in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0036] The following is an explanation of the technical terms used in this application:
[0037] BERT: A pre-trained language model that uses a bidirectional encoder representation from a Transformer. Pre-trained BERT representations can be fine-tuned with an additional output layer, making them suitable for building state-of-the-art models across a wide range of tasks.
[0038] NER (Named Entity Recognition) refers to the identification of entities with specific meanings in text, mainly including names of people, places, organizations, proper nouns, etc.
[0039] Word2vec: Represents words as vectors of a certain dimension.
[0040] Sensitive words: These refer to words or phrases that violate business rules and need to be detected in actual business operations.
[0041] Sensitive word detection refers to detecting whether a voice conversation contains specific words that violate company regulations and national laws and regulations. It is an important component of intelligent quality inspection.
[0042] As internet information technology is increasingly applied in the financial sector and companies continue to strengthen their innovation efforts, market competition is becoming increasingly fierce. In this intense market competition, customer service has become an important measure to reflect competitive differentiation, enhance corporate image, and increase customer satisfaction. Therefore, the management and control of customer service quality has become an important daily task for business managers, and intelligent voice quality inspection is a major component of this.
[0043] The daily customer service system generates a large amount of voice data. If this data can be used effectively and intelligent quality inspection can be carried out in accordance with the standard requirements to detect non-standard points in customer service calls, the quality of customer service and user satisfaction can be greatly improved, manual work can be reduced, and customer service personnel can be evaluated to improve the performance evaluation system.
[0044] Currently, in the field of AI (Artificial Intelligence), the main methods for sensitive word detection are rule-based and model-based approaches. Rule-based methods are relatively simple and easy to maintain, but have low accuracy and lack semantic understanding capabilities. Model-based sensitive word detection, on the other hand, is gaining popularity due to its high accuracy and ability to fully understand semantics.
[0045] However, sensitive word detection based on models is a supervised learning method that relies on a large amount of labeled data. Data labeling is the process of producing labeled data, and the amount of data has a great impact on the final effect of the model. Therefore, in practical applications, a large amount of labeled data is needed to improve the model's performance.
[0046] Currently, the main methods for augmenting labeled data are: manual annotation-based methods and deep learning model-based methods. Manual annotation-based methods are simple, easy to understand, and produce high-quality augmented corpora. However, these methods require significant manpower, are inefficient, and demand that annotators be very familiar with the business logic. Furthermore, manual annotation-based methods lack generalization ability and can only be used with a limited sensitive word database. Deep learning model-based methods can reduce human workload and have some generalization ability, but their accuracy is low with limited labeled data, and their generalization ability is also limited.
[0047] To solve the above problems, the inventive concept of this application is as follows:
[0048] When augmenting the labeled data for target words, feature information is extracted from multiple dimensions to obtain multiple feature vectors. The similarity between the feature vectors of the original data (filtered by preset filtering rules) and the unfiltered original data is compared to select the labeled data that can be used for augmentation. Preset filtering rules improve the generalization ability of this method, while automatically selecting labeled data based on similarity greatly reduces manual workload and significantly improves accuracy.
[0049] Figure 1 This is a schematic diagram illustrating an application scenario for a data augmentation processing method provided in this application. For example... Figure 1 As shown, before conducting intelligent voice quality inspection, server 30 imports recording file 40 to expand the labeled data relied upon for quality inspection supervision. This enriches the labeled data in the expanded database used for voice quality inspection, making the intelligent voice quality inspection more intelligent in recognizing target words and improving the accuracy of the inspection. The audio of the conversation between user 10 and customer service representative 20 can be transmitted to server 30 in real time, or stored in the database first, and then periodically extracted and transmitted to server 30 for intelligent voice quality inspection, assisting in supervising and evaluating the service quality of customer service representative 20. It should be noted that customer service representative 20 can also be an artificial intelligence voice system, and server 30 has a pre-installed data expansion processing program to provide data expansion processing services.
[0050] It should be noted that the data augmentation processing method provided in this application can be applied to intelligent voice quality inspection systems, automatic augmentation systems for manually annotated voice data, and automatic augmentation systems for annotated data.
[0051] The data augmentation processing method provided in this application is described in detail below:
[0052] Figure 2 This is a flowchart illustrating a data augmentation processing method provided in an embodiment of this application. Figure 2 As shown, the specific steps of this data augmentation processing method include:
[0053] S201. Obtain the audio file and perform speech recognition on the audio file to obtain the speech recognition result.
[0054] In this step, the audio file is pre-prepared, unannotated audio material.
[0055] In this embodiment, ASR (Automatic Speech Recognition) technology is used to automatically recognize the acquired audio file and convert the speech data into text data, thus obtaining the speech recognition result.
[0056] It should be noted that the speech recognition result includes multiple sentences, and during speech-to-text translation, different speaker information, or speaker information, can be identified based on audio or timbre. Each speaker is assigned a number, and a correspondence is established between each speaker and each sentence. In other words, the speech recognition result also includes speaker information.
[0057] S202. Filter the speech recognition results using the target vocabulary to obtain the filtered results. Extract features from the filtered results and the speech recognition results in multiple dimensions using various expression methods to determine multiple feature vectors.
[0058] In this step, the filtering results include at least one word or phrase that matches a target word in the target thesaurus. Further, in one possible implementation, it is also required that the word or phrase in the filtering results has the same sentiment polarity as the target word in the target thesaurus.
[0059] Each feature vector contains feature information in at least one dimension. The target vocabulary includes various target words used for voice quality inspection or other recognition purposes, such as sensitive words that violate laws and regulations, uncivilized language, discriminatory language, or sensitive words used in voice service quality inspection, such as "complaint," etc. It should be noted that target words can be set according to different quality inspection objectives or different recognition objectives, and this application does not impose any limitations.
[0060] It is worth noting that each expression method can yield a feature vector set, or the speech recognition result and the filtering result can each correspond to a feature vector set. All the feature vectors in these feature vector sets can reflect the same or similar features as the target word. Furthermore, each feature vector can be a multi-dimensional vector, and one-dimensional feature information may correspond to a multi-dimensional vector, or a multi-dimensional vector may correspond to feature information of multiple dimensions.
[0061] For example, multi-dimensional feature extraction can be performed on the filtering results and speech recognition results using three different representation methods. The feature vector set corresponding to the filtering results can be represented as {w1_vec1, w1_vec2, w1_vec3}, where w1_vec1 represents the feature vector obtained by feature extraction under the first representation method, w1_vec2 represents the feature vector obtained by feature extraction under the second representation method, and w1_vec3 represents the feature vector obtained by feature extraction under the third representation method. Similarly, the feature vector set corresponding to the speech recognition results can be represented as {w2_vec1, w2_vec2, w2_vec3}.
[0062] The speech recognition results are filtered using a target word library, including at least one of the following aspects: (1) removing sentences that do not contain the target word; (2) removing short sentences that do not contain actual meaning; (3) removing sentences that are semantically meaningless due to noise interference; (4) merging consecutive adjacent sentences of the same speaker; (5) removing polite or formulaic sentences with neutral emotional tendencies.
[0063] Understandably, users can set filtering rules according to their actual needs, without being limited to the filtering methods listed above.
[0064] After filtering the speech recognition results based on the target vocabulary and preset filtering rules, the filtered results are determined. These filtered results still contain a large number of sentences to be recognized. The next step is feature extraction.
[0065] It is worth noting that in this step, feature extraction is not only performed on the filtered results, but also directly on the original data, i.e., the speech recognition results. The objects of the two feature extractions are different, which prepares for the next step, S203, to automatically select the amplified labeled data, i.e., the amplified dataset.
[0066] It should also be noted that, during feature extraction, the embodiments of this application obtain different feature vectors through different methods. Each feature vector expression method, or generation method, corresponds to an information extraction dimension, including: word density dimension (also known as word usage frequency dimension), contextual understanding dimension, and whether the semantics are changed after the target word is replaced, etc. Those skilled in the art can employ multiple different feature vector generation methods, or utilize different models, such as the BERT model, neural network model, self-learning model, etc., to obtain different feature vectors from different dimensions according to the needs of actual applications, so as to more accurately reflect the information contained in the speech recognition results and filtering results.
[0067] Suppose that after feature extraction, the filtering result yields multiple feature vectors that can be represented by a first vector set W1, i.e., W1 = {W1_vec1, W1_vec2, W1_vec3, ..., W1_vecn}. Similarly, after feature extraction, the speech recognition result yields multiple feature vectors that can be represented by a second vector set W2, i.e., W2 = {W2_vec1, W2_vec2, W2_vec3, ..., W2_vecn}. It is worth noting that both vector sets W1 and W2 contain the exact same number of feature vectors.
[0068] S203. Determine the augmented dataset from the sentences corresponding to the speech recognition results and / or filtering results based on multiple feature vectors, multiple weight values and preset similarity thresholds, and add the augmented dataset to the augmented database.
[0069] In this step, the weight values correspond to the feature vectors; for example, each feature vector corresponds to a different weight value. It should be noted that the weight values can be flexibly adjusted according to the needs of the actual application scenario, or a preset adjustment model, such as a neural network model, can be used to adjust the weight values.
[0070] Specifically, the similarity, such as cosine similarity, of each corresponding feature vector in the first vector set W1 and the second vector set W2 is calculated. Then, the weight value corresponding to each feature vector is multiplied by the similarity, and the resulting products are summed to calculate the weighted average of the similarities, which is used as the target similarity. Then, it is determined whether each target similarity is greater than a preset similarity threshold. If so, the sentences corresponding to the speech recognition results and / or filtering results are determined to be the automatically selected labeled data. These labeled data are then packaged into an augmented dataset.
[0071] In this step, the expanded database is used to perform quality inspection on the voice service data. It should be noted that the expanded database may only include the sentences corresponding to the aforementioned speech recognition results and / or filtering results (i.e., the expanded dataset), or it may include the expanded dataset and all target words in the target vocabulary. Therefore, this step can be implemented in several ways:
[0072] (1) Add the augmented dataset to a separately configured augmented database, and use the augmented database and the target lexicon for speech quality supervision or speech quality inspection processing.
[0073] (2) Use the target lexicon in S202 as an expanded database and add the expanded dataset directly to the target lexicon.
[0074] Then, using the trained voice quality inspection model, the extended database is called to conduct quality inspection and supervision on the real-time voice communication audio data between business personnel and customers. Alternatively, the voice dialogue audio between business personnel and customers can be stored in a temporary database first. At each interval, a portion of the dialogue audio is extracted from the temporary database according to preset rules and input into the voice quality inspection model. Based on the extended database, quality inspection and supervision are carried out, and the supervision results are fed back to business management personnel in the form of reports or early warning prompts.
[0075] This application provides a data augmentation processing method that extracts features from speech recognition results and their filtering results in multiple dimensions through various expression methods, thereby improving the accuracy of feature information extraction for target words.
[0076] Figure 3 A flowchart illustrating another data augmentation processing method provided for the implementation of this application. (e.g.) Figure 3 As shown, the specific steps of this data augmentation processing method include:
[0077] S301. Obtain the audio file and perform speech recognition on the audio file to obtain the speech recognition result.
[0078] S302. Obtain the original lexicon and preset scene training data.
[0079] In this step, the original lexicon includes: the target lexicon to be updated, the initial version of the target lexicon, or a version of the target lexicon from a previous update, etc. The target lexicon contains a large number of words and / or sentences, which are referred to as target words in this application. Target words are textual expressions with a preset text or speech recognition target as the semantic understanding target.
[0080] The preset scenario training data is used to expand the original lexicon. That is, the data is collected in one or more specific business scenarios with preset targets. Optionally, it can also be screened manually or screened and verified by the training model to finally obtain the scenario training data.
[0081] S303. Based on each target word in the original lexicon, train word vectors on the scene training data to determine multiple word vectors.
[0082] In this step, using Word2vec technology and pre-set training models, such as neural network models and self-learning models, based on a large amount of scene training data and combined with various target words in the original lexicon, a portion of the sentences in the scene training data is converted into multiple word vectors.
[0083] S304. Based on multiple word vectors, traverse the original vocabulary and determine the similarity between each word vector and the target word.
[0084] In this step, a preset similarity algorithm, such as the formula for calculating cosine similarity, is used to traverse the original vocabulary based on all word vectors obtained in S303, and calculate the similarity between each word vector and each target word.
[0085] S305. Add the scene training data corresponding to the top N similarity scores to the original vocabulary to obtain the target vocabulary.
[0086] In this step, the similarity of each word vector in S304 is sorted from largest to smallest. Then, the text corresponding to the top N word vectors, i.e. the text in the scene training data, is selected and added to the original vocabulary.
[0087] Optionally, after being added to the original lexicon, since there may be some errors in the similar words calculated based on word vectors, manual judgment and verification are required to confirm that the words added to the original lexicon are semantically similar to the target words before they are finally retained in the original lexicon to obtain the expanded target lexicon.
[0088] S306. Use the target vocabulary to filter the speech recognition results and determine the filtering results.
[0089] In this embodiment, the words with the same emotional polarity as the target words are identified as the filtering results, that is, the filtering results contain at least one target word from the target word library.
[0090] By filtering, text statements that do not contain the target words are initially removed, reducing interference with subsequent filtering and dataset expansion, and improving generalization ability.
[0091] S307. Perform sentiment-based filtering on the screening results and determine the filtering results.
[0092] In this step, the sentiment polarity of the filtered results is the same as that of the target words. The sentiment polarity includes: negative sentiment, neutral sentiment, and positive sentiment.
[0093] In this embodiment, the filtering results are determined by selecting those with the same sentiment polarity as the target word. Sentiment filtering further removes useless data without sentiment bias, further reducing interference with subsequent dataset expansion. After expanding the target word library in steps S302-S305, followed by two steps in S306 (coarse recall filtering) and S307 (fine sentiment recognition filtering), the generalization ability when expanding the labeled data for the target word is improved.
[0094] Specifically, in the quality inspection of voice services, touching sensitive words, i.e., target words, generally carries a certain emotional bias. For example, swear words will mostly have a negative emotional bias. Another example is the sensitive word "complaint," which might be expressed in two ways: "Call me again and I'll complain about you" and "Hello sir, our complaint hotline is 12315." The former has business value, meaning its emotional polarity is negative, and it's what quality inspection aims to detect. The latter, however, has little business value, as its emotional polarity is neutral, representing only a procedural or polite exchange. By comparing the two sentences, the former clearly has a negative emotional bias, while the latter is more neutral. Based on this, when expanding text, using a sentiment model for text filtering can improve the accuracy of the augmented data and reduce the computational load of subsequent steps.
[0095] In this step, traditional deep learning models such as TextCNN and LSTM can be used for sentiment filtering, or pre-trained models such as the BERT model can be used for sentiment recognition.
[0096] Figure 4 This is a schematic diagram illustrating the process of identifying and filtering sentiment tendencies, as provided in an embodiment of this application. Figure 4 As shown, the labeled text 41 is input into the preset model 42 for sentiment filtering training. After training, the sentiment model 43 is obtained. The sentiment model 43 performs sentiment recognition on the input data to be recognized 44 and finally outputs the sentiment polarity 45 of the text to be recognized 44.
[0097] S308. Multiple feature vectors are determined by extracting features from the filtering results and speech recognition results in multiple dimensions through various expression methods.
[0098] In this embodiment, the expression methods include: a first expression method, a second expression method, and a third expression method. The first expression method is used to express contextual semantics that are the same as or similar to the target word in the target lexicon; that is, the first expression method is an expression based on the first dimension of the target word's influence on contextual semantics. The second expression method is used to express the hyperordinate context corresponding to the target word in the target lexicon; that is, the second expression method is an expression based on the second dimension of contextual hyperordinate understanding. The third expression method is used to express the word density of the target word appearing in the filtering results and speech recognition results; that is, the third expression method is an expression based on the third dimension of word density. Correspondingly, the feature vector includes: a first feature vector, a second feature vector, and a third feature vector.
[0099] It should be noted that the first feature vector, the second feature vector, and the third feature vector each contain two parts: one part is the first vector corresponding to the filtering result, and the other part is the second vector corresponding to the speech recognition result.
[0100] The specific implementation methods for extracting the first dimension features from the filtering results and speech recognition results using the first expression method include:
[0101] Using the first expression method, add a sentence beginning marker to each sentence in the filtering results and speech recognition results to obtain at least one first sentence;
[0102] Perform multi-layer feature extraction on at least one first sentence, and determine the first latent vector corresponding to the sentence beginning identifier of each first sentence in the extraction results;
[0103] The target words in each first statement are masked or removed to determine at least one second statement.
[0104] Multi-level feature extraction is performed on each second statement, and the second latent vector corresponding to the sentence beginning identifier of each second statement is determined from the extraction results;
[0105] The first feature vector is determined based on the first latent vector and the second latent vector, and the feature vector includes the first feature vector.
[0106] Specifically, the text containing sensitive words (target words) in each sentence of the filtering results is processed by a pre-trained model, such as the BERT model, to extract features. The BERT model includes multiple feature extraction layers. In the last layer, the hidden states of each token corresponding to each text in the sentence are obtained. The vector representation of the target word is the feature vector, which is the average of the vectors corresponding to each token, denoted as W1_vec1.
[0107] Figure 5 This is a schematic diagram illustrating the first feature vector generation method provided in this application embodiment. For example... Figure 5 As shown, assuming the sensitive word or target word in "If you don't pay back the money, I'll file a complaint against you" is "complaint", then the vector corresponding to "complaint" is the sum of the latent vectors output by the two tokens "complaint" and "complaint" in the last layer of the BERT model.
[0108] Similarly, for the original data corresponding to the speech recognition result and the filtering result, the first expression method described above is also used to perform multi-layer feature extraction. The resulting feature vector is represented as W2_vec1, which will not be elaborated here.
[0109] The specific implementation methods for extracting second-dimensional features from the filtering results and speech recognition results using the second expression method include:
[0110] Using the second expression method, multi-layer feature extraction is performed on the filtering results and speech recognition results respectively to determine the third latent vector of each tag corresponding to the target word included in the filtering results and speech recognition results;
[0111] The mean of each third latent vector is determined as the second eigenvector, and the eigenvectors include the second eigenvector.
[0112] Specifically, for a complete sentence containing sensitive words (i.e., target words) in the filtering results, special characters [cls] and [sep] are added before and after the sentence is input into the model. The processed sentence is then sent to the pre-trained model, where feature extraction is performed. The latent vector h1_cls corresponding to [cls] in the last layer of the model can be regarded as the semantic representation of the entire sentence, which contains the semantic information of the entire sentence.
[0113] Figure 6This is a schematic diagram illustrating a second feature vector generation method provided in an embodiment of this application. For example... Figure 6 As shown, to measure the role of sensitive words, i.e. target words, in a sentence, we can perform a masking operation on the sensitive word "complain" in the sentence "If you don't pay back the money, I'll complain about you," or remove the sensitive word, i.e. target word, to obtain the sentence "If you don't pay back the money, I'll complain about you." Then, we input it into the BERT model and still take the latent vector h2_cls corresponding to the last layer [cls]. This vector can represent the semantic representation of the sentence after masking the sensitive word, i.e. target word. The difference between the two represents the role of the sensitive word in the sentence, which can be regarded as the vector representation of the sensitive word, denoted as W1_vec2 = h1_cls - h2_cls.
[0114] Similarly, for the original data corresponding to the speech recognition result, i.e. the filtering result, the second expression method described above is also used to perform multi-layer feature extraction. The resulting feature vector is represented as W2_vec2, which will not be elaborated here.
[0115] Specific implementation methods for extracting third-dimensional features from both the filtering results and speech recognition results using a third expression method include:
[0116] Using the target vocabulary, word vectors are trained on both the filtering results and the speech recognition results;
[0117] The word vectors in the training results are determined as the third feature vectors, and the feature vectors include the third feature vector.
[0118] Specifically, similar to S303, Word2vec technology is used to utilize a pre-trained model, such as a neural network model or a self-learning model, to convert the sentences in the filtered results into multiple word vectors based on the filtering results and the target words in the target lexicon. These vector groups are then combined into a feature vector in a pre-defined manner, denoted as W1_vec3.
[0119] Similarly, for the original data corresponding to the speech recognition result and the filtering result, the third expression method mentioned above is also used to perform multi-layer feature extraction. The resulting feature vector is represented as W2_vec3, which will not be elaborated here.
[0120] S309. Determine the augmented dataset from the sentences corresponding to the speech recognition results based on multiple feature vectors, multiple weight values and preset similarity thresholds, and add the augmented dataset to the augmented database.
[0121] In this step, multiple similarities are determined, and each similarity corresponds one-to-one with a multiple feature vector. The specific implementation of determining the similarity includes: calculating the similarity between the first vector and the second vector included in the feature vector.
[0122] The target similarity is determined based on multiple similarity scores and multiple weight values. For example, the target similarity is equal to the sum of the products of multiple similarity scores and their corresponding weight values.
[0123] If the target similarity is greater than the preset similarity threshold, the sentences corresponding to the speech recognition results and / or filtering results will be added to the augmented dataset.
[0124] Calculate the similarity between the first and second vectors in each feature vector; determine the target similarity by summing the products of each similarity and its corresponding weight value; if the target similarity is greater than a preset similarity threshold, add the sentence corresponding to the speech recognition result to the augmented dataset.
[0125] In this embodiment, the first vector in the first feature vector is represented by W1_vec1, and the second vector in the first feature vector is represented by W2_vec1. Then, the similarity score1 corresponding to the first feature vector can be calculated according to formula (1), which is shown below:
[0126]
[0127] The first vector in the second feature vector is represented by W1_vec2, and the second vector in the second feature vector is represented by W2_vec2. Then the similarity score2 corresponding to the second feature vector can be calculated according to formula (2), which is shown below:
[0128]
[0129] The first vector in the third feature vector is represented by W1_vec3, and the second vector in the third feature vector is represented by W2_vec3. Then, the similarity score3 corresponding to the third feature vector can be calculated according to formula (3), which is shown below:
[0130]
[0131] Assuming the first weight of the first feature vector is α1, the second weight of the second feature vector is α2, and the third weight of the third feature vector is α3, then the target similarity score... sum It can be calculated according to formula (4), which is shown below:
[0132] score sum =α1*score1+α2*score2+α3*score3 (4)
[0133] Where α1+α2+α3=1, and the values of α1, α2, and α3 can be flexibly configured according to the actual situation.
[0134] Next, the target similarity is compared with a preset similarity threshold. If it is greater than the preset similarity threshold, it is considered to meet the requirements and is added to the augmented dataset.
[0135] In one possible design, the preset similarity thresholds include an optimal threshold θ1 and a minimum hit threshold θ2, and each value of the preset similarity threshold can be flexibly configured according to specific usage scenarios. The similarity scores of each target are compared with the optimal threshold θ1 and the minimum hit threshold θ2 respectively. If the target similarity score is higher than the optimal threshold θ1, the sensitive words (target words) in the sentence are considered very similar to the sensitive words in the original seed corpus data, and the sentence can be directly added to the expanded database. If the similarity score is lower than the minimum hit threshold θ2, the similarity value is considered particularly low, and the sentence can be discarded. If the similarity value is greater than the minimum hit threshold θ2 but less than the optimal threshold θ1, the sensitive words in the sentence are considered to have some similarity to the sensitive words in the seed corpus, but not enough to directly determine whether it can be directly added to the expanded database; manual verification is required.
[0136] Figure 7 This is a schematic diagram of a data augmentation processing device provided in an embodiment of this application. The data augmentation processing device 700 can be implemented through software, hardware, or a combination of both.
[0137] like Figure 7 As shown, the data augmentation processing apparatus 700 includes:
[0138] Module 701 is used to acquire audio recording files;
[0139] Processing module 702 is used for:
[0140] Perform speech recognition on the audio file to obtain the speech recognition results;
[0141] The speech recognition results are filtered using a target vocabulary to obtain the filtered results. Multiple feature vectors are determined by extracting features from the filtered results and the speech recognition results in multiple dimensions through various expression methods. Each feature vector contains feature information in at least one dimension.
[0142] The augmented dataset is determined from the sentences corresponding to the speech recognition results and / or filtering results based on multiple feature vectors, multiple weight values, and preset similarity thresholds, and then added to the expanded database. The weight values correspond to the feature vectors. The expanded database is used for quality inspection of speech service data.
[0143] In one possible design, the filtering results include at least one word or phrase that matches the target words in the target thesaurus.
[0144] In one possible design, the words or phrases included in the filtering results have the same sentiment polarity as the target words included in the target lexicon.
[0145] In one possible design, the various expression methods include a first expression method, which is used to express contextual semantics that are the same as or similar to target words in the target lexicon; the processing module 702 is used for:
[0146] The specific implementation methods for extracting the first dimension features from the filtering results and speech recognition results using the first expression method include:
[0147] Using the first expression method, add a sentence beginning marker to each sentence in the filtering results and speech recognition results to obtain at least one first sentence;
[0148] Perform multi-layer feature extraction on at least one first sentence, and determine the first latent vector corresponding to the sentence beginning identifier of each first sentence in the extraction results;
[0149] The target words in each first statement are masked or removed to determine at least one second statement.
[0150] Multi-level feature extraction is performed on each second statement, and the second latent vector corresponding to the sentence beginning identifier of each second statement is determined from the extraction results;
[0151] The first feature vector is determined based on the first latent vector and the second latent vector, and the feature vector includes the first feature vector.
[0152] In one possible design, multiple expression methods include: a second expression method, used to express the hyperordinate context corresponding to the target word in the target lexicon; and processing module 702, used for:
[0153] The specific implementation methods for extracting second-dimensional features from the filtering results and speech recognition results using the second expression method include:
[0154] Using the second expression method, multi-layer feature extraction is performed on the filtering results and speech recognition results respectively to determine the third latent vector of each tag corresponding to the target word included in the filtering results and speech recognition results;
[0155] The mean of each third latent vector is determined as the second eigenvector, and the eigenvectors include the second eigenvector.
[0156] In one possible design, multiple expression methods include a third expression method, which is used to express the filtering results and the word density of the target word appearing in the speech recognition results; the processing module 702 is used for:
[0157] Specific implementation methods for extracting third-dimensional features from both the filtering results and speech recognition results using a third expression method include:
[0158] Using the target vocabulary, word vectors are trained on both the filtering results and the speech recognition results;
[0159] The word vectors in the training results are determined as the third feature vectors, and the feature vectors include the third feature vector.
[0160] In one possible design, the processing module 702 is used for:
[0161] Multiple similarities are determined, and each similarity corresponds one-to-one with a multiple feature vector; the specific implementation of determining the similarity includes: calculating the similarity between the first vector and the second vector included in the feature vector;
[0162] Target similarity is determined based on multiple similarity scores and multiple weight values;
[0163] If the target similarity is greater than the preset similarity threshold, the sentences corresponding to the speech recognition results and / or filtering results will be added to the augmented dataset.
[0164] In one possible design, the processing module 702 is used for:
[0165] The speech recognition results are filtered using a target lexicon to determine the filtering results. The filtering results contain at least one target word from the target lexicon. The filtering results include the filtering results.
[0166] In one possible design, the processing module 702 is also used for:
[0167] The screening results are filtered based on sentiment tendency to determine the filtered results. The sentiment polarity of the filtered results is the same as that of the target words.
[0168] In one possible design, the expression also includes: a first expression based on the first dimension of the target word's influence on contextual semantics, processing module 702, used for:
[0169] Using the first expression method, add sentence beginning and sentence ending identifiers to each sentence in the filtering results and speech recognition results to identify each first sentence;
[0170] Multi-level feature extraction is performed on each first sentence, and the first latent vector corresponding to the sentence beginning identifier of each first sentence is determined from the extraction results;
[0171] The target words in each first statement are masked or removed to determine each second statement;
[0172] Multi-level feature extraction is performed on each second statement, and the second latent vector corresponding to the sentence beginning identifier of each second statement is determined from the extraction results;
[0173] The first feature vector is determined based on the first latent vector and the second latent vector, and the feature vector includes the first feature vector.
[0174] In one possible design, the expression method also includes: a second expression method based on the second dimension of contextualized understanding; the processing module 702 is further used for:
[0175] Using the second expression method, multi-layer feature extraction is performed on the filtering results and speech recognition results respectively to determine the third latent vector of each tag corresponding to the target word;
[0176] The mean of each third latent vector is determined as the second eigenvector, and the eigenvectors include the second eigenvector.
[0177] In one possible design, the expression method also includes: a third expression method based on the third dimension of word density; the processing module 702 is further used for:
[0178] Using the target vocabulary, word vectors are trained on both the filtering results and the speech recognition results;
[0179] The word vectors in the training results are determined as the third feature vectors, and the feature vectors include the third feature vector.
[0180] In one possible design, the feature vector includes: a first vector corresponding to the filtering result and a second vector corresponding to the speech recognition result. The processing module 702 is used for:
[0181] Calculate the similarity between the first and second vectors in each feature vector;
[0182] The sum of the products of each similarity score and its corresponding weight value is determined as the target similarity score.
[0183] If the target similarity is greater than the preset similarity threshold, the sentence corresponding to the speech recognition result will be added to the augmented dataset.
[0184] In one possible design, the acquisition module 701 is also used to acquire the original lexicon and preset scene training data;
[0185] Processing module 702 is also used for:
[0186] Based on each target word in the original lexicon, word vectors are trained on the scene training data to determine multiple word vectors;
[0187] Based on multiple word vectors, the original vocabulary is traversed to determine the similarity between each word vector and the target word;
[0188] The training data of the scenes corresponding to the top N similarities are added to the original vocabulary to obtain the target vocabulary.
[0189] It is worth noting that, Figure 7 The apparatus provided in the illustrated embodiments can execute the methods provided in any of the above method embodiments. Their specific implementation principles, technical features, explanations of technical terms, and technical effects are similar and will not be repeated here.
[0190] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 8 As shown, the electronic device 800 may include at least one processor 801 and a memory 802. Figure 8 The example shown is an electronic device using a processor.
[0191] The memory 802 is used to store programs. Specifically, the program may include program code, which includes computer operation instructions.
[0192] The memory 802 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.
[0193] The processor 801 is used to execute computer execution instructions stored in the memory 802 to implement the methods described in the above embodiments.
[0194] The processor 801 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application.
[0195] Optionally, the memory 802 can be either standalone or integrated with the processor 801. When the memory 802 is a device independent of the processor 801, the electronic device 800 may further include:
[0196] Bus 803 is used to connect the processor 801 and the memory 802. The bus can be an industry standard architecture (ISA) bus, a peripheral component (PCI) bus, or an extended industry standard architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc., but this does not mean there is only one bus or one type of bus.
[0197] Optionally, in a specific implementation, if the memory 802 and the processor 801 are integrated on a single chip, the memory 802 and the processor 801 can communicate through an internal interface.
[0198] This application also provides a computer-readable storage medium, which may include various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a disk, or an optical disk. Specifically, the computer-readable storage medium stores program instructions, which are used in the methods described in the above-mentioned method embodiments.
[0199] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the methods described in the above-described method embodiments.
[0200] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the claims.
[0201] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
Claims
1. A data augmentation processing method, characterized in that, include: Acquire an audio recording file and perform speech recognition on the audio recording file to obtain the speech recognition result; The speech recognition results are filtered using a target vocabulary to obtain a filtered result; multiple feature extractions are performed on the filtered result and the speech recognition result using various expression methods to determine multiple feature vectors, each feature vector including a first vector corresponding to the filtered result and a second vector corresponding to the speech recognition result; Multiple similarities are determined, each of which is determined based on the first vector and the second vector included in the corresponding feature vector; The target similarity is determined based on the multiple similarities and multiple weight values. When the target similarity is greater than a preset similarity threshold, the sentences corresponding to the speech recognition results and / or the filtering results are added to the augmented dataset, and the augmented dataset is added to the expanded database. The weight values correspond to the feature vectors, and the expanded database is used for quality inspection of the speech service data.
2. The data augmentation processing method according to claim 1, characterized in that, The filtering result includes at least one word or phrase, and the word or phrase included in the filtering result matches the target words included in the target thesaurus.
3. The data augmentation processing method according to claim 2, characterized in that, The words or phrases included in the filtering results have the same sentiment polarity as the target words included in the target lexicon.
4. The data augmentation processing method according to claim 1, characterized in that, The multiple expression methods include a first expression method, which is used to express contextual semantics that are the same as or similar to target words in the target lexicon; The specific implementation methods for extracting features of the first dimension from the filtering result and the speech recognition result using the first expression method include: Using the first expression method, add a sentence beginning identifier to each sentence in the filtering result and the speech recognition result to obtain at least one first sentence; Multi-layer feature extraction is performed on at least one of the first statements, and a first latent vector corresponding to the sentence beginning identifier of each of the first statements is determined from the extraction results; The target words in each of the first statements are masked or removed to determine at least one second statement; Multi-level feature extraction is performed on each of the second statements, and the second latent vector corresponding to the sentence beginning identifier of each of the second statements is determined from the extraction results; A first feature vector is determined based on the first latent vector and the second latent vector, wherein the feature vector includes the first feature vector.
5. The data augmentation processing method according to claim 1, characterized in that, The multiple expression methods include: a second expression method, which is used to express the hyperordinate context corresponding to the target word in the target lexicon; The specific implementation methods for extracting second-dimensional features from the filtering results and the speech recognition results using the second expression method include: Using the second expression method, multi-layer feature extraction is performed on the filtering result and the speech recognition result respectively to determine the third latent vector of each tag corresponding to the target word included in the filtering result and the speech recognition result; The mean of each of the third latent vectors is determined as a second feature vector, the feature vector including the second feature vector.
6. The data augmentation processing method according to claim 1, characterized in that, The multiple expression methods include a third expression method, which is used to express the filtering result and the word density of the target word appearing in the speech recognition result; The specific implementation methods for extracting third-dimensional features from the filtering results and the speech recognition results using a third expression method include: Using the target vocabulary, word vectors are trained on the filtering results and the speech recognition results, respectively; The word vectors in the training results are determined as the third feature vector, and the feature vector includes the third feature vector.
7. A data augmentation processing device, characterized in that, include: The acquisition module is used to acquire audio files; Processing module, used for: The audio file is subjected to speech recognition to obtain the speech recognition result; The speech recognition results are filtered using a target vocabulary to obtain a filtered result; multiple feature extractions are performed on the filtered result and the speech recognition result using various expression methods to determine multiple feature vectors, each feature vector including a first vector corresponding to the filtered result and a second vector corresponding to the speech recognition result; Multiple similarities are determined, each based on the first vector and the second vector included in the corresponding feature vector; a target similarity is determined based on the multiple similarities and multiple weight values; when the target similarity is greater than a preset similarity threshold, the sentences corresponding to the speech recognition results and / or the filtering results are added to the augmented dataset, and the augmented dataset is added to the expanded database. The weight values correspond to the feature vectors, and the expanded database is used for quality inspection of the speech service data.
8. An electronic device, characterized in that, include: processor; as well as, Memory for storing the computer program of the processor; The processor is configured to perform the data augmentation processing method according to any one of claims 1 to 6 by executing the computer program.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the data augmentation processing method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Extended corpus generation method and device in target field and electronic equipment
CN112541076A
Corpus processing method, related device and equipment
CN113821593A