Synthetic data driven paralinguistic labeling method, apparatus, device and storage medium
By using a dual-supervised model training mode that integrates speech and text features, the problem of noise and accent interference in different scenarios for paralanguage recognition models is solved, and accurate recognition and annotation of paralanguage is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING HAITIAN RUISHENG SCI TECH CO LTD
- Filing Date
- 2025-11-14
- Publication Date
- 2026-04-10
AI Technical Summary
Existing paralinguistic recognition models are inaccurate in recognizing paralinguistic features in speech under the influence of noise and user accents in different scenarios.
By training a speech recognition model using both speech and text as input data, and employing a multimodal fusion mechanism to integrate speech and text features, a dual-supervised model training mode is achieved, thereby improving the accuracy of paralinguistic annotation.
It achieves accurate identification and annotation of sub-language in different scenarios, improving the robustness and accuracy of the model.
Smart Images

Figure CN121122249B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of speech recognition, and particularly relates to a synthetic data driven paralinguistic labeling method and device, equipment and a storage medium. BACKGROUND
[0002] With the acoustic model providing more and more diversified functions, the acoustic model supports recognizing paralinguistic information in speech.
[0003] However, with the emergence of paralinguistic recognition models, when using the paralinguistic recognition model to recognize paralinguistic information in speech, due to different noises, languages, and user accents in speech in different scenarios, the paralinguistic recognition model does not accurately recognize paralinguistic information in speech. SUMMARY
[0004] To overcome the problems in the related art, the present disclosure provides a synthetic data driven paralinguistic labeling method, device, equipment and storage medium.
[0005] According to a first aspect of an embodiment of the present disclosure, a paralinguistic labeling method is provided, comprising: obtaining language information of paralinguistic information to be labeled; identifying the language information through a preset paralinguistic labeling model to obtain target text, the target text containing paralinguistic labeling, the paralinguistic labeling model being obtained by adjusting a speech recognition model architecture based on output text output by the speech recognition model architecture and label data included in a sample pair, the output text being obtained by fusing a speech feature vector and a semantic feature vector and converting the obtained fused vector into text, the speech feature vector and the semantic feature vector being obtained by respectively extracting features of speech and text in the sample pair and converting the obtained speech features and semantic features into vectors.
[0006] In an implementation, the speech recognition model architecture is adjusted to obtain the paralinguistic labeling model in the following manner: extracting speech features of speech in a sample pair and converting the speech features into a speech feature vector; obtaining text corresponding to the speech in the sample pair and determining prompt text corresponding to the text, the prompt text not including paralinguistic information; extracting semantic features of the prompt text and converting the semantic features into a semantic feature vector; cross-fusing the speech feature vector and the semantic feature vector and converting the obtained fused vector into output text; and adjusting parameters of the speech recognition model architecture based on the output text and label data included in the sample pair to obtain the paralinguistic labeling model.
[0007] In an implementation, the speech recognition model architecture is adjusted to obtain the paralinguistic annotation model in the following manner: speech features of speech or semantic features of text in a sample pair are extracted; the speech features are converted into speech feature vectors or the semantic features are converted into semantic feature vectors; output text corresponding to the speech feature vectors or the semantic feature vectors is output; and parameters of the speech recognition model architecture are adjusted based on the output text and label data included in the sample pair to obtain the paralinguistic annotation model.
[0008] In an implementation, the sample pair is determined in at least one of the following manners: text is generated based on a model, speech is generated, and the text and the speech are combined into a sample pair; paralinguistic speech streams are embedded in speech to obtain synthesized speech, and paralinguistic text is embedded in text to obtain synthesized text, and the synthesized speech and the synthesized text are combined into a sample pair; paralinguistic speech streams in speech are annotated to obtain annotated speech, and the annotated speech is converted into annotated text, and the annotated speech and the annotated text are combined into a sample pair.
[0009] In an implementation, the language information includes text and / or speech.
[0010] According to a second aspect of the embodiments of the present disclosure, a paralinguistic annotation device driven by synthesized data is provided, and the device comprises:
[0011] An acquisition unit is configured to acquire language information of a paralinguistic language to be annotated.
[0012] A processing unit is configured to recognize the language information by using a preset paralinguistic annotation model to obtain target text, the target text containing paralinguistic annotation, the paralinguistic annotation model being obtained by adjusting a speech recognition model architecture based on output text output by the speech recognition model architecture and label data included in a sample pair, the output text being obtained by fusing speech feature vectors and semantic feature vectors and converting the obtained fused vectors into text, the speech feature vectors and the semantic feature vectors being obtained by extracting features of speech and text included in the sample pair respectively and converting the obtained speech features and semantic features into vectors.
[0013] According to some embodiments of the present disclosure, the auxiliary language labeling model is obtained by adjusting the speech recognition model architecture in the following manner: extracting speech features of the speech in the sample pair and converting the speech features into a speech feature vector; obtaining text corresponding to the speech in the sample pair and determining prompt text corresponding to the text, the prompt text not including auxiliary language; extracting semantic features of the prompt text and converting the semantic features into a semantic feature vector; cross-fusing the speech feature vector and the semantic feature vector and converting the obtained fused vector into output text; and adjusting parameters of the speech recognition model architecture based on the output text and label data included in the sample pair to obtain the auxiliary language labeling model.
[0014] According to some embodiments of the present disclosure, the auxiliary language labeling model is obtained by adjusting the speech recognition model architecture in the following manner: extracting speech features of the speech in the sample pair or semantic features of the text; converting the speech features into a speech feature vector or the semantic features into a semantic feature vector; outputting output text corresponding to the speech feature vector or the semantic feature vector; and adjusting parameters of the speech recognition model architecture based on the output text and label data included in the sample pair to obtain the auxiliary language labeling model.
[0015] According to some embodiments of the present disclosure, the sample pair is determined in at least one of the following manners: generating text based on a model and generating speech, and combining the text and the speech into a sample pair; embedding an auxiliary language speech stream in the speech to obtain a synthesized speech and embedding auxiliary language text in the text to obtain synthesized text, and combining the synthesized speech and the synthesized text into a sample pair; labeling an auxiliary language speech stream in the speech to obtain labeled speech and converting the labeled speech into labeled text, and combining the labeled speech and the labeled text into a sample pair.
[0016] According to some embodiments of the present disclosure, the language information includes text and / or speech.
[0017] According to a third aspect of the embodiments of the present disclosure, an electronic device is provided, including: a processor; a memory for storing processor-executable instructions; and wherein the processor is configured to perform the synthetic data-driven auxiliary language labeling method in any one of the first aspect or the implementation manners thereof.
[0018] According to a fourth aspect of the embodiments of the present disclosure, a storage medium is provided, and the storage medium stores instructions, when the instructions in the storage medium are executed by a processor, the processor can perform the synthetic data-driven auxiliary language labeling method in the first aspect or any one of the implementation manners thereof.
[0019] The technical scheme provided by the embodiment of the present disclosure can include the following beneficial effects: the voice and text in the sample pair are simultaneously used as input data to train an existing speech recognition model architecture, and a new dual-supervision model training mode for paralinguistic labeling is implemented. Based on a multi-modal fusion mechanism, the features corresponding to the voice and text are fused, the multi-modal data of the voice and text can be recognized, and the target text containing paralanguage can be accurately obtained. It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF DRAWINGS
[0020] The accompanying drawings, which are incorporated into and form part of the specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0021] Figure 1 is a flowchart of a paralinguistic labeling method driven by synthetic data according to an exemplary embodiment of the present disclosure.
[0022] Figure 2 is a flowchart of a paralinguistic labeling model training method according to an exemplary embodiment of the present disclosure.
[0023] Figure 3 is a schematic block diagram of a Whisper model architecture according to an exemplary embodiment of the present disclosure.
[0024] Figure 4 is a schematic diagram of a decoder output target text sequence according to an exemplary embodiment of the present disclosure.
[0025] Figure 5 is a flowchart of another paralinguistic labeling model training method according to an exemplary embodiment of the present disclosure.
[0026] Figure 6 is a block diagram of a paralinguistic labeling device driven by synthetic data according to an exemplary embodiment of the present disclosure.
[0027] Figure 7 is a block diagram of a device for paralinguistic labeling driven by synthetic data according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION
[0028] Some embodiments of the present disclosure will be described in detail herein, with examples shown in the drawings. The following description is made with reference to the accompanying drawings, in which like reference numerals represent like elements or similar elements unless otherwise stated. Various changes, modifications, and equivalents of the methods, devices, and / or systems described herein will become apparent to those skilled in the art after a study of the following description. For instance, the order in which operations are described is merely an example, and is not intended to be limiting, as the operations can be performed in any order, except where otherwise required by the particular order of operations. Additionally, for the sake of brevity and clarity, descriptions of features known to one skilled in the art can be omitted.
[0029] The implementations described in some embodiments of the present disclosure below do not represent all implementations consistent with the present disclosure. Instead, they are merely examples of apparatuses and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0030] For ease of understanding, some technical terms appearing in the embodiments of the present disclosure will be illustratively explained below.
[0031] Open-source speech recognition model (Whisper), including a Transformer architecture.
[0032] Deep learning architecture (Transformer): a mechanism that uses self-attention as its core, widely used in natural language processing (NLP) field. Its core architecture consists of an encoder and a decoder, which realizes information interaction through attention mechanism.
[0033] Mel-frequency cepstral coefficients (Mel-Frequency Cepstral Coefficients, MFCC): a speech feature extraction method designed based on human auditory characteristics, which converts speech signals into low-dimensional feature representation by simulating human ear's nonlinear perception of frequency, widely used in speech recognition, speaker recognition and speech synthesis fields. Its core steps include pre-emphasis, framing and windowing, spectral analysis, Mel filter bank mapping, logarithmic transformation and discrete cosine transformation, and finally the feature coefficients are extracted.
[0034] Feature extraction method in speech signal processing (FilterBank, FBank): a commonly used feature extraction method that simulates human auditory characteristics to improve speech recognition performance.
[0035] Pre-trained language models (Bidirectional Encoder Representations from Transformers, BERT) are based on the Transformer architecture, which generates dynamic word embeddings through bidirectional context understanding, significantly improving the performance of natural language processing (NLP) tasks.
[0036] Sinusoidal Positional Encoding (SPE) is a technique used in Transformer models to introduce position information, generating unique encodings for each position in the sequence using sine and cosine functions, allowing the model to capture word order relationships.
[0037] Time-to-frequency domain conversion methods (Log-Mel spectrogram): a common method for converting sound signals from time domain to frequency domain, widely used in audio processing and speech recognition.
[0038] Learned Positional Encoding (LPE) is a method of automatically generating position information through machine learning parameters, allowing the model to understand the position relationships of elements in the sequence.
[0039] Next-Token Prediction (NTP) is a core technology in natural language processing, which refers to predicting the next Token (such as a word, phrase, or punctuation mark) after knowing the previous part of the text.
[0040] Multitask learning method (Tokens in Multitask Training Format) is a machine learning method that allows the model to simultaneously learn to perform multiple related tasks.
[0041] The paralinguistic labeling method provided by some embodiments of the present disclosure is applied to different dialogue scenarios to identify paralinguistic information in dialogue speech.
[0042] In related technologies, as the functions provided by the acoustic model become more diverse, the acoustic model supports the identification of paralinguistic information in speech. Paralinguistic information, for example, the patient's breathing sound in a medical dialogue scenario, the laughter of friends during conversation, and the animal's call during conversation on the plains, etc. However, with the emergence of paralinguistic recognition models, when using paralinguistic recognition models to identify paralinguistic information in speech, due to different noise and user accents in different scenarios (e.g., noisy environment, speech including dialect speech, and multiple phonetic characters in speech, etc.), the paralinguistic recognition model does not accurately identify paralinguistic information in speech.
[0043] Therefore, some embodiments of the present disclosure provide a paralinguistic labeling method. A pre-existing speech recognition model architecture is trained by taking the speech and text in a sample pair as input data, thereby implementing a new dual-supervised model training mode for paralinguistic labeling. The paralinguistic labeling model obtained based on pre-training can fuse the features corresponding to the speech and text by using a multi-modal fusion mechanism, thereby achieving recognition of multi-modal data of the speech and text and accurately obtaining target text containing paralanguage.
[0044] Figure 1 is a flowchart of a synthetic data-driven paralinguistic labeling method according to an exemplary embodiment, as shown in Figure 1 The method includes steps S11 to S12.
[0045] In step S11, language information of paralanguage to be labeled is obtained.
[0046] In the embodiments of the present disclosure, paralanguage includes sounds other than artificially created dialogue language, which are information carriers perceived by hearing. For example, breathing sounds, crying sounds, laughing sounds, coughing sounds, animal sounds, and the like.
[0047] In the embodiments of the present disclosure, the language information includes at least one of, but is not limited to, speech and text. The speech can contain paralanguage speech streams, and the text can not contain paralanguage corresponding labeled text.
[0048] In the embodiments of the present disclosure, the language information can be dialogue speech in different scenarios. For example, dialogue scenarios of multiple languages, dialects, medical care, and customer service.
[0049] In step S12, the language information is recognized by a preset paralanguage labeling model to obtain target text containing paralanguage labeling. The paralanguage labeling model is obtained by adjusting a speech recognition model architecture based on output text output by the speech recognition model architecture and label data included in a sample pair. The output text is obtained by fusing a speech feature vector and a semantic feature vector, and converting the obtained fusion vector into text. The speech feature vector and the semantic feature vector are obtained by respectively extracting features of the speech and the text included in the sample pair, and converting the obtained speech features and semantic features into vectors.
[0050] In the embodiments of the present disclosure, the speech recognition model architecture is trained based on a sample pair combining speech and text, thereby effectively enhancing the language detection and recognition capability of the paralanguage labeling model. The features of the speech and the text can be recognized and fused based on the obtained paralanguage labeling model. The multi-modal information fusion mechanism can more accurately recognize and label the paralanguage of the language information in the form of feature fusion of multiple types of information.
[0051] In the embodiments of the present disclosure, the language information includes text and / or voice. The paralinguistic annotation model can identify voice, text, and the combination of voice and text respectively, realize the input of different language information, and improve the practicability and universality of the paralinguistic annotation model. The multi-modal information can be fused, and the target text containing paralinguistic annotation can be accurately obtained.
[0052] In the embodiments of the present disclosure, the target text is embedded with paralinguistic text. For example, “Hello [laughter] today the weather is really good”, the embedded paralinguistic “laughter” can be directly obtained based on the target text, and no post-processing or additional annotation is required.
[0053] Figure 2 FIG. 1 is a flowchart of a paralinguistic annotation model training method according to an example embodiment. As shown in FIG. 1, the method comprises the following steps. Figure 2
[0054] In step S21, the speech features of the voice in the sample pair are extracted, and the speech features are converted into a speech feature vector.
[0055] In the embodiments of the present disclosure, the sample pair includes a voice and text pair. The sample pair also includes label data. The label data contains paralinguistic annotated text.
[0056] In the embodiments of the present disclosure, the speech features of the voice in the sample pair can be extracted by the encoder of the Transformer architecture of the Whisper model (an open source speech recognition model) through MFCC (a speech feature extraction method based on human auditory characteristics) or FBank (a feature extraction method in speech signal processing). By using the feature extraction method through the encoder of the speech recognition architecture, the speech features can be accurately extracted.
[0057] In step S22, the text corresponding to the voice in the sample pair is obtained, and the prompt text corresponding to the text is determined. The prompt text does not include paralinguistic information.
[0058] In the embodiments of the present disclosure, the prompt text can be directly obtained by recognizing the voice dialogue content, and can also be obtained by removing the paralinguistic text after recognizing the voice.
[0059] In the embodiments of the present disclosure, by determining the prompt text, another input data can be provided for the speech recognition model architecture. The prompt text and the voice can be simultaneously used as input data to realize double supervision model training.
[0060] In step S23, the semantic features of the prompt text are extracted, and the semantic features are converted into a semantic feature vector.
[0061] In the embodiments of the present disclosure, the semantic features of the prompt text can be extracted by the decoder of the speech recognition architecture. By taking the text prompt word as the initial condition of the decoder, context guidance is provided for the paralinguistic generation, improving the text semantic consistency and the generation accuracy of the final paralinguistic labeling model.
[0062] In step S24, the speech feature vector and the semantic feature vector are cross-fused, and the obtained fused vector is converted into output text.
[0063] In the embodiments of the present disclosure, the speech feature vector and the semantic feature vector can be fused in the manner of multi-modal cross-fusion operation of the decoder of the speech recognition architecture.
[0064] In step S25, the parameters of the speech recognition model architecture are adjusted based on the output text and the label data included in the sample pair, to obtain a paralinguistic labeling model.
[0065] In the embodiments of the present disclosure, the loss of the speech recognition model architecture can be calculated through the output text and the label data, the parameters of the speech recognition model architecture are adjusted based on the loss, and the paralinguistic labeling model is obtained. By training the paralinguistic labeling model through the dual-supervised model, the correlation between the speech and the text can be utilized to reduce the misrecognition and the missed recognition. The paralinguistic labeling model has stronger generalization ability and can accurately generate text containing paralinguistic labeling based on the features of the speech and the text.
[0066] In the embodiments of the present disclosure, test data can also be obtained based on the sample pair. The test data includes input data and label data. The input data includes at least one of the speech, the text, and the combination of the speech and the text. The label data includes text containing paralinguistic labeling. The paralinguistic labeling model is tested through the test data, and the parameters of the paralinguistic labeling model are continuously adjusted based on the test result, so as to optimize the paralinguistic labeling model and improve the recognition accuracy of the paralinguistic labeling model.
[0067] In the embodiments of the present disclosure, the paralinguistic labeling model can be integrated into a real-time voice assistant or a transcription system, etc., and the paralinguistic in the speech and / or text can be recognized and labeled through the real-time voice assistant or the transcription system, etc.
[0068] Figure 3 is a schematic block diagram of a Whisper model architecture according to an example embodiment, as shown in Figure 3 as shown, including.
[0069] Deep learning model based on encoder and decoder architecture (Sequence-to-Sequence Learning).
[0070] In the embodiments of the present disclosure, Sequence-to-Sequence Learning includes Transformer Encoder Blocks (encoder), Transformer Decoder Blocks (decoder) and (two one-dimensional convolutional layers and an activation function). In the figure represents cross multiplication.
[0071] In the embodiments of the present disclosure, the Transformer Encoder Blocks include multiple encoding layers, and each encoding layer includes an MLP (Multi-Layer Perceptron) and a Self attention (self-attention mechanism).
[0072] In the embodiments of the present disclosure, the Transformer Decoder Blocks include multiple decoding layers, and each decoding layer includes a Cross attention (cross-attention mechanism), an MLP and a Self attention.
[0073] In the embodiments of the present disclosure, Sequence-to-Sequence Learning also includes various deep learning algorithms, such as Sinusoidal Positional Encoding (a technique for introducing position information), Log-Mel spectrogram (a method of converting time domain to frequency domain), Tokens in Multitask Training Format (a multi-task learning method), Learned Positional Encoding (a position encoding method), and Next-Token Prediction (natural language processing).
[0074] In the embodiments of the present disclosure, for the encoder part in the figure, the speech is converted from time domain to frequency domain by the method of converting time domain to frequency domain, and the speech that the Whisper model can recognize is obtained. Then the converted speech is preliminarily extracted by two one-dimensional convolutional layers and an activation function. The preliminarily extracted speech features can be located in the paralinguistic features in the speech features by cross multiplication through the technique of introducing position information. Finally, the speech features of the speech are extracted and converted into a speech feature vector by the perception layer of the encoder and the self-attention mechanism.
[0075] In the embodiments of the present disclosure, for the decoder part in the figure, the decoder part can extract features from text through multi-task learning, and the text features are preliminarily extracted by cross multiplication through the position encoding method. Then the semantic features of the text are extracted and converted into a semantic feature vector by the perception layer of the decoder, the cross-attention mechanism and the self-attention mechanism.
[0076] In the embodiments of the present disclosure, the cross-attention mechanism of the deep learning model is used to fuse the semantic feature vector and the speech feature vector, and the obtained fusion vector is subjected to natural language processing to output text with paralinguistic labels.
[0077] In the embodiments of the present disclosure, when the dual-supervised model training is performed using the Transformer architecture of Whisper, the speech features of the speech in the sample pair can be extracted by the encoder, and the speech features are converted into a speech feature vector. The semantic features of the prompt text are extracted by the decoder, and the semantic features are converted into a semantic feature vector. The speech feature vector and the semantic feature vector are cross-fused by the decoder, and the obtained fusion vector is converted into an output text. Based on the output text and the label data included in the sample pair, the parameters of the speech recognition model architecture are adjusted to obtain a paralinguistic labeling model. By fusing the semantic features corresponding to the prompt text and the speech features extracted by the encoder through the decoder, the dual-supervised model training mode can be implemented to improve the labeling accuracy of the paralinguistic labeling model. The paralinguistic recognition and text conversion are processed uniformly by using Whisper, which can avoid the collaboration difficulty between independent models.
[0078] Figure 4 is a schematic diagram of a decoder outputting a target text sequence according to an example embodiment, as shown in Figure 4 , which includes.
[0079] SOS (input sequence), LANGUAGE (language sequence), REGION (region sequence), TRANSCRIBE (transcription), Begin time (start time), Text token (text token), End time (end time), NO TIMESTAMPS (no timestamp), and EOS (output sequence).
[0080] In the embodiments of the present disclosure, the prompt text is input into the decoder, and the language sequence extraction and region sequence extraction methods are used to extract features from the speech sequence, region sequence, etc. after labeling, and the extracted features are converted into a feature vector. Finally, the ext-only tranacription (pure text translation) output target text sequence without a timestamp is returned. The Tranacription allgned in utterance leve (utterance level strategy translation) with a timestamp is performed to obtain the target text sequence corresponding to the start time, end time, and text token.
[0081] Figure 5 is a flowchart of another paralinguistic labeling model training method according to an example embodiment, as shown in Figure 5 , which includes the following steps.
[0082] In step S51, speech features of speech or semantic features of text in the sample pair are extracted.
[0083] In the embodiments of the present disclosure, the speech features of the sample pair can be extracted by an encoder. The semantic features of the text in the sample pair can be extracted by a decoder.
[0084] In step S52, the speech features are converted into a speech feature vector or the semantic features are converted into a semantic feature vector.
[0085] In the embodiments of the present disclosure, the speech feature vector and the semantic feature vector can be realized by a vector conversion algorithm. The features of the speech and the text can be represented by vectors.
[0086] In step S53, output text corresponding to the speech feature vector or the semantic feature vector is output.
[0087] In the embodiments of the present disclosure, when the input is speech, the speech can be directly converted into text and the paralinguistic language in the text is labeled. The text converted from the speech and the text prompt corresponding to the training data in the training process of the paralinguistic language labeling model can be cross-feature fused to obtain output text containing paralinguistic language labeling.
[0088] In the embodiments of the present disclosure, when the input is text, the text and the text prompt corresponding to the training data in the training process of the paralinguistic language labeling model can be cross-feature fused to obtain output text containing paralinguistic language labeling.
[0089] In step S54, parameters of a speech recognition model architecture are adjusted based on the output text and label data included in the sample pair to obtain a paralinguistic language labeling model.
[0090] In the embodiments of the present disclosure, the loss of the speech recognition model architecture can be calculated based on the output text and the label data, and the parameters of the speech recognition model architecture are adjusted through the loss to obtain the paralinguistic language labeling model.
[0091] In the embodiments of the present disclosure, a large amount of data can be used to train the paralinguistic language labeling model. For example, a speech recognition architecture with 1B parameters is trained for millions of hours to obtain the paralinguistic language labeling model.
[0092] In the embodiments of the present disclosure, the speech recognition model architecture is trained by single supervision through the input of text or speech, so that the final paralinguistic language labeling model can recognize speech or text alone, and can accurately obtain target text containing paralinguistic language labeling.
[0093] In the embodiments of the present disclosure, the sample pair is determined in at least one of the following manners: generating text and speech based on a model, combining the text and the speech into a sample pair; embedding a paralinguistic speech stream in the speech to obtain synthesized speech, and embedding paralinguistic text in the text to obtain synthesized text, and combining the synthesized speech and the synthesized text into a sample pair; and labeling the paralinguistic speech stream in the speech to obtain labeled speech, and converting the labeled speech into labeled text, and combining the labeled speech and the labeled text into a sample pair.
[0094] In the embodiments of the present disclosure, the generation of the text and the speech based on the model can be implemented by a preset generation model. The generation model can be obtained by training an existing synthesis model based on a sample pair composed of the text and paralinguistic text, or a sample pair composed of the speech or a paralinguistic speech stream.
[0095] In the embodiments of the present disclosure, the synthesized speech can be obtained by embedding the paralinguistic speech stream in the speech through forced alignment and language analysis, and the synthesized text can be obtained by embedding the paralinguistic text in the text.
[0096] In the embodiments of the present disclosure, the labeled speech can be obtained by labeling the paralinguistic speech stream in the speech through a preset paralinguistic speech labeling model. The paralinguistic speech labeling model can be obtained by training an existing labeling model based on a sample pair composed of the speech and paralinguistic labeling.
[0097] In the embodiments of the present disclosure, the sample pair is obtained by combining the three methods of generation, splicing and mining, which solves the problem of data scarcity and enriches the paralinguistic types and context scenarios in the sample pair.
[0098] The paralinguistic labeling method provided by the embodiments of the present disclosure inputs the speech and the text as input data of the model training sample into the speech recognition model architecture, adopts a double supervision model training manner to train the speech recognition model architecture, and improves the robustness of the paralinguistic labeling model. The paralinguistic labeling model can identify the speech features and text semantics of the speech and the text, and accurately output the output text containing paralinguistic labeling after feature fusion of the speech features and the text semantics.
[0099] The paralinguistic labeling method provided by the embodiments of the present disclosure inputs the speech and the text as input data of the model training sample into the speech recognition model architecture, adopts a double supervision model training manner to train the speech recognition model architecture, and improves the robustness of the paralinguistic labeling model. The paralinguistic labeling model can identify the speech features and text semantics of the speech and the text, and accurately output the output text containing paralinguistic labeling after feature fusion of the speech features and the text semantics. Figure 6 is a block diagram of a synthesis data-driven paralinguistic labeling apparatus 100 according to an example embodiment. Referring to Figure 6 The apparatus includes an acquisition unit 101 and a processing unit 102.
[0100] The acquisition unit 101 is configured to acquire language information of a paralinguistic language to be labeled.
[0101] The processing unit 102 is configured to identify the language information by using a preset secondary language labeling model to obtain a target text containing secondary language labeling, wherein the secondary language labeling model is obtained by adjusting a speech recognition model architecture based on output text output by the speech recognition model architecture and label data included in a sample pair, the output text is obtained by fusing a speech feature vector and a semantic feature vector and converting the fused vector into text, and the speech feature vector and the semantic feature vector are obtained by respectively extracting features of speech and text included in the sample pair and converting the obtained speech features and semantic features into vectors.
[0102] According to some embodiments of the present disclosure, the secondary language labeling model is obtained by adjusting the speech recognition model architecture in the following manner: extracting speech features of speech in a sample pair and converting the speech features into a speech feature vector; obtaining text corresponding to the speech in the sample pair and determining prompt text corresponding to the text, wherein the prompt text does not include secondary language; extracting semantic features of the prompt text and converting the semantic features into a semantic feature vector; cross-fusing the speech feature vector and the semantic feature vector and converting the fused vector into output text; and adjusting parameters of the speech recognition model architecture based on the output text and label data included in the sample pair to obtain the secondary language labeling model.
[0103] According to some embodiments of the present disclosure, the secondary language labeling model is obtained by adjusting the speech recognition model architecture in the following manner: extracting speech features of speech or semantic features of text in a sample pair; converting the speech features into a speech feature vector or the semantic features into a semantic feature vector; outputting output text corresponding to the speech feature vector or the semantic feature vector; and adjusting parameters of the speech recognition model architecture based on the output text and label data included in the sample pair to obtain the secondary language labeling model.
[0104] According to some embodiments of the present disclosure, the sample pair is determined in at least one of the following manners: generating text based on a model and generating speech, and combining the text and the speech into a sample pair; embedding a secondary language speech stream in speech to obtain synthesized speech and embedding secondary language text in text to obtain synthesized text, and combining the synthesized speech and the synthesized text into a sample pair; labeling a secondary language speech stream in speech to obtain labeled speech and converting the labeled speech into labeled text, and combining the labeled speech and the labeled text into a sample pair.
[0105] According to some embodiments of the present disclosure, the language information includes text and / or speech.
[0106] With reference to the apparatus in the above-described embodiments, specific manner in which various modules perform operations has been described in detail in embodiments relating to the method, and thus will not be described in detail here.
[0107] Figure 7 is a block diagram of an apparatus 200 for synthesizing data-driven paralinguistic labeling according to an exemplary embodiment. The apparatus 200 can be provided as a terminal. For example, the apparatus 200 can be a mobile phone, a computer, a digital broadcasting terminal, a message transmitting / receiving device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0108] Referring to Figure 7 , the apparatus 200 can include one or more of the following components: a processing component 202, a memory 204, a power supply component 206, a multimedia component 208, an audio component 210, an input / output (I / O) interface 212, a sensor component 214, and a communication component 216.
[0109] The processing component 202 typically controls overall operations of the apparatus 200, such as operations associated with displaying, making phone calls, data communications, camera operations, and recording operations. The processing component 202 can include one or more processors 220 to execute instructions stored in the memory 204 to complete all or part of steps of the above-described methods. In addition, the processing component 202 can include one or more modules to facilitate interaction between the processing component 202 and other components. For example, the processing component 202 can include a multimedia module to facilitate the interaction between the multimedia component 208 and the processing component 202.
[0110] The memory 204 is configured to store various types of data to support operations of the apparatus 200. Examples of these data include instructions for any applications or methods operating on the apparatus 200, contact data, phonebook data, messages, pictures, videos, etc. The memory 204 can be implemented by any type of volatile or nonvolatile memory devices or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0111] The power supply component 206 supplies electric power to various components of the apparatus 200. The power supply component 206 can include a power management system, one or more power sources, and other components associated with generating, managing and distributing electric power for the apparatus 200.
[0112] The multimedia component 208 includes a display for the device 200 and a screen for providing an output interface between the device 200 and a user. In some embodiments, the screen can include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from a user. The touch panel includes one or more touch sensors for sensing touch, swiping and gestures on the touch panel. The touch sensors can not only sense a boundary of a touching or swiping action, but also detect duration and pressure associated with the touching or swiping action. In some embodiments, the multimedia component 208 includes a front camera and / or a rear camera. When the device 200 is in an operation mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front and rear cameras can be a fixed optical lens system or have a focal length and optical zoom capability.
[0113] The audio component 210 is configured to output and / or input audio signals. For example, the audio component 210 includes a microphone (MIC) for receiving an external audio signal when the device 200 is in an operation mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in the memory 204 or transmitted via the communication component 216. In some embodiments, the audio component 210 also includes a speaker for outputting audio signals.
[0114] The I / O interface 212 provides an interface between the processing component 202 and peripheral interface modules, which can be a keypad, a click wheel, buttons, and the like. The buttons can include, but are not limited to, a home button, a volume button, a start button, and a lock button.
[0115] The sensor component 214 includes one or more sensors for providing status assessments of various aspects of the device 200. For example, the sensor component 214 can detect an open / closed position of the device 200, relative positioning of components, such as a display and a keypad of the device 200, a change in position of the device 200 or a component of the device 200, the presence or absence of user contact with the device 200, the orientation or acceleration / deceleration of the device 200, and a temperature change of the device 200, among a plethora of other examples. The sensor component 214 can include a proximity sensor configured to detect presence of a nearby object without any physical touch. The sensor component 214 can further include a light sensor (e.g., a CMOS or CCD image sensor) configured to work in conjunction with the camera component 208. In some embodiments, the sensor component 214 can also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0116] The communication component 216 is configured to facilitate wired or wireless communication between the apparatus 200 and other devices. The apparatus 200 can access a wireless network based on a communication standard, such as WiFi, 2G, or 3G, or a combination thereof. In an example embodiment, the communication component 216 receives a broadcast signal or broadcast related information from an external broadcast management system via a broadcast channel. In an example embodiment, the communication component 216 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) techniques, infrared data association (IrDA) techniques, ultra-wideband (UWB) techniques, Bluetooth (BT) techniques, and other techniques.
[0117] In an example embodiment, the apparatus 200 can be implemented using one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, micro-controllers, microprocessors or other electronic devices, to perform the above-described methods.
[0118] In an example embodiment, a non-transitory computer-readable storage medium, such as the memory 204 including instructions, is also provided, which can be executed by the processor 220 of the apparatus 200 to perform the above-described methods. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disc, and an optical data storage device, etc.
[0119] In some embodiments of the present disclosure, a storage medium, which can be a non-transitory computer-readable storage medium, is provided.
[0120] In some embodiments of the present disclosure, when instructions in the storage medium are executed by a processor of an electronic device, the processor of the electronic device is enabled to perform the above-described related secondary language labeling method.
[0121] Embodiments of the present disclosure further provide a computer program product, including a computer program, which, when executed by a processor, implements the secondary language labeling method related in any of the above-described embodiments.
[0122] Those of skill would further appreciate that the various illustrative logical blocks, modules, and steps described in connection with the disclosure herein can be implemented by electronic hardware, computer software, or combinations of both. The various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality, without limitation. Depending upon the implementation, the techniques described herein can be implemented in various ways, consistent with the disclosure herein.
[0123] In the detailed description above, reference is made to the accompanying drawings, which form a part hereof, and in which are shown by way of illustration various aspects of the disclosure. In this regard, directional terminology, such as "center," "longitudinal," "lateral," "length," "width," "thickness," "upper," "lower," "front," "back," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," "outer," "clockwise," "counterclockwise," "axial," "radial," "circumferential," and the like, is used herein for the purpose of illustration only. Because the described devices can be positioned in a number of different orientations, the directional terminology can be used for illustration purposes only and is not intended to be limiting. It is to be understood that other aspects can be utilized and structural or logical changes can be made without departing from the scope of the disclosure. The following detailed description, therefore, is not to be taken in a limiting sense, as the scope of the disclosure is defined by the appended claims.
[0124] It is to be understood that the features of various aspects described herein can be combined with each other, unless specifically noted otherwise. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items; it will be understood that the terms "coupled," "connected," "attached," "joined," "connected," "fixed" and the like, as used herein, are not limited to direct connections only. For example, such terms can also pertain to an indirect connection or an integral connection, and can pertain to a mechanical or electrical connection or an interaction between two elements, unless otherwise explicitly stated. The specific meaning of the above terms should be understood in light of the specific context in which they are used, as will be apparent to those of ordinary skill in the art.
[0125] Further, the word "over" used in the description of components, elements or material layers "over" other components, elements or material layers means that the component, element or material layer is positioned or disposed on the other component, element or material layer such that one or more additional components, elements or layers can be disposed between the surface and the component, element or material layer. However, the word "over" can optionally also mean that the component, element or material layer is directly positioned or disposed on the other component, element or material layer, e.g., in direct contact with the surface.
[0126] It should be understood that spatially relative terms, such as "above", "upper", "below", and "lower" and the like, can be used herein for ease of description to describe one element's or feature's relationship to another element(s) or feature(s) as illustrated in the figures. Such spatially relative terms can be intended to encompass different orientations of the device in use or operation in addition to the orientation depicted in the figures. For example, if the device in the figures is turned over, elements described as above other elements or features would then be oriented below the other elements or features. Thus, the terms "above" and "below" can encompass both an orientation of above and below. The device can be otherwise oriented (e.g., rotated 90 degrees or at other orientations) and the spatially relative descriptors used herein interpreted accordingly.
[0127] Although terms such as "first", "second", and "third" can be used herein to describe various elements, components, regions, layers or sections, these elements, components, regions, layers or sections are not limited by these terms. Instead, these terms are only used to distinguish one element, component, region, layer or section from another element, component, region, layer or section. Thus, the first element, component, region, layer or section mentioned in the examples described herein can also be referred to as a second element, component, region, layer or section without departing from the teachings of the examples. In addition, the terms "first", "second" are used for description purposes only and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the technical features indicated. Thus, the features defined with "first", "second" can explicitly or implicitly include at least one of the features.
[0128] It can be further understood that the terms "first", "second", and the like are used to describe various information, but the information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other, and do not indicate a specific order or importance. In fact, the expressions "first", "second", and the like can be used interchangeably. For example, the first information can also be referred to as the second information, and similarly, the second information can also be referred to as the first information without departing from the scope of the disclosure.
[0129] In the description herein, the meaning of "a plurality" is at least two, referring to two or more, e.g., two, three, etc., unless the context clearly dictates otherwise. Similar for other quantifiers. The singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. Further, the terms "comprise," "comprising," "comprises," and the like can be used interchangeably with the term "include," "including," "includes," and the like. The terms "comprise," "comprising," "comprises," and the like can each be construed as permitting as set out herein rather than other inhibitions.
[0130] It should be understood that features of some of the various embodiments of the present disclosure described herein can be combined with one another, unless specifically noted otherwise. As used in this document, the term "and / or" includes any (and all) one or more of the associated listed items and the term "comprises / comprising" means "including without limitation" to the extent that the term "comprises / comprising" is used in the corresponding claim. The term "and / or" describes association between associated objects and indicates that three relationships can exist, for example, A and / or B can mean: A exists alone, A and B exist together, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects. Similarly, "at least one of... " includes any (and all) one or more of the associated listed items.
[0131] Further, the word "exemplary" is used herein to mean serving as an example, instance, or illustration. Any aspect or design described herein as "exemplary" is not necessarily to be construed as advantageous over other aspects or designs. Rather, the exemplary aspects are used to present concepts in a concrete fashion. As used in this document, the term "or" is intended to mean an inclusive "or" rather than an exclusive "or." That is, unless specified otherwise, or clear from context, "X employs A or B" is intended to mean any of the natural inclusive permutations. That is, if X employs A; X employs B; or X employs both A and B, then "X employs A or B" is satisfied under any of the foregoing instances. In addition, the articles "a" and "an" as used in this document are generally taken to mean "one or more" unless otherwise indicated.
[0132] Similarly, although this disclosure has been shown and described with respect to one or more implementations, equivalent variations and modifications will occur to those skilled in the art upon reading and understanding the specification and drawings. This disclosure includes all such modifications and variations and is limited only by the scope of the claims. In particular, with respect to the various functions performed by the components described above (e.g., elements, resources, etc.), unless otherwise indicated, the terminology used to describe such components is intended to correspond to any component (functionally equivalent) that performs the specific function of the described component, even if structurally not equivalent to the disclosed structure. Furthermore, although specific features of this disclosure may have been disclosed with respect to only one of several implementations, such features may be combined with one or more other features of other implementations, as may be desired and advantageous to any given or particular application. Moreover, with regard to the terms “comprising,” “owning,” “having,” “having,” or variations thereof as used in this disclosure, such terms are intended to be inclusive in a manner similar to the term “including.”
[0133] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein.
[0134] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. A synthetic data-driven paralinguistic labeling method, characterized by, The method comprises: acquiring language information of a sub-language to be labeled, the language information comprising text and / or voice; recognizing the language information through a preset sub-language labeling model to obtain target text, the target text containing sub-language labeling, the sub-language labeling model being obtained by adjusting a speech recognition model architecture through output text output by the speech recognition model architecture and label data included in a sample pair, the output text being obtained by fusing a voice feature vector and a semantic feature vector through a cross-attention layer of a decoder and converting the obtained fused vector into text, the voice feature vector being obtained by extracting voice features of voice in the sample pair through an encoder and converting the voice features into a voice feature vector, the semantic feature vector being obtained by extracting semantic features of a prompt text through the decoder and converting the semantic features into a semantic feature vector, the prompt text being text included in the sample pair from which sub-language is removed; the sample pair being determined in at least one of the following ways: generating text and voice based on a generation model, and combining the text and the voice into a sample pair; embedding a sub-language voice stream in voice to obtain synthesized voice and embedding sub-language text in text to obtain synthesized text, and combining the synthesized voice and the synthesized text into a sample pair; labeling a sub-language voice stream in voice to obtain labeled voice and converting the labeled voice into labeled text, and combining the labeled voice and the labeled text into a sample pair; wherein the generation model is trained based on a sample composed of text and sub-language text; the synthesized text is obtained by embedding a sub-language voice stream in voice to obtain synthesized voice and embedding sub-language text in text; and the labeled voice is obtained based on a preset sub-language voice labeling model that labels a sub-language voice stream in voice, the sub-language voice labeling model being obtained by training a labeling model through a sample pair composed of voice and sub-language labeling.
2. The method of claim 1, wherein, The sub-language labeling model adjusts the speech recognition model architecture in the following way: extracting voice features of voice or semantic features of text in a sample pair; converting the voice features into a voice feature vector or the semantic features into a semantic feature vector; outputting output text corresponding to the voice feature vector or the semantic feature vector; adjusting parameters of the speech recognition model architecture based on the output text and label data included in the sample pair to obtain the sub-language labeling model.
3. A synthetic data driven paralinguistic labeling apparatus, comprising: comprises: an acquisition unit configured to acquire language information of a sub-language to be labeled, the language information comprising text and / or voice; The processing unit is configured to identify the language information by using a preset minor language labeling model to obtain a target text containing minor language labeling, wherein the minor language labeling model is obtained by adjusting a speech recognition model architecture based on output text output by the speech recognition model architecture and label data included in a sample pair, the output text is obtained by fusing a speech feature vector and a semantic feature vector through a cross-attention layer of a decoder and converting the obtained fused vector into text, the speech feature vector is obtained by extracting speech features of speech in the sample pair through an encoder and converting the speech features into a speech feature vector, and the semantic feature vector is obtained by extracting semantic features of a prompt text through the decoder and converting the semantic features into a semantic feature vector, and the prompt text is obtained by removing minor language from text included in the sample pair; The sample pair is determined in at least one of the following manners: generating text and speech based on a generation model, and combining the text and the speech into a sample pair; embedding a minor language speech stream in the speech to obtain synthesized speech, and embedding minor language text in the text to obtain synthesized text, and combining the synthesized speech and the synthesized text into a sample pair; labeling a minor language speech stream in the speech to obtain labeled speech, and converting the labeled speech into labeled text, and combining the labeled speech and the labeled text into a sample pair. The generation model is trained based on a sample composed of text and minor language text; the synthesized text is obtained by embedding a minor language speech stream in the speech to obtain synthesized speech, and embedding minor language text in the text; and the labeled speech is obtained by labeling a minor language speech stream in the speech based on a preset minor language speech labeling model, and the minor language speech labeling model is obtained by training a labeling model based on a sample pair composed of speech and minor language labeling.
4. An electronic device, comprising: Comprise: a processor; a memory for storing computer programs or instructions executable by the processor; wherein the processor is configured to execute the computer programs or instructions to implement the steps of the synthetic data-driven minor language labeling method of any one of claims 1 to 2.
5. A storage medium, characterized by The storage medium stores computer programs or instructions, which, when executed by the processor of the electronic device, enable the processor of the electronic device to execute the synthetic data-driven minor language labeling method of any one of claims 1 to 2.
Citation Information
Patent Citations
Voice processing method and device, equipment and storage medium
CN116913278A