Data labeling method and device, electronic device, and storage medium

By performing speech-to-text and paralinguistic information recognition on speech data, and combining Whisper and Large Language Model (LLM) for automated annotation, the problem of low efficiency in speech data annotation is solved, achieving efficient and accurate multi-dimensional annotation and reducing manual intervention.

CN119889324BActive Publication Date: 2025-10-28CHINA TELECOM ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411805039.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-09
Publication Date
2025-10-28
Estimated Expiration
2044-12-09

AI Technical Summary

Technical Problem

Existing technologies for voice data annotation are inefficient and inaccurate. In particular, the annotation of complex paralinguistic information relies on manual labor, which is time-consuming, labor-intensive, and difficult to process effectively.

Method used

The initial annotation results are obtained by performing speech-to-text and paralinguistic information recognition on the speech data. These results are then input into a pre-trained natural language processing model for correction and optimization to generate target annotation results. The Whisper model is used for multi-task parallel processing, and the Large Language Model (LLM) is combined with context analysis and label optimization.

Benefits of technology

It automates speech data annotation, improving efficiency and accuracy, reducing the cost and workload of manual annotation, and generating multi-dimensional paralinguistic information tags such as emotion, tone, and speaker variations, thereby improving annotation quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119889324B_ABST
    Figure CN119889324B_ABST
Patent Text Reader

Abstract

This invention provides a data annotation method, apparatus, electronic device, and storage medium, comprising: acquiring speech data to be annotated in response to a data annotation request; performing speech-to-text processing and paralinguistic information recognition processing on the speech data to obtain an initial annotation result; wherein the initial recognition result includes multiple sub-texts and paralinguistic information tags corresponding to each sub-text; inputting the initial annotation result into a pre-trained natural language processing model to correct and / or optimize the paralinguistic information tags to obtain a target annotation result, and feeding back the target annotation result. This invention automates speech data annotation, annotating paralinguistic information tags with higher efficiency and accuracy, and reducing the cost and workload of data annotation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence, and in particular to data annotation methods and apparatus, electronic devices, and storage media. Background Technology

[0002] With the continuous development of artificial intelligence technology, various models are being applied more and more widely in various fields.

[0003] Data annotation is an important step before training a model. It involves adding labels or annotations to the raw sample data so that the model can understand and process the data.

[0004] In related technologies, speech data annotation is typically done manually. This method is time-consuming and labor-intensive, especially for speech data with complex information. Relying solely on manual annotation is not only inefficient but may also affect the accuracy of the annotation. Summary of the Invention

[0005] In view of the above problems, data annotation methods, apparatuses, electronic devices, and storage media are proposed to overcome or at least partially solve the above problems, including:

[0006] A data annotation method, the method comprising:

[0007] In response to a data annotation request, obtain the speech data to be annotated;

[0008] The speech data is processed for speech-to-text conversion and paralinguistic information recognition to obtain initial annotation results; wherein, the initial recognition results include multiple sub-texts and paralinguistic information tags corresponding to each sub-text;

[0009] The initial annotation results are input into a pre-trained natural language processing model to correct and / or optimize the sub-language information labels, thereby obtaining the target annotation results, and the target annotation results are fed back.

[0010] Optionally, the step of performing speech-to-text processing and paralinguistic information recognition processing on the speech data to obtain the initial annotation result includes:

[0011] The speech data is input into a pre-trained speech recognition model for parallel speech-to-text processing and paralinguistic information recognition processing to obtain the initial annotation result; wherein, the paralinguistic information labels in the initial recognition result include emotion information labels, tone information labels, and role information labels.

[0012] Optionally, inputting the initial annotation results into a pre-trained natural language processing model to correct and / or optimize the paralinguistic information labels includes:

[0013] Determine the associated subtexts for each subtext, and correct and / or optimize the secondary language information tags corresponding to each subtext based on the associated subtexts, including:

[0014] Based on the associated subtext, determine whether the secondary language information tag is incorrect; if the secondary language information tag is incorrect, correct the secondary language information tag.

[0015] And / or, based on the associated sub-text, determine whether there is a change in tone in the speech source of each sub-text; if there is a change in tone in the speech source, update the tone information label to a tone change information label.

[0016] And / or, based on the associated subtext, determine whether there is an emotional change in the voice source. If there is an emotional change in the voice source, update the emotional information tag to an emotional change information tag.

[0017] Optionally, the feedback of the target annotation results includes:

[0018] Determine the confidence level corresponding to the target annotation result;

[0019] If the confidence level is greater than or equal to the preset confidence level, the target labeling result is fed back.

[0020] If the confidence level is less than the preset confidence level, a request is initiated to correct the target labeling result.

[0021] Optionally, after initiating a request to correct the target annotation results, the method further includes:

[0022] Receive the corrected annotation result after correcting the target annotation result, and feed back the corrected annotation result.

[0023] Optionally, before performing speech-to-text processing and paralinguistic information recognition processing on the speech data, the method further includes:

[0024] The voice data is subjected to one or more of the following preprocessing methods: noise reduction, format conversion, cropping, and ambient noise muting.

[0025] Optionally, after feeding back the corrected annotation results, the method further includes:

[0026] Based on the corrected annotation results, the natural language processing model and / or the speech recognition model are optimized.

[0027] A data annotation device, the device comprising:

[0028] The unannotated speech data acquisition module is used to acquire the speech data to be annotated in response to the data annotation request;

[0029] The initial annotation result generation module is used to perform speech-to-text processing and paralinguistic information recognition processing on the speech data to obtain initial annotation results; wherein, the initial recognition results include multiple sub-texts and paralinguistic information tags corresponding to each sub-text;

[0030] The target annotation result processing module is used to input the initial annotation result into a pre-trained natural language processing model to correct and / or optimize the sub-language information labels, obtain the target annotation result, and feed back the target annotation result.

[0031] An electronic device includes a processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein the computer program, when executed by the processor, implements the data annotation method as described above.

[0032] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the data annotation method described above.

[0033] The embodiments of the present invention have the following advantages: by responding to a data annotation request, the speech data to be annotated is obtained; the speech data is processed by speech-to-text conversion and paralinguistic information recognition to obtain an initial annotation result; wherein, the initial recognition result includes multiple sub-texts and paralinguistic information labels corresponding to each sub-text; the initial annotation result is input into a pre-trained natural language processing model to correct and / or optimize the paralinguistic information labels to obtain the target annotation result, and the target annotation result is fed back, thereby realizing the automation of speech data annotation, annotating paralinguistic information labels with higher efficiency and accuracy, and reducing the cost and workload of data annotation. Attached Figure Description

[0034] To more clearly illustrate the technical solution of the present invention, the accompanying drawings used in the description of the present invention will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0035] Figure 1 This is a flowchart of the steps of a data annotation method provided in an embodiment of the present invention;

[0036] Figure 2This is a diagram illustrating an implementation process of data annotation using a model, according to an embodiment of the present invention.

[0037] Figure 3 This is a structural block diagram of a data annotation device provided in an embodiment of the present invention;

[0038] Figure 4 This is a schematic diagram of the structure of an electronic device for data annotation provided in an embodiment of the present invention. Detailed Implementation

[0039] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0040] In related technologies, the annotation of speech data mainly relies on manual work, especially the annotation of complex paralinguistic information (such as emotion, tone, speaker switching, etc.), which is time-consuming and labor-intensive.

[0041] While some related technologies have achieved automated annotation, most are limited to speech-to-text conversion and cannot effectively handle complex paralinguistic information such as emotion and tone, resulting in low annotation quality. Furthermore, even existing multi-task annotation systems are technically imperfect, often requiring multiple steps, each with its own independent model and processing module. Finally, the results from different modules need to be integrated and aligned, which is inefficient and prone to errors.

[0042] Therefore, this invention is based on the core idea of ​​performing speech recognition and paralinguistic information recognition on the speech data to be labeled to obtain initial labeling results, and then inputting the initial labeling results into a pre-trained natural language processing model for correction and / or optimization to obtain the target labeling results, thereby achieving automated labeling and reducing the complexity and time cost of manual labeling.

[0043] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0044] Reference Figure 1 The diagram illustrates a flowchart of a data annotation method according to an embodiment of the present invention, which may specifically include the following steps:

[0045] Step 101: In response to the data annotation request, obtain the speech data to be annotated;

[0046] The speech data to be labeled refers to the original speech sample data that needs to be labeled based on the needs of model training, data analysis, etc.

[0047] For example, voice data can be raw waveform data stored as audio files, such as WAV, MP3, and FLAC formats. Specifically, it can also be recordings used in certain application scenarios, such as conversation recordings or meeting recordings.

[0048] Step 102: Perform speech-to-text processing and paralinguistic information recognition processing on the speech data to obtain initial annotation results; wherein, the initial recognition results include multiple sub-texts and paralinguistic information tags corresponding to each sub-text;

[0049] Speech-to-text processing involves analyzing and recognizing data stored in audio format and then converting it into text format.

[0050] Paralinguistic information refers to information conveyed in language communication beyond actual vocabulary and grammatical structures, through sound, intonation, volume, rhythm, pauses, and other means. This information can influence the emotional tone, attitude, and intention of communication, conveying rich meaning even without actual linguistic content. Paralinguistic information recognition involves identifying paralinguistic information in speech data and forwarding it as corresponding paralinguistic information tags.

[0051] Subtext refers to the subtext within the complete text obtained after speech data has been converted to text. This can include subtexts such as sentences, phrases, and words. Each subtext can correspond to multiple paralinguistic information tags to reflect the paralinguistic information of that subtext.

[0052] For example, preset algorithms or models and data analysis can be used to perform speech-to-text processing and paralinguistic information recognition processing on speech data.

[0053] In some embodiments of the present invention, before performing speech-to-text processing and paralinguistic information recognition processing on the speech data, the method further includes:

[0054] The voice data is subjected to one or more of the following preprocessing methods: noise reduction, format conversion, cropping, and ambient noise muting.

[0055] In practical applications, data preprocessing before speech-to-text and paralinguistic information recognition can improve the accuracy of initial annotation results. Specifically, noise reduction improves speech clarity and enhances speech features; format conversion standardizes data formats and improves compatibility; cropping removes useless data and focuses on relevant information in the speech data; and ambient noise muting preserves only the speaker's voice, enhancing speech features.

[0056] In some embodiments of the present invention, the step of performing speech-to-text processing and paralinguistic information recognition processing on the speech data to obtain initial annotation results includes:

[0057] The speech data is input into a pre-trained speech recognition model for parallel speech-to-text processing and paralinguistic information recognition processing to obtain the initial annotation result; wherein, the paralinguistic information labels in the initial recognition result include emotion information labels, tone information labels, and role information labels.

[0058] Speech recognition models are models that can convert speech into text.

[0059] Emotional information tags can reflect the emotional information of the voice source corresponding to the labeled subtext, such as anger, calmness, tension, satisfaction, etc.

[0060] Tone information tags can reflect the tone information of the speaker, such as questioning, sarcasm, affirmation, etc.

[0061] Role information tags can reflect the role information of the voice source, such as customer, customer service, etc.

[0062] For example, the Whisper speech recognition model can be used to perform multi-task, parallel processing of speech data for speech-to-text conversion and paralinguistic information recognition.

[0063] Whisper is a model based on the Transformer architecture (a deep learning model architecture) specifically designed for multilingual speech-to-text tasks. Whisper performs exceptionally well in handling multiple languages ​​and dialects and maintains high recognition accuracy even in noisy environments.

[0064] In practical applications, the Whisper model needs to be fine-tuned and trained to enable it to perform multiple tasks such as speech-to-text conversion and paralinguistic information recognition, as detailed below:

[0065] 1) Data Preparation: First, a speech dataset containing multi-dimensional labels needs to be prepared to train the Whisper model. This speech dataset includes pre-annotated speech files, text transcriptions, emotion information labels, tone information labels, role information labels, etc. For example, for a customer service scenario, the speech dataset could be an audio recording of a conversation between a customer and a customer service representative.

[0066] 2) Model Fine-tuning: The Whisper model was fine-tuned using the aforementioned speech dataset to enable multi-task learning capabilities. During training, the model not only learns speech-to-text transcription but also learns to generate additional sentiment, tone, and role labels. For example, in a customer service scenario, the output obtained after inputting an audio file into the Whisper model is as follows:

[0067] {

[0068]

[0069] Here, "text" refers to the subtext, "emotion" is "neutral", "tone" is "polite", and "speaker" is "agent".

[0070] Furthermore, after the Whisper model is trained to convergence, the speech data to be labeled is input into the Whisper model.

[0071] For example, after inputting a recording of a conversation between a customer and a customer service representative into the Whisper model, the output is as follows:

[0072] Customer: "I've been waiting for a long time, why hasn't it been processed yet?"

[0073]

[0074] In this dialogue segment, the model identified the customer's emotion as "angry," their tone as "questioning," and their role as "customer."

[0075] Customer service: "I'm so sorry for the inconvenience. I'll check it for you right away."

[0076]

[0077] In this dialogue segment, the model identified the customer service representative's emotion as "calm," their tone as "apologetic," and their role as "agent."

[0078] Step 103: Input the initial annotation result into the pre-trained natural language processing model, correct and / or optimize the sub-language information labels to obtain the target annotation result, and feed back the target annotation result.

[0079] Natural language processing (NLP) models are computer models used to process and understand human language. By inputting the initial annotation results into a pre-trained NLP model, the paralinguistic information labels can be corrected and / or optimized, making the paralinguistic information labels more accurate and reflecting more dimensions of information.

[0080] In practical applications, natural language processing models can be large language models (LLMs), which are deep learning models with a large number of parameters and complex structures, capable of processing and generating natural language text. LLM models can further expand paralinguistic information labels to reflect more dimensions of information.

[0081] For example, the paralinguistic information labels can be further refined to identify potential paralinguistic information; for instance, a tone label that appears to be "peaceful" can be analyzed using a natural language processing model combined with contextual analysis to identify its potential "sarcasm" tone label, thereby optimizing the paralinguistic information labels.

[0082] In this way, by further refining and optimizing the initial annotation results using the natural language processing model, multi-dimensional labels including speech content, emotion, tone, and speaker are automatically generated, thereby significantly improving annotation efficiency and accuracy and reducing the cost and workload of manual annotation.

[0083] In some embodiments of the present invention, the step of inputting the initial annotation result into a pre-trained natural language processing model to correct and / or optimize the paralinguistic information labels includes:

[0084] Determine the associated subtexts for each subtext, and correct and / or optimize the secondary language information tags corresponding to each subtext based on the associated subtexts, including:

[0085] Based on the associated subtext, determine whether the secondary language information tag is incorrect; if the secondary language information tag is incorrect, correct the secondary language information tag.

[0086] Associated subtexts are other subtexts that are related to the current subtext, such as the context of the current subtext, other subtexts from the same language source, etc.

[0087] For example, if the role information label of a certain subtext is "user", and the natural language processing model is used to analyze the context, it is found that the actual voice message was sent by a customer service representative, then the role information label of the text is corrected to "customer service representative".

[0088] And / or, based on the associated sub-text, determine whether there is a change in tone in the speech source of each sub-text; if there is a change in tone in the speech source, update the tone information label to a tone change information label.

[0089] For example, if a customer's tone information label is "affirmative", and the tone changes from "questioning" to "affirmative" after the customer service representative answers the question, then the tone information label "affirmative" will be updated to the tone change information label "questioning to affirmative".

[0090] And / or, based on the associated subtext, determine whether there is an emotional change in the voice source. If there is an emotional change in the voice source, update the emotional information tag to an emotional change information tag.

[0091] For example, if a customer's tone information label is "calm", and the customer is reassured by customer service and then changes from "angry" to "calm", the tone information label "calm" will be updated to the emotional change information label "angry to calm".

[0092] In some examples, the Whisper model can initially label different speakers and their transition points, and the LLM can further supplement the labeling with the identities of different speakers and the emotional changes between the dialogues based on the text dialogue relationships. By combining Whisper's multi-task learning capabilities with the LLM's contextual understanding capabilities, multi-dimensional labels including speech content, emotion, tone, and speaker are automatically generated, thereby significantly improving labeling efficiency and accuracy and reducing the cost and workload of manual labeling.

[0093] In some embodiments of the present invention, the feedback of the target annotation result includes:

[0094] Determine the confidence level corresponding to the target annotation result;

[0095] If the confidence level is greater than or equal to the preset confidence level, the target labeling result is fed back.

[0096] If the confidence level is less than the preset confidence level, a request is initiated to correct the target labeling result.

[0097] Confidence level is used to measure the certainty or reliability of the model's output results. When the confidence level is greater than or equal to the preset confidence level, the target labeling results can be considered reliable, and feedback is then provided. When the confidence level is less than the preset confidence level, the target labeling results are reliable, but the target labeling results need to be corrected.

[0098] For example, after a request to correct the target annotation results is initiated, the recipient can correct the target annotation results through manual identification.

[0099] In some embodiments of the present invention, after initiating a request to correct the target annotation result, the method further includes:

[0100] Receive the corrected annotation result after correcting the target annotation result, and feed back the corrected annotation result.

[0101] In this embodiment, once the corrected annotation result is received, it is considered reliable, and feedback can be provided on the corrected annotation result.

[0102] In some embodiments of the present invention, after feeding back the corrected annotation result, the method further includes:

[0103] Based on the corrected annotation results, the natural language processing model and / or the speech recognition model are optimized.

[0104] In practical applications, the corrected annotation results can be manually corrected target annotation results. These corrected annotation results can then be used again as sample data for training natural language processing models (such as LLM models) and speech recognition models (such as Whisper models) to perform iterations and optimizations.

[0105] Specifically, through the system's feedback mechanism, if annotators or users find that the target annotation results do not match the actual situation, they can provide feedback through a reinforcement learning framework. The natural language processing model and speech recognition model will then use this feedback for iterative optimization. For example, if the model misclassifies a sentiment tag, the annotator can provide corrections, and the model will optimize its future annotation capabilities based on the corrected annotation results, thereby achieving continuous optimization and iteration of the model.

[0106] In some examples, such as Figure 2 As shown, a diagram illustrating the implementation process of labeling language data and providing reinforcement learning feedback is also provided, as detailed below:

[0107] 1) Input the speech data to be labeled and perform preprocessing;

[0108] 2) Input the data into the Whisper model for speech-to-text conversion and paralinguistic information recognition to obtain initial annotation results;

[0109] 3) Input the initial annotation results into the LLM model for correction and analysis to obtain the optimized target annotation results;

[0110] 4) Directly integrate and output the target annotation results with high confidence, and output the target annotation results with low confidence after manual annotation and feedback, and use the results of manual annotation and feedback to provide reinforcement learning feedback to Whisper and LLM.

[0111] 5) Generate results and provide feedback.

[0112] In some examples, the target annotation results can also be used as pre-annotated results in the manual data annotation process, which can help annotators process large amounts of dialogue data more efficiently, reduce manual intervention, and improve annotation efficiency.

[0113] It should be noted that, for the sake of simplicity, the method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments of the present invention are not limited to the described order of actions, because according to the embodiments of the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present invention.

[0114] The embodiments of the present invention have the following advantages: by responding to a data annotation request, the speech data to be annotated is obtained; the speech data is processed by speech-to-text conversion and paralinguistic information recognition to obtain an initial annotation result; wherein, the initial recognition result includes multiple sub-texts and paralinguistic information labels corresponding to each sub-text; the initial annotation result is input into a pre-trained natural language processing model to correct and / or optimize the paralinguistic information labels to obtain the target annotation result, and the target annotation result is fed back, thereby realizing the automation of speech data annotation, annotating paralinguistic information labels with higher efficiency and accuracy, and reducing the cost and workload of data annotation.

[0115] Furthermore, some embodiments of the present invention also have the following advantages:

[0116] Improve annotation efficiency: Automated annotation reduces a lot of manual annotation work, which is especially advantageous when processing large-scale speech data.

[0117] Multi-dimensional information generation: It not only annotates the speech content, but also generates complex paralinguistic information, covering emotions, tone, speaker changes, etc., which greatly enriches the annotation dimensions of speech data.

[0118] High-quality annotation: Leveraging the text understanding and semantic inference capabilities of LLM, more accurate and detailed labels such as sentiment and tone are generated, improving the overall quality of annotation.

[0119] Reference Figure 3 The diagram shows a structural schematic of a data annotation device according to an embodiment of the present invention, which may specifically include the following modules:

[0120] The unannotated speech data acquisition module 301 is used to acquire the speech data to be annotated in response to a data annotation request;

[0121] The initial annotation result generation module 302 is used to perform speech-to-text processing and paralinguistic information recognition processing on the speech data to obtain initial annotation results; wherein, the initial recognition results include multiple sub-texts and paralinguistic information tags corresponding to each sub-text;

[0122] The target annotation result processing module 303 is used to input the initial annotation result into a pre-trained natural language processing model, correct and / or optimize the sub-language information labels, obtain the target annotation result, and feed back the target annotation result.

[0123] In some embodiments of the present invention, the initial annotation result generation module 302 includes:

[0124] The speech recognition model processing submodule is used to input the speech data into a pre-trained speech recognition model for parallel speech-to-text processing and paralinguistic information recognition processing to obtain the initial annotation result; wherein, the paralinguistic information labels in the initial recognition result include emotion information labels, tone information labels, and role information labels.

[0125] In some embodiments of the present invention, the target annotation result processing module 303 includes:

[0126] The target annotation result processing submodule is used to determine the associated subtexts corresponding to each subtext, and to correct and / or optimize the secondary language information tags corresponding to each subtext based on the associated subtexts, including:

[0127] Based on the associated subtext, determine whether the secondary language information tag is incorrect; if the secondary language information tag is incorrect, correct the secondary language information tag.

[0128] And / or, based on the associated sub-text, determine whether there is a change in tone in the speech source of each sub-text; if there is a change in tone in the speech source, update the tone information label to a tone change information label.

[0129] And / or, based on the associated subtext, determine whether there is an emotional change in the voice source. If there is an emotional change in the voice source, update the emotional information tag to an emotional change information tag.

[0130] In some embodiments of the present invention, the target annotation result processing module 303 includes:

[0131] The confidence level determination submodule is used to determine the confidence level corresponding to the target annotation result;

[0132] The target annotation result feedback submodule is used to feed back the target annotation result if the confidence level is greater than or equal to a preset confidence level.

[0133] The target annotation result correction request submodule is used to initiate a request to correct the target annotation result if the confidence level is less than the preset confidence level.

[0134] In some embodiments of the present invention, the apparatus further includes:

[0135] The correction annotation result receiving module is used to receive the correction annotation result after correcting the target annotation result, and to feed back the correction annotation result.

[0136] In some embodiments of the present invention, the apparatus further includes:

[0137] The preprocessing module is used to perform one or more of the following data preprocessing operations on the voice data: noise reduction, format conversion, cropping, and ambient noise muting.

[0138] In some embodiments of the present invention, the apparatus further includes:

[0139] The model optimization module is used to optimize the natural language processing model and / or the speech recognition model based on the corrected annotation results.

[0140] Some embodiments of the present invention also provide an electronic device, which may include a processor, a memory, and a computer program stored in the memory and capable of running on the processor. When the computer program is executed by the processor, it implements the data annotation method described above.

[0141] Specifically, such as Figure 4 As shown, the electronic device includes a processor 401, a communication interface 402, a memory 403, and a communication bus 404, wherein the processor 401, the communication interface 402, and the memory 403 communicate with each other through the communication bus 404.

[0142] Memory 403 is used to store computer programs;

[0143] When processor 401 executes the program stored in memory 403, it performs the following steps:

[0144] In response to a data annotation request, obtain the speech data to be annotated;

[0145] The speech data is processed for speech-to-text conversion and paralinguistic information recognition to obtain initial annotation results; wherein, the initial recognition results include multiple sub-texts and paralinguistic information tags corresponding to each sub-text;

[0146] The initial annotation results are input into a pre-trained natural language processing model to correct and / or optimize the sub-language information labels, thereby obtaining the target annotation results, and the target annotation results are fed back.

[0147] In some embodiments of the present invention, the step of performing speech-to-text processing and paralinguistic information recognition processing on the speech data to obtain initial annotation results includes:

[0148] The speech data is input into a pre-trained speech recognition model for parallel speech-to-text processing and paralinguistic information recognition processing to obtain the initial annotation result; wherein, the paralinguistic information labels in the initial recognition result include emotion information labels, tone information labels, and role information labels.

[0149] In some embodiments of the present invention, the step of inputting the initial annotation result into a pre-trained natural language processing model to correct and / or optimize the paralinguistic information labels includes:

[0150] Determine the associated subtexts for each subtext, and correct and / or optimize the secondary language information tags corresponding to each subtext based on the associated subtexts, including:

[0151] Based on the associated subtext, determine whether the secondary language information tag is incorrect; if the secondary language information tag is incorrect, correct the secondary language information tag.

[0152] And / or, based on the associated sub-text, determine whether there is a change in tone in the speech source of each sub-text; if there is a change in tone in the speech source, update the tone information label to a tone change information label.

[0153] And / or, based on the associated subtext, determine whether there is an emotional change in the voice source. If there is an emotional change in the voice source, update the emotional information tag to an emotional change information tag.

[0154] In some embodiments of the present invention, the feedback of the target annotation result includes:

[0155] Determine the confidence level corresponding to the target annotation result;

[0156] If the confidence level is greater than or equal to the preset confidence level, the target labeling result is fed back.

[0157] If the confidence level is less than the preset confidence level, a request is initiated to correct the target labeling result.

[0158] In some embodiments of the present invention, after initiating a request to correct the target annotation result, the method further includes:

[0159] Receive the corrected annotation result after correcting the target annotation result, and feed back the corrected annotation result.

[0160] In some embodiments of the present invention, before performing speech-to-text processing and paralinguistic information recognition processing on the speech data, the method further includes:

[0161] The voice data is subjected to one or more of the following preprocessing methods: noise reduction, format conversion, cropping, and ambient noise muting.

[0162] In some embodiments of the present invention, after feeding back the corrected annotation result, the method further includes:

[0163] Based on the corrected annotation results, the natural language processing model and / or the speech recognition model are optimized.

[0164] Some embodiments of the present invention also provide a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, it implements the data annotation method described above.

[0165] Some embodiments of the present invention also provide a computer program product, including a computer program that, when executed by a processor, implements the data annotation method described above.

[0166] As the device embodiment is basically similar to the method embodiment, the description is relatively simple, and relevant parts can be found in the description of the method embodiment.

[0167] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.

[0168] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0169] The communication bus mentioned above can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.

[0170] The communication interface is used for communication between the aforementioned terminal and other devices.

[0171] The memory may include random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0172] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0173] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, apparatus, or computer program products. Therefore, embodiments of the present invention can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of the present invention can take the form of computer program products implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0174] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0175] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0176] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal equipment, causing a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0177] Although preferred embodiments of the present invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present invention.

[0178] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes the aforementioned element.

[0179] The above provides a detailed description of the data annotation method, apparatus, electronic device, and storage medium. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A data annotation method, characterized in that, The method includes: In response to a data annotation request, obtain the speech data to be annotated; The speech data is processed for speech-to-text conversion and paralinguistic information recognition to obtain initial annotation results; wherein, the initial recognition results include multiple sub-texts and paralinguistic information tags corresponding to each sub-text; The initial annotation results are input into a pre-trained natural language processing model to correct and / or optimize the sub-language information labels, thereby obtaining the target annotation results, and the target annotation results are fed back.

2. The method according to claim 1, characterized in that, The process of performing speech-to-text processing and paralinguistic information recognition on the speech data to obtain initial annotation results includes: The speech data is input into a pre-trained speech recognition model for parallel speech-to-text processing and paralinguistic information recognition processing to obtain the initial annotation result; wherein, the paralinguistic information labels in the initial recognition result include emotion information labels, tone information labels, and role information labels.

3. The method according to claim 2, characterized in that, The step of inputting the initial annotation results into a pre-trained natural language processing model to correct and / or optimize the sub-language information labels includes: Determine the associated subtexts for each subtext, and correct and / or optimize the secondary language information tags corresponding to each subtext based on the associated subtexts, including: Based on the associated subtext, determine whether the secondary language information tag is incorrect; if the secondary language information tag is incorrect, correct the secondary language information tag. And / or, based on the associated sub-text, determine whether there is a change in tone in the speech source of each sub-text; if there is a change in tone in the speech source, update the tone information label to a tone change information label. And / or, based on the associated subtext, determine whether there is an emotional change in the voice source. If there is an emotional change in the voice source, update the emotional information tag to an emotional change information tag.

4. The method according to claim 2, characterized in that, The feedback of the target annotation results includes: Determine the confidence level corresponding to the target annotation result; If the confidence level is greater than or equal to the preset confidence level, the target labeling result is fed back. If the confidence level is less than the preset confidence level, a request is initiated to correct the target labeling result.

5. The method according to claim 4, characterized in that, After initiating a request to correct the target annotation results, the method further includes: Receive the corrected annotation result after correcting the target annotation result, and feed back the corrected annotation result.

6. The method according to claim 1, characterized in that, Before performing speech-to-text processing and paralinguistic information recognition processing on the speech data, the method further includes: The voice data is subjected to one or more of the following preprocessing methods: noise reduction, format conversion, cropping, and ambient noise muting.

7. The method according to claim 5, characterized in that, After providing the corrected annotation results, the method further includes: Based on the corrected annotation results, the natural language processing model and / or the speech recognition model are optimized.

8. A data annotation device, characterized in that, The device includes: The unannotated speech data acquisition module is used to acquire the speech data to be annotated in response to the data annotation request; The initial annotation result generation module is used to perform speech-to-text processing and paralinguistic information recognition processing on the speech data to obtain initial annotation results; wherein, the initial recognition results include multiple sub-texts and paralinguistic information tags corresponding to each sub-text; The target annotation result processing module is used to input the initial annotation result into a pre-trained natural language processing model, correct and / or optimize the sub-language information labels, obtain the target annotation result, and feed back the target annotation result.

9. An electronic device, characterized in that, It includes a processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein the computer program, when executed by the processor, implements the data annotation method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, which, when executed by a processor, implements the data annotation method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Text-to-speech conversion method, device, electronic equipment and storage medium

    CN112765971A

  • Speech synthesis method and device, equipment and storage medium

    CN113990286A