Audio data processing method and system, electronic equipment and medium

By extracting and analyzing intermediate features in dialogue audio, using the end-to-end method or large model of Attention Encoder-Decoder for emotional annotation, the problem of high cost and low efficiency in traditional audio emotional annotation methods is solved, and more efficient and accurate emotional annotation is achieved.

CN120108430APending Publication Date: 2025-06-06UNICOM WOYUEDU TECH CULTURE CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510214438.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Traditional audio sentiment labeling methods have high cost and low efficiency, making it difficult to ensure accuracy and quality, and hinder the development of the field of deep learning.

Method used

An audio data processing method is proposed, including obtaining dialogue audio, standardizing processing, extracting intermediate features (such as emotional tendency features and acoustic features), and emotional annotation through the end-to-end method or large model of Attention Encoder-Decoder.

Benefits of technology

This method can effectively improve the accuracy of audio emotional labeling, reduce labeling costs, and improve efficiency, solving the problems of high and low labeling costs in traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108430A_ABST
    Figure CN120108430A_ABST
Patent Text Reader

Abstract

The invention relates to an audio data processing method and system, electronic equipment and a medium. The method comprises the following steps: acquiring a dialogue audio; performing standardization processing on the dialogue audio; extracting intermediate features from the dialogue audio; and according to the intermediate feature, determining an emotion label of the dialogue audio. According to the method, the emotion in the dialogue audio can be automatically labeled, and the problems of high labeling cost and low efficiency in a traditional method are effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the technical field of audio data processing, and in particular, to an audio data processing method, system, electronic device and medium. Background Art

[0002] With the booming development of artificial intelligence technology, especially the breakthrough progress of deep learning technology, in the third wave of artificial intelligence, speech recognition and natural language processing technologies have made significant leaps, opening up a broad market space for the audio data annotation industry and providing solid technical support. As an emerging practice in this field, the technical process of audio data annotation has gone through the stage of gradual evolution from early manual annotation to semi-automatic and even automatic annotation. This transformation has greatly improved the efficiency and accuracy of annotation work. Relying on machine learning and deep learning algorithms, automated annotation technology can now realize automatic classification and annotation of audio data. Although it cannot completely replace manual annotation, it has greatly reduced the manual burden and become an important force to promote the progress of the industry.

[0003] A variety of audio data annotation tools have emerged in the market. These tools integrate a variety of annotation functions, covering speaker role recognition, environmental scene description, multilingual processing, etc., fully responding to the needs of different application scenarios. Annotation methods are also becoming more diversified, including rhythm feature annotation, systematic annotation, and emotion recognition annotation, providing a powerful tool for in-depth analysis of audio data.

[0004] Faced with the growing requirements for annotation quality and efficiency, the audio data annotation industry has actively built a strict quality control system, striving for excellence in every link from data preprocessing, annotation process monitoring to the review and verification of annotation results. At the same time, by continuously optimizing the annotation process and introducing cutting-edge annotation tools and technologies, the industry has effectively reduced annotation costs while improving annotation efficiency, further promoting the maturity and development of audio data annotation technology. However, with the development of technology in the field of deep learning, emotional annotation has gradually become the underlying support capability for further optimization of AI-generated content. Large amounts of data require extremely high-cost manpower, and the accuracy and quality under manual operation are difficult to guarantee, which to a certain extent hinders the development of the deep learning field. Summary of the invention

[0005] The following is a summary of the subject matter described in detail herein. This summary is not intended to limit the scope of the claims.

[0006] The main purpose of the embodiments of the present disclosure is to propose an audio data processing method, system, electronic device and medium, which can improve the accuracy of audio emotion labeling.

[0007] A first aspect of an embodiment of the present application provides an audio data processing method, the method comprising:

[0008] Get the conversation audio;

[0009] Performing standardization processing on the conversation audio;

[0010] Extracting intermediate features from the conversation audio;

[0011] Determine the emotion label of the conversation audio according to the intermediate features.

[0012] An audio data processing method provided by the present disclosure has at least the following beneficial effects:

[0013] The method includes obtaining conversational audio; performing standardization on the conversational audio; extracting intermediate features from the conversational audio; and determining the emotional annotation of the conversational audio based on the intermediate features. The method can automatically annotate the emotions in the conversational audio, effectively solving the problems of high annotation cost and low efficiency in traditional methods.

[0014] In some embodiments, the step of obtaining the conversation audio includes the following steps:

[0015] Get first conversation samples on multiple platforms;

[0016] Screening the first dialogue samples based on automatic speech recognition technology to select the first dialogue samples with speech recognition rates higher than a first threshold as second dialogue samples;

[0017] Parameters for capturing target features are annotated in the second conversation sample to obtain the conversation-like audio.

[0018] In some embodiments, the step of normalizing the conversation audio comprises the following steps:

[0019] Adjusting the volume of the conversation-type audio according to the standard filter so that the volume of the conversation-type audio is within a preset range;

[0020] The sampling rate of the conversation audio after the volume is adjusted is adjusted to a second threshold.

[0021] In some embodiments, the intermediate features include sentiment tendency features and acoustic features;

[0022] The step of extracting intermediate features from the conversation audio comprises the following steps:

[0023] Extracting acoustic features from the conversational audio;

[0024] Extracting emotional tendency features from the conversation audio;

[0025] Determining the emotion label of the conversation audio according to the intermediate features includes:

[0026] Based on the acoustic features and the emotional tendency features, a first emotional label of the conversation audio is determined.

[0027] In some embodiments, the step of extracting emotional tendency features from the conversation audio comprises the following steps:

[0028] The end-to-end method based on Attention Encoder-Decoder performs emotional tendency recognition on the standardized conversation audio to obtain emotional tendency features; the end-to-end method of Attention Encoder-Decoder includes: LSTM based on attention mechanism, or Transformer.

[0029] In some embodiments, the intermediate features include style features and acoustic features;

[0030] The step of extracting intermediate features from the conversation audio comprises the following steps:

[0031] Converting the conversational audio into conversational text;

[0032] Extracting style features from the conversation text based on the large model;

[0033] Extracting acoustic features from the conversational audio;

[0034] Determining the emotion label of the conversation audio according to the intermediate features includes:

[0035] A second emotion annotation of the conversation audio is determined from the style features and the acoustic features based on the large model.

[0036] In some embodiments, after determining the emotion annotation of the conversation audio according to the intermediate features, the method further includes:

[0037] Comparing the first emotion annotation with the second emotion annotation to obtain a comparison result;

[0038] According to the comparison result, the first emotion annotation or the second emotion annotation is selected as the final emotion annotation.

[0039] A second aspect of an embodiment of the present application provides an audio data processing system, the system comprising:

[0040] An audio acquisition module, used to acquire conversation audio;

[0041] A standard processing module, used for performing standard processing on the conversation audio;

[0042] A feature extraction module, used to extract intermediate features from the conversation audio;

[0043] The emotion annotation module is used to determine the emotion annotation of the dialogue audio according to the intermediate features.

[0044] A third aspect of an embodiment of the present application proposes an electronic device, at least one controller and a memory for communicating with the at least one controller; the memory stores instructions that can be executed by the at least one controller, and the instructions are executed by the controller so that the controller performs the audio data processing method described in the first aspect.

[0045] A fourth aspect of an embodiment of the present application proposes a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the audio data processing method as described above.

[0046] It can be understood that the beneficial effects of the second to fourth aspects compared with the related art are the same as the beneficial effects of the first aspect compared with the related art. Please refer to the relevant description in the first aspect, and no further details will be given here. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or related technical descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0048] Figure 1 It is a flowchart of the audio data processing method provided by the present application;

[0049] Figure 2 is a structural diagram of an audio data processing system proposed in an embodiment of the present application;

[0050] Figure 3 It is a schematic diagram of the structure of an electronic device proposed in an embodiment of the present application. DETAILED DESCRIPTION

[0051] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0052] It should be noted that, although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first", "second", etc. in the specification, claims and the above drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.

[0053] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0054] With the booming development of artificial intelligence technology, especially the breakthrough progress of deep learning technology, in the third wave of artificial intelligence, speech recognition and natural language processing technologies have made significant leaps, opening up a broad market space for the audio data annotation industry and providing solid technical support. As an emerging practice in this field, the technical process of audio data annotation has gone through the stage of gradual evolution from early manual annotation to semi-automatic and even automatic annotation. This transformation has greatly improved the efficiency and accuracy of annotation work. Relying on machine learning and deep learning algorithms, automated annotation technology can now realize automatic classification and annotation of audio data. Although it cannot completely replace manual annotation, it has greatly reduced the manual burden and become an important force to promote the progress of the industry.

[0055] A variety of audio data annotation tools have emerged in the market. These tools integrate a variety of annotation functions, covering speaker role recognition, environmental scene description, multilingual processing, etc., fully responding to the needs of different application scenarios. Annotation methods are also becoming more diversified, including rhythm feature annotation, systematic annotation, and emotion recognition annotation, providing a powerful tool for in-depth analysis of audio data.

[0056] Faced with the growing requirements for annotation quality and efficiency, the audio data annotation industry has actively built a strict quality control system, striving for excellence in every link from data preprocessing, annotation process monitoring to the review and verification of annotation results. At the same time, by continuously optimizing the annotation process and introducing cutting-edge annotation tools and technologies, the industry has effectively reduced annotation costs while improving annotation efficiency, further promoting the maturity and development of audio data annotation technology. However, with the development of technology in the field of deep learning, emotional annotation has gradually become the underlying support capability for further optimization of AI-generated content. Large amounts of data require extremely high-cost manpower, and the accuracy and quality under manual operation are difficult to guarantee, which to a certain extent hinders the development of the deep learning field.

[0057] An embodiment of the present application provides an audio data processing method, the method comprising the following steps S110 to S140:

[0058] Step S110, obtaining conversation audio;

[0059] Step S120, normalizing the conversation audio;

[0060] Step S130, extracting intermediate features from the conversation audio;

[0061] Step S140: Determine the emotion labeling of the conversation audio based on the intermediate features.

[0062] In step S110, the conversation audio may come from public conversation records on social media platforms, conversation clips in movies and TV series, professional emotion databases, and conversation samples recorded in specific situations, etc.;

[0063] In step S120, the standardization process includes but is not limited to: denoising and sampling rate unification;

[0064] In step S130, the intermediate features include but are not limited to: emotional tendency features, acoustic features, style features, etc.;

[0065] In step S140, a large model (such as Chat-GPT, DeepSeek) can be used to determine the emotional annotation of the conversation audio. For example, the intermediate features are input into the large model to generate the emotional annotation of the conversation audio according to the large model.

[0066] Similarly, an end-to-end method based on Attention Encoder-Decoder (such as a long short-term memory network based on an attention mechanism) can be used to determine the emotional annotation of conversational audio. For example, the intermediate features are input into the long short-term memory network to generate the emotional annotation of the conversational audio according to the long short-term memory network.

[0067] This method has at least the following advantages:

[0068] The method includes obtaining conversational audio; performing standardization on the conversational audio; extracting intermediate features from the conversational audio; and determining the emotional annotation of the conversational audio based on the intermediate features. The method can automatically annotate the emotions in the conversational audio, effectively solving the problems of high annotation cost and low efficiency in traditional methods.

[0069] In addition, obtaining the dialogue audio in step S110 includes the following steps S210 to S230:

[0070] Step S210, obtaining a first conversation sample of multiple platforms;

[0071] Step S220, screening the first dialogue samples based on the automatic speech recognition technology to select the first dialogue samples with a speech recognition rate higher than a first threshold as the second dialogue samples;

[0072] Step S230, marking parameters for capturing target features in the second conversation sample to obtain conversation-like audio.

[0073] In step S210, the platform may be a social media platform, a movie and TV series platform, a conversation recording platform (such as customer service), etc.

[0074] In step S220 and step S230, the automatic speech recognition (ASR) technology can be used to initially filter out the audio with a speech recognition rate lower than a preset threshold to reduce noise interference in subsequent processing. In the initial stage, a small amount of data needs to be manually annotated to capture the Mel-frequency cepstral coefficients (MFCCs), spectral centroid, spectral flatness, etc. of the target features, so as to mark the frequency domain characteristics and dynamic characteristics for the model, which is convenient for subsequent processing.

[0075] In addition, the step S120 of normalizing the conversation audio includes the following steps S310 and S320:

[0076] Step S310, adjusting the volume of the conversation-type audio according to the standard filter so that the volume of the conversation-type audio is within a preset range;

[0077] Step S320, adjusting the sampling rate of the conversation audio after the volume is adjusted to a second threshold.

[0078] You can first use a standard filter (such as DSP) to adjust the volume level of all audio files, normalizing the loudness of all audio to a preset and uniform dynamic range, thereby eliminating the uneven volume caused by differences in the original recording conditions (such as recording equipment, environmental noise, etc.). At the same time, the recorded audio needs to maintain a uniform sampling rate.

[0079] On the first aspect, in addition, the intermediate features include sentiment tendency features and acoustic features;

[0080] Extracting intermediate features from the conversation audio in step S130 includes the following steps S410 to S420:

[0081] Step S410, extracting acoustic features from the conversation audio;

[0082] Step S420, extracting emotional tendency features from the conversation audio;

[0083] Determining the emotion labeling of the dialogue audio according to the intermediate features in step S140 includes the following step S430:

[0084] Step S430: determining a first emotion label of the dialogue audio based on the acoustic features and the emotion tendency features.

[0085] Among them, acoustic features can be intonation, rhythm, volume changes, etc.

[0086] This method is based on acoustic features and emotional tendency features and can produce more accurate emotional labeling.

[0087] In addition, extracting emotional tendency features from the dialogue audio in step S420 includes the following steps:

[0088] The end-to-end method based on Attention Encoder-Decoder recognizes the emotional tendency of the standardized conversation audio to obtain the emotional tendency features; the end-to-end method of Attention Encoder-Decoder includes: LSTM based on the attention mechanism, or Transformer.

[0089] On the second aspect, the intermediate features include style features and acoustic features;

[0090] Extracting intermediate features from the conversation audio in step S130 includes the following steps S510 to S530:

[0091] Step S510, converting the conversation audio into conversation text;

[0092] Step S520, extracting style features from the dialogue text based on the large model;

[0093] Step S530, extracting acoustic features from the conversation audio;

[0094] In step S140, the emotion annotation of the dialogue audio is determined according to the intermediate features, including step S540:

[0095] Step S540: Determine a second emotion label for the dialogue audio from the style features and the acoustic features based on the large model.

[0096] The large model can be Chat-GPT or DeepSeek model.

[0097] Among them, style features can be high-pitched, low-pitched, calm, etc. Acoustic features can be intonation, rhythm, volume changes, etc.

[0098] This method can obtain more accurate sentiment annotations based on the context analysis capabilities of the large model.

[0099] After determining the emotion annotation of the conversation audio according to the intermediate features, the method further includes steps S610 and S620:

[0100] Step S610, comparing the first emotion annotation and the second emotion annotation to obtain a comparison result;

[0101] Step S620: Select the first emotion annotation or the second emotion annotation as the final emotion annotation according to the comparison result.

[0102] The comparison result may be a comparison of the emotion and intonation style of the first emotion annotation and the second emotion annotation. This method may compare the judgment result of the large model with the emotion and intonation style recognized by the attention mechanism to determine which emotion recognition result is more appropriate, and thus obtain a more accurate emotion annotation.

[0103] To facilitate understanding by those skilled in the art, an embodiment of the present application provides an audio data processing method, the method comprising the following steps S910 to S930:

[0104] Step S910: data set collection.

[0105] During the dataset collection phase, we focused on building a comprehensive and diverse conversational audio dataset.

[0106] This dataset needs to cover a wide range of emotion categories (such as happy, sad, angry, surprised, calm, etc.) and a variety of conversation styles (such as formal, casual, humorous, serious, etc.).

[0107] Data sources include but are not limited to public conversation records on social media platforms, dialogue clips from movies and TV series, professional emotion databases, and conversation samples recorded in specific situations.

[0108] To ensure the quality and accuracy of the data, automatic speech recognition (ASR) technology is used to initially filter out audio with a speech recognition rate below a preset threshold to reduce noise interference in subsequent processing.

[0109] When performing sentiment labeling, although the goal is to achieve automatic labeling, in the initial stage, a small amount of data needs to be manually labeled to capture the Mel-frequency cepstral coefficients (MFCCs), spectral centroid, spectral flatness, etc. of the target features, so as to mark the frequency domain characteristics and dynamic characteristics for the model.

[0110] Step S920: audio standardization processing.

[0111] Use the Standard filter to adjust the volume levels of all audio files, normalizing the loudness of all audio to a preset, uniform dynamic range.

[0112] At the same time, the recorded audio must maintain a uniform sampling rate. For example, the sampling rate of emotion annotation must be maintained above 44100HZ after testing, otherwise there will be obvious deviations in subsequent annotations.

[0113] Step S930: audio regularization and feature recognition.

[0114] First, audio emotion recognition is based on the standard emotion frequency domain. By learning using an end-to-end method based on Attention Encoder-Decoder, an estimate of the labeled discrete emotion features can be obtained, thereby realizing the recognition of sentence emotions.

[0115] First, we use sentiment dictionaries and sentiment analysis models (such as LSTM and Transformer based on attention mechanisms) to identify the sentiment in the text. At the same time, we combine the acoustic features of the audio signal, such as intonation, rhythm, and volume changes, to comprehensively analyze and obtain more accurate sentiment annotations.

[0116] Secondly, in order to improve accuracy, we added estimates based on large language models, using language style models to identify style elements such as vocabulary selection, sentence structure, and tone in the text, combined with features such as speaking speed and pronunciation habits in the audio to achieve automatic labeling of conversation style.

[0117] PROMPT needs to clearly express the format and tone, and the definition of emotion in order to generate accurately. The example PROMPT is: Please judge the tone style and emotion of the sentence I will provide next. Each line is regarded as a sentence. Emotions are selected from "happy", "surprised", "sad", "disgusted", "angry", "fear", "neutral", and "doubtful". The tone style is selected from "steady", "high", "low", "soothing", "whispering", "gentle", "playful", "serious", "loud", "cheerful", "nervous", "confident", and "stuttering".

[0118] The focus of emotion is on the feeling, while the focus of tone style is on the expression. You can mark 2 or more labels, put the important ones in front, and separate them with minus signs. In the order of sentences, only list the emotion and tone style columns in Excel. The following is the sentence that needs to be judged:

[0119] Use / n to separate sentences. It is also recommended to prohibit the large model from modifying the sentence separation in the system PROMPT. In the large model with intelligent agents, tabular output results can be obtained. Non-tabular output can be requested to print the file. The graphical results are shown in the following table:

[0120] Table 1

[0121] mood Tone and style anger High Tired - Neutral Low Confuse smooth Surprise-Confusion High neutral smooth Confuse smooth Serious-Neutral smooth Comfort-Neutral gentle Confuse smooth Doubt-Nervousness nervous Worry-helplessness Low

[0122] Then, the judgment result of the second aspect of the big model is compared with the emotion and intonation style recognized by the attention mechanism of the first aspect, and it is required to force the context to make a second judgment on which emotion recognition result is more appropriate, and to give the thinking process. This thinking process is not used for manual verification, but to improve accuracy. Then, according to the previous PROMPT example, the big model is required to give a table file corresponding to the text.

[0123] This application realizes the automatic annotation of emotions and conversation styles in dialogue audio by constructing a set of standard data annotation method processes, effectively solving the problems of high annotation cost and low efficiency in traditional methods, and providing strong support for the development of sentiment analysis and dialogue systems.

[0124] like Figure 2 One embodiment of the present application provides an audio data processing system, the system comprising:

[0125] The audio acquisition module 1100 is used to acquire conversation audio;

[0126] The standard processing module 1200 is used to perform standardization processing on the conversational audio;

[0127] The feature extraction module 1300 is used to extract intermediate features from the conversation audio;

[0128] The emotion annotation module 1400 is used to determine the emotion annotation of the dialogue audio according to the intermediate features.

[0129] It should be noted that the audio data processing system provided in this embodiment and the above-mentioned audio data processing method are based on the same inventive concept, so the relevant content of the above-mentioned audio data processing method is also applicable to the content of the audio data processing system, and therefore, it will not be repeated here.

[0130] like Figure 3 The embodiment of the present application further provides an electronic device, the electronic device comprising:

[0131] at least one memory;

[0132] at least one processor;

[0133] at least one program;

[0134] The programs are stored in the memory, and the processor executes at least one program to implement the audio data processing method described above in the present disclosure.

[0135] The electronic device may be any intelligent terminal including a mobile phone, a tablet computer, a personal digital assistant (PDA), a vehicle-mounted computer, etc.

[0136] The electronic device according to the embodiment of the present application is described in detail below.

[0137] The processor 1600 may be implemented by a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present invention.

[0138] The memory 1700 may be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1700 may store an operating system and other application programs. When the technical solution provided in the embodiment of this specification is implemented by software or firmware, the relevant program code is stored in the memory 1700, and the processor 1600 calls and executes the audio data processing method of the embodiment of the present invention.

[0139] Input / output interface 1800, used to implement information input and output;

[0140] The communication interface 1900 is used to realize the communication interaction between the device and other devices. The communication can be realized through a wired manner (such as USB, network cable, etc.) or a wireless manner (such as mobile network, WIFI, Bluetooth, etc.);

[0141] Bus 2000 , which transmits information between various components of the device (e.g., processor 1600 , memory 1700 , input / output interface 1800 , and communication interface 1900 );

[0142] The processor 1600 , the memory 1700 , the input / output interface 1800 , and the communication interface 1900 are connected to each other in communication within the device via the bus 2000 .

[0143] An embodiment of the present invention further provides a storage medium, which is a computer-readable storage medium. The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the above-mentioned audio data processing method.

[0144] The memory is a non-transitory computer-readable storage medium that can be used to store non-transitory software programs and non-transitory computer executable programs. In addition, the memory can include a high-speed random access memory and can also include a non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage device.

[0145] In some embodiments, the memory may include a memory remotely located relative to the processor, and the remote memory may be connected to the processor via a network. Examples of the above network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0146] The embodiments described in the present invention are intended to more clearly illustrate the technical solutions of the embodiments of the present invention and do not constitute a limitation on the technical solutions provided by the embodiments of the present invention. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present invention are also applicable to similar technical problems.

[0147] Those skilled in the art will appreciate that the technical solutions shown in the figures do not limit the embodiments of the present invention and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.

[0148] The device embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0149] Those skilled in the art will appreciate that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices may be implemented as software, firmware, hardware, or a suitable combination thereof.

[0150] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0151] It should be understood that in the present application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0152] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0153] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0154] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0155] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including multiple instructions to enable an electronic device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of each embodiment method of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, ROM), random access memory (Random Access Memory, RAM), disk or optical disk and other media that can store programs.

[0156] The above is a specific description of the preferred implementation of the embodiments of the present application, but the embodiments of the present application are not limited to the above-mentioned implementation methods. Technical personnel familiar with the field can also make various equivalent modifications or substitutions without violating the spirit of the embodiments of the present application. These equivalent modifications or substitutions are all included in the scope defined by the claims of the embodiments of the present application.

Claims

1. An audio data processing method, characterized in that: The method comprises: Get the conversation audio; Performing standardization processing on the conversation audio; Extracting intermediate features from the conversation audio; Determine the emotion label of the conversation audio according to the intermediate features.

2. The audio data processing method according to claim 1, characterized in that: The step of obtaining the conversation audio includes the following steps: Get first conversation samples on multiple platforms; Screening the first dialogue samples based on automatic speech recognition technology to select the first dialogue samples with speech recognition rates higher than a first threshold as second dialogue samples; Parameters for capturing target features are annotated in the second conversation sample to obtain the conversation-like audio.

3. The audio data processing method according to claim 2, characterized in that: The step of performing standardization processing on the conversation audio comprises the following steps: Adjusting the volume of the conversation-type audio according to the standard filter so that the volume of the conversation-type audio is within a preset range; The sampling rate of the conversation audio after the volume is adjusted is adjusted to a second threshold.

4. The audio data processing method according to claim 1, characterized in that: The intermediate features include sentiment tendency features and acoustic features; The step of extracting intermediate features from the conversation audio comprises the following steps: Extracting acoustic features from the conversational audio; Extracting emotional tendency features from the conversation audio; Determining the emotion label of the conversation audio according to the intermediate features includes: Based on the acoustic features and the emotional tendency features, a first emotional label of the conversation audio is determined.

5. The audio data processing method according to claim 4, characterized in that: The step of extracting emotional tendency features from the conversation audio comprises the following steps: An end-to-end method based on Attention Encoder-Decoder performs emotional tendency recognition on the standardized conversation audio to obtain emotional tendency features; The end-to-end method of the Attention Encoder-Decoder includes: LSTM based on the attention mechanism, or Transformer.

6. The audio data processing method according to claim 4, characterized in that: The intermediate features include style features and acoustic features; The step of extracting intermediate features from the conversation audio comprises the following steps: Converting the conversational audio into conversational text; Extracting style features from the conversation text based on the large model; Extracting acoustic features from the conversational audio; Determining the emotion label of the conversation audio according to the intermediate features includes: A second emotion annotation of the conversation audio is determined from the style features and the acoustic features based on the large model.

7. The audio data processing method according to claim 6, characterized in that: After determining the emotion label of the conversation audio according to the intermediate features, the method further includes: Comparing the first emotion annotation with the second emotion annotation to obtain a comparison result; According to the comparison result, the first emotion annotation or the second emotion annotation is selected as the final emotion annotation.

8. An audio data processing system, characterized in that: The system comprises: An audio acquisition module, used to acquire conversation audio; A standard processing module, used for performing standard processing on the conversation audio; A feature extraction module, used to extract intermediate features from the conversation audio; The emotion annotation module is used to determine the emotion annotation of the dialogue audio according to the intermediate features.

9. An electronic device, characterized in that: include: at least one controller and a memory for communicatively coupling with the at least one controller; The memory stores instructions that can be executed by the at least one controller, and the instructions are executed by the controller to enable the controller to perform the audio data processing method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the audio data processing method according to any one of claims 1 to 7.