Audio data labeling method, device, equipment and medium
By analyzing audio data using gender and accent classifiers, multiple attributes are obtained and discrete category names are generated. This solves the problem of inaccurate accent and feature labeling in existing technologies, and improves the efficiency and accuracy of audio data processing.
Patent Information
- Application Number
- CN202411494276.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-24
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2044-10-24
AI Technical Summary
Existing speech recognition and synthesis systems suffer from decreased accuracy and naturalness when processing accents from users in different regions. They also struggle to effectively annotate various features in audio data, such as pitch and speech rate, thus failing to meet the financial sector's demand for high-precision speech processing.
A gender classifier is used to determine gender attributes, an accent classifier to determine accent attributes, audio quality is analyzed to obtain signal-to-noise ratio and early-to-late reflection ratio, pitch features are analyzed to obtain speaker average pitch and pitch standard deviation, the ratio of phoneme number to audio length is calculated to determine speaking speed attributes, and these attributes are mapped to discrete category names to generate keywords.
It improves the efficiency and accuracy of audio data annotation, ensures the consistency of data annotation, simplifies the processing of complex data, and is suitable for large-scale audio data scenarios.
Smart Images

Figure CN119360881B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence and the technical field of financial technology, and in particular to an audio data labeling method, device and equipment and a storage medium. BACKGROUND
[0002] With the continuous development of financial technology, voice technology is increasingly widely used in financial services, including intelligent customer service, voice identity authentication, automated transaction instructions, etc. Existing voice recognition and synthesis technology can handle basic voice tasks, but there are obvious shortcomings when dealing with diversified audio data.
[0003] Although existing voice recognition and synthesis systems can recognize the pronunciation of standard languages, the accuracy and naturalness decrease significantly when dealing with the accents of users from different regions. Since accents have both discrete and continuous characteristics, traditional discrete label labeling methods cannot comprehensively and accurately reflect the characteristics of different accents, resulting in large differences in voice recognition effects, especially in financial scenarios, where the differences in accents of different users may affect the quality of services.
[0004] At the same time, existing technologies cannot effectively label various features in audio data, such as pitch and speech rate. These features have an important influence on voice recognition and synthesis, but due to the lack of fine processing capabilities for these detailed features in existing systems, the effects of voice processing in some complex scenarios are not ideal, which cannot meet the demand for high-precision voice processing in the financial field. SUMMARY
[0005] The main purpose of the present application is to provide an audio data labeling method, device, equipment and storage medium, which aims to solve the technical problem that existing technologies cannot accurately label and distinguish various key features such as accents, pitches and speech rates in audio data.
[0006] To achieve the above-mentioned purpose, the present application provides an audio data labeling method, comprising:
[0007] obtaining audio data;
[0008] determining a gender attribute using a gender classifier based on the audio data;
[0009] determining an accent attribute using an accent classifier based on the audio data;
[0010] analyzing the audio quality of the audio data to obtain a signal-to-noise ratio attribute and an early-to-late reflection ratio attribute;
[0011] analyzing the pitch characteristics of the audio data to obtain a speaker average pitch attribute and a pitch standard deviation attribute;
[0012] determining a speaking speed attribute by analyzing a ratio of a number of phonemes to a length of audio in the audio data;
[0013] taking the category name of the gender attribute as a keyword of the gender attribute and taking the category name of the accent attribute as a keyword of the accent attribute;
[0014] mapping the speaking speed attribute, the signal-to-noise ratio attribute, the early-to-late reflection ratio attribute, the speaker average pitch attribute and the pitch standard deviation attribute to corresponding preset discrete category names, and taking the preset discrete category names as keywords of the corresponding attributes respectively;
[0015] generating the labeling information of the audio data based on the keyword of each attribute.
[0016] Further, to achieve the above object, the present application provides an audio data labeling device, comprising:
[0017] a data acquisition module configured to acquire audio data;
[0018] a gender recognition module configured to determine a gender attribute by using a gender classifier based on the audio data;
[0019] an accent recognition module configured to determine an accent attribute by using an accent classifier based on the audio data;
[0020] an audio quality analysis module configured to analyze audio quality of the audio data and acquire a signal-to-noise ratio attribute and an early-to-late reflection ratio attribute;
[0021] a pitch feature analysis module configured to analyze pitch features of the audio data and acquire a speaker average pitch attribute and a pitch standard deviation attribute;
[0022] a speaking speed analysis module configured to determine a speaking speed attribute by analyzing a ratio of a number of phonemes to a length of audio in the audio data;
[0023] a gender and accent classification module configured to take the category name of the gender attribute as a keyword of the gender attribute and take the category name of the accent attribute as a keyword of the accent attribute;
[0024] an attribute mapping module configured to map the speaking speed attribute, the signal-to-noise ratio attribute, the early-to-late reflection ratio attribute, the speaker average pitch attribute and the pitch standard deviation attribute to corresponding preset discrete category names, and take the preset discrete category names as keywords of the corresponding attributes respectively;
[0025] a labeling generation module configured to generate the labeling information of the audio data based on the keyword of each attribute.
[0026] Further, to achieve the above object, the present application also provides a computer device, comprising a memory, a processor, and an audio data labeling program stored in the memory and executable on the processor, wherein the audio data labeling program implements the steps of the audio data labeling method when executed by the processor.
[0027] Further, to achieve the above object, the present application also provides a computer readable storage medium, wherein the storage medium stores an audio data labeling program, and the audio data labeling program implements the steps of the audio data labeling method when executed by a processor.
[0028] Beneficial effects: The present application relates to the fields of artificial intelligence and financial technology, and discloses an audio data labeling method. The method comprises the following steps: obtaining audio data, determining gender attributes by using a gender classifier, determining accent attributes by using an accent classifier, analyzing audio quality to obtain signal-to-noise ratio and early-to-late reflection ratio, analyzing pitch characteristics to obtain average pitch and pitch standard deviation of a speaker, and calculating the ratio of phoneme quantity to audio length to determine speaking speed attributes. Corresponding keywords are generated for each attribute, and labeling information of the audio data is generated based on the keywords. The present application automatically analyzes various attributes (such as gender, accent, and pitch) of the audio data, generates corresponding keywords, and significantly improves the efficiency and accuracy of audio data labeling. Discretization processing converts continuous attributes (such as speaking speed and signal-to-noise ratio) into simple category names, simplifies the processing of complex data, and makes the system more efficient in classification, retrieval, and scalability. At the same time, the automated process reduces manual operations, ensures the consistency and accuracy of data labeling, and is particularly suitable for scenarios that require large-scale processing of audio data. BRIEF DESCRIPTION OF DRAWINGS
[0029] The present application will be further described below with reference to the accompanying drawings and embodiments. In the drawings:
[0030] Figure 1 An application environment diagram of the audio data labeling method in an embodiment of the present application;
[0031] Figure 2 A flow diagram of the audio data labeling method in an embodiment of the present application;
[0032] Figure 3 A functional module diagram of a preferred embodiment of the audio data labeling device of the present application;
[0033] Figure 4 A structure diagram of a computer device in an embodiment of the present application;
[0034] Figure 5 Another structure diagram of a computer device in an embodiment of the present application. DETAILED DESCRIPTION
[0035] It should be understood that the specific embodiments described herein are merely exemplary and not intended to limit the application.
[0036] The audio data labeling method provided by the embodiments of the present application can be applied in an application environment as shown in Figure 1 , wherein a user end communicates with a service end through a network. The service end can obtain audio data through the user end, determine a gender attribute by using a gender classifier, determine a dialect attribute by using a dialect classifier, analyze audio quality to obtain a signal-to-noise ratio and an early-to-late reflection ratio, analyze a pitch feature to obtain an average pitch and a pitch standard deviation of a speaker, and calculate a ratio of a phoneme quantity to an audio length to determine a speaking speed attribute. Corresponding keywords are generated for each attribute, and labeling information of the audio data is generated based on the keywords. The present application can accurately label various key features in the audio data by analyzing and labeling various attributes in the audio data, ensure comprehensive expression of each attribute, and improve the processing efficiency and accuracy of the speech data. The user end can be, but is not limited to, various personal computers, notebook computers, smart phones, tablet computers and portable wearable devices. The service end can be implemented by an independent server or a server cluster composed of multiple servers. The present application will be described in detail through specific embodiments.
[0037] Please refer to Figure 2 , Figure 2 for a flowchart of an embodiment of the audio data labeling method provided by the present application. It should be noted that although a logical order is shown in the flowchart, in some cases, the steps shown or described herein can be performed in an order different from that shown herein.
[0038] As shown in Figure 2 , the audio data labeling method provided by the present application includes the following steps:
[0039] S10, obtaining audio data;
[0040] In the present embodiment, obtaining audio data means collecting speech signals or audio samples from a specified data source or device. This can be achieved through multiple channels, such as direct collection by a recording device, extraction from an audio database, or downloading existing audio data from a cloud service platform. A microphone or other recording device can be used to collect audio, or an API interface can be used to call audio files from a cloud service. The collected audio data can be saved in formats such as WAV, MP3, etc., and needs to be preprocessed, such as noise reduction and background sound removal, to ensure that the quality of the audio data is suitable for subsequent analysis. Audio files can also be extracted from existing audio databases or public data sets (such as LibriVox).
[0041] In one specific implementation, the transaction dialogue between the customer and the customer service is collected through a dedicated telephone line, and the dialogue content is automatically recorded using the built-in recording function of the device. The system automatically identifies key financial transaction words in the dialogue, such as "transfer" and "remittance", and automatically marks these parts for subsequent processing when collecting.
[0042] Example: In the identity verification process of a certain financial institution, the system ensures the legality of the transaction by recognizing and analyzing the customer's voice data. By obtaining the customer's voice data, the system can use a gender classifier to determine the gender attribute of the customer, ensuring consistency with their registered information; use an accent classifier to identify the customer's accent and determine whether the customer's voice is abnormal; obtain signal-to-noise ratio, pitch, speech rate, and other features to ensure the clarity and naturalness of the voice.
[0043] Through the above steps, the system can ensure that the audio data obtained is extensive and of controllable quality, adapting to the needs of different application scenarios. Not only does it improve the usability of audio data, but it also ensures the accuracy and efficiency of subsequent processing (such as audio feature analysis and labeling). In the financial field, it can achieve real-time collection of customer voice data or quickly obtain customer voice information from historical data, thereby improving the quality of voice interaction services.
[0044] S20, based on the audio data, using a gender classifier to determine a gender attribute;
[0045] In this embodiment, by processing the audio data, a specific gender classifier is used to determine the gender attribute of the speaker. The gender classifier is a model based on machine learning or deep learning, which is specifically used to analyze the features (such as frequency, pitch, timbre, etc.) in the audio and predict the gender of the speaker based on these features. First, extract the audio features that can distinguish gender from the audio data, including pitch, spectral features, and the volatility of the voice. Use a pre-trained gender classifier model that can classify the gender of the speaker based on the audio features. According to the results of the gender classifier, mark the speaker as "male" or "female" and generate the corresponding gender attribute label.
[0046] By preprocessing the audio data, the system extracts relevant frequency information and pitch features, especially the main frequency and formant of the audio. Based on these features, the gender classifier can identify the gender difference of the speaker. Use the trained gender classification model to input the extracted audio features into the model. Common gender classification models include models based on convolutional neural networks (CNN) or models based on naive Bayes classifiers. The classifier processes the input audio features and outputs a label of "male" or "female", which is used as the gender attribute.
[0047] In one specific implementation, frequency components in the audio, especially pitch and formants, are extracted through spectral analysis techniques, and a CNN model is used for gender classification. The classifier can adapt to pronunciation differences for audio data with different accents, ensuring the accuracy of classification, especially in different tones and speech rates.
[0048] Example: In a smart financial assistant application, the system interacts with the customer through voice, determines the gender of the customer, and provides customized financial advice and financial product recommendations. For example, the system can recommend appropriate financial products based on the preferences of customers of different genders. The customer communicates with the smart financial assistant system through voice, and the system collects voice data in real time. The gender classifier is used to analyze the customer's voice and identify the customer's gender. Based on the identified gender attribute, the system automatically provides personalized financial advice and financial product recommendations for the customer.
[0049] Through the above steps, the system can quickly and accurately identify the gender attribute of the speaker, providing accurate basic data for subsequent voice processing.
[0050] S30, based on the audio data, using an accent classifier to determine an accent attribute;
[0051] In this embodiment, the accent attribute of the speaker is determined using an accent classifier based on audio data. The accent classifier can identify different regional dialects or accents, such as Mandarin, Cantonese, Sichuanese, and Minnanese. By extracting audio features related to accents, the classifier can effectively distinguish the pronunciation differences of different places and determine the accent attribute of the speaker. The system extracts features from audio data that can distinguish different accents, including pitch, tone, speech rate, formant frequency, and prosody. These features have significant differences in pronunciation in different places. Using a specially trained accent classifier model, the model is trained on various local dialects and accents, and can identify common accents such as Mandarin, Cantonese, Minnanese, Northeastern Mandarin, and Sichuanese. According to the identification result of the accent classifier, the corresponding accent attribute label is generated, which identifies the specific accent characteristics of the speaker.
[0052] In one specific implementation, the system extracts pitch and spectral features from audio data, and uses an LSTM (Long Short-Term Memory) based accent classifier model to identify the difference between Mandarin and local dialects (such as Cantonese and Minnanese). In the classifier model, the evaluation of the standard degree of Mandarin is added to judge whether the Mandarin has local accent, such as "Mandarin with Guangdong accent" or "Mandarin with Sichuan accent".
[0053] By using the accent classifier, the system can accurately distinguish the accent attributes of the speaker, and has high precision in the recognition of Mandarin and dialects. It can help financial institutions better understand the language characteristics of customers and improve the performance of voice recognition systems in multi- regional and multi- accent scenarios.
[0054] S40, analyze the audio quality of the audio data, obtain signal-to-noise ratio attribute and early and late reflection ratio attribute;
[0055] In this embodiment, by analyzing the audio quality attributes in the audio data, the signal-to-noise ratio (SNR) and the early and late reflection ratio are obtained. These two attributes directly affect the intelligibility and recognizability of the audio. The signal-to-noise ratio is used to measure the intelligibility of the audio signal relative to the background noise, while the early and late reflection ratio is used to describe the reverberation characteristics in the audio, especially the proportion of early reflection and late reflection, which plays an important role in determining the audibility and spatial sense of the audio.
[0056] The signal-to-noise ratio (SNR) refers to the ratio of signal strength to background noise intensity, reflecting the intelligibility of the speech signal relative to the noise. By separating the speech and background noise parts, the system calculates the signal-to-noise ratio.
[0057] The early and late reflection ratio reflects the reverberation characteristics in the audio. Early reflection is a short reflection after direct sound, and late reflection is a long-delayed reverberation reflection. By measuring the energy of different reflection sounds in the audio, the spatial sense and reverberation degree of the audio are determined.
[0058] Using spectral analysis or waveform separation technology, the effective speech part and background noise part in the audio data are separated. By calculating the power of the effective speech signal and the power of the background noise, and taking the ratio of the two, the signal-to-noise ratio attribute is obtained. According to the predetermined signal-to-noise ratio threshold range, the audio data is classified into different quality levels (such as "high quality", "medium quality", "low quality").
[0059] Through time-frequency analysis technology, the time points and energy of early reflection and late reflection in the audio are identified. The energy of early reflection sound and late reflection sound is measured respectively, and the energy ratio of the two is taken as the early and late reflection ratio attribute. According to the preset ratio range, the audio is classified into different reverberation levels (such as "strong reverberation", "moderate reverberation", "weak reverberation").
[0060] Through the analysis of the signal-to-noise ratio and the early and late reflection ratio of the audio data, the system can accurately judge the intelligibility and spatial sense of the audio, thereby improving the processing quality and recognition accuracy of the audio data.
[0061] S50, analyze the pitch characteristics of the audio data, obtain the speaker average pitch attribute and the pitch standard deviation attribute;
[0062] In this embodiment, the average pitch attribute and the fluctuation of the pitch (standard deviation attribute) of the speaker are obtained by analyzing the pitch characteristics in the audio data. The pitch is an important feature of speech, especially for the identification of the speaker's emotion and identity. The average pitch is used to reflect the basic tone of the speaker, while the pitch standard deviation is used to measure the fluctuation of the speaker's voice. By extracting all the pitch values in the audio data, the average value of these values is calculated to obtain the basic pitch attribute of the speaker. By analyzing the amplitude and distribution of the pitch values, the standard deviation of the pitch values is calculated to reflect the fluctuation of the speaker's voice.
[0063] The audio data is processed using a pitch detection tool to extract the pitch values (such as F0 frequency) in the audio, which reflect the basic frequency of the speech. The system records the changes of the pitch according to the time axis to form a pitch sequence. All the pitch values in the pitch sequence are summarized to calculate the average value of these pitch values as the average pitch attribute of the speaker.
[0064] The distribution of the pitch values is analyzed to calculate the deviation of each pitch value from the average pitch, and then the standard deviation is calculated to obtain the range and characteristics of the pitch fluctuation.
[0065] In one specific implementation, the system identifies the periodic signal in the audio through the autocorrelation function to extract the pitch value. The autocorrelation function can suppress the noise part in the audio signal, so it is particularly suitable for use in environments with high background noise, such as telephone calls or voice message systems. The system combines the formant analysis and autocorrelation pitch detection to further improve the accuracy of pitch detection in high noise environments. It is suitable for audio processing in financial institution telephone customer service systems, so that the system can still accurately analyze the voice characteristics of the customer in the case of high background noise.
[0066] Example: In a voice transaction system of a certain financial institution, customers make large amount transfer instructions through the telephone. The system analyzes the pitch characteristics in the customer's voice to extract the average pitch and pitch standard deviation. If it is detected that the customer's pitch fluctuation is large and does not match the usual voice data, the system will automatically determine that the customer is in a state of tension or anxiety, triggering an additional identity verification step or requiring manual customer service intervention to confirm the transaction instruction. This process not only improves the security of the transaction, but also prevents potential fraudulent behavior.
[0067] Through the above steps, the pitch characteristics of the speaker can be accurately analyzed, and the average pitch and pitch standard deviation, two key attributes, can be extracted. This helps better understand the identity characteristics, emotional fluctuations, and speech patterns of the speaker in speech analysis. In the financial field, pitch analysis can be used for customer identity verification and emotional analysis. For example, in telephone transactions, the system can analyze the pitch fluctuations of the customer to determine whether they are in a state of tension or anxiety, thereby enhancing risk management and customer service responsiveness.
[0068] S60, analyzing the ratio of the number of phonemes to the length of the audio in the audio data to determine the speaking speed attribute;
[0069] In this embodiment, the speaking speed attribute of the speaker is calculated by analyzing the ratio of the number of phonemes to the total length of the audio in the audio data. Phonemes are the smallest units of speech, while speaking speed reflects the amount of speech output by the speaker within a certain time. By calculating the ratio of the number of phonemes to the length of the audio, the speaker's speech speed can be accurately measured and classified into different speech speed levels (such as "fast", "medium", "slow"). Through speech recognition technology, the system can identify each phoneme in the audio and count the total number of phonemes. The length of the audio is the total duration of the audio, i.e. the actual speaking time of the speaker in a particular audio. Through the ratio calculation of the number of phonemes to the length of the audio, the system obtains the quantitative result of the speaking speed, and further determines the speech speed attribute of the speaker.
[0070] Using speech recognition technology to process the audio data, extract the phonemes in each utterance, and count the number of phonemes in the entire audio. Common techniques include HMM (Hidden Markov Model) and deep neural network (DNN) models. The system accurately calculates the total duration of the audio by analyzing the length of the audio file, which can be done by directly reading the metadata of the audio file. The system divides the number of phonemes by the length of the audio to obtain the number of phoneme utterances per second (phonemes / second), which is the speech speed attribute of the speaker. According to the calculated speech speed value, the system classifies it into different speech speed levels, such as "fast" (high number of phonemes per second), "medium" (moderate phoneme utterance), and "slow" (fewer phoneme utterances).
[0071] In one specific implementation, the audio data is analyzed using a Hidden Markov Model (HMM) to extract the phoneme sequence in the audio. The system captures the dynamic changes in the speech signal through the HMM model, counts the frequency of each phoneme, and calculates the speaking speed in combination with the audio length. The system can adjust the parameters of the HMM model according to the context of different audios to ensure accurate phoneme recognition in noisy environments, especially in the application of financial industry telephone calls or conference recordings.
[0072] By analyzing the ratio of the number of phonemes to the length of the audio, more accurate speech rate recognition and analysis can be achieved in the financial field, helping financial institutions improve overall efficiency and security in customer service, transaction security, and risk control.
[0073] S70, the category name of the gender attribute is used as the keyword of the gender attribute, and the category name of the accent attribute is used as the keyword of the accent attribute;
[0074] In this embodiment, the gender attribute is determined by analyzing the speech signal in the audio data to determine the gender of the speaker. Common gender categories include "male" and "female". The category name is the text label of the classification result, i.e. "male" or "female". Such names are usually generated by a gender classifier based on speech features such as pitch, frequency, etc. The keyword is the generated category name used as an identifier for subsequent data processing. The category name of the gender attribute is directly used as the keyword of the gender attribute, i.e. "Male" or "Female" can be used as a label to mark the gender information of the audio data.
[0075] The accent attribute is determined by analyzing specific features of pronunciation in the audio data, such as speech patterns, pronunciation habits, rhythm, and tone, etc.
[0076] The accent attribute usually includes the differences between various regional Mandarin and local dialects. The main categories can be divided according to specific application requirements:
[0077] Standard Mandarin: Official standard Mandarin with clear pronunciation and no obvious local accent.
[0078] Mandarin with dialect: Mandarin mixed with local accents, such as northern accent, southern accent, etc.
[0079] Specific dialect accent: such as Northeastern accent, Cantonese accent, Minnan accent, Sichuan accent, Hunan accent, Hakka accent, etc.
[0080] Category name: The classifier outputs the category name by analyzing the pronunciation features in the audio data, such as "standard Mandarin", "Mandarin with Northeastern accent", or "Cantonese".
[0081] The category name is the result of the classifier's judgment based on the pronunciation features of the audio data, such as "Mandarin", "Mandarin with Northeastern accent", or "Cantonese". The keyword is the category name of the accent attribute as an identifier for this audio data. By generating keywords such as "Mandarin" or "Cantonese", classification and retrieval of audio data can be achieved.
[0082] For example, the audio data is input into the accent classifier, which extracts pronunciation features such as syllables, tones, pronunciation rhythm, etc. from the audio for identifying the accent type. Using a trained accent classification model (such as a deep learning-based model), by analyzing the pronunciation patterns of the audio, it is determined whether the speaker is using standard Mandarin, Mandarin with local accent, or a certain specific dialect such as Cantonese, Minnan, etc. According to the classification result, the accent category name is generated, such as "Mandarin", "Mandarin with Northeastern accent" or "Cantonese". The generated accent category name is used as the keyword of the accent attribute to label the audio data.
[0083] By automatically classifying and labeling the gender and accent attributes of the audio data, supporting Mandarin and dialects in various regions, and generating corresponding keywords, the accuracy and efficiency of data labeling are significantly improved, the manual cost is reduced, and the retrieval and management capabilities of audio data are enhanced, which is suitable for scenarios that require processing of large amounts of audio data.
[0084] S80, mapping the speaking speed attribute, signal-to-noise ratio attribute, early and late reflection ratio attribute, speaker average pitch attribute and pitch standard deviation attribute to corresponding preset discrete category names, and taking the preset discrete category names as keywords of the corresponding attributes respectively;
[0085] In this embodiment, the speaking speed attribute of the audio data is analyzed, and according to the number of phonemes per minute, it is mapped to a preset discrete category, such as "fast", "medium" and "slow". By counting the number of phonemes in a unit of time, and according to the preset threshold, the speaking speed is divided into "fast", "medium" or "slow". The corresponding generated keywords are "fast", "medium" or "slow".
[0086] The signal-to-noise ratio attribute measures the proportion of audio signal and background noise, and according to the value of signal-to-noise ratio, it is mapped to a preset discrete category, such as "high", "medium" and "low". By performing spectral analysis on the audio data, the intensities of signal and noise are calculated, and according to the preset signal-to-noise ratio interval, it is divided into "high", "medium" and "low". The corresponding generated keywords are "high signal-to-noise ratio", "medium signal-to-noise ratio" or "low signal-to-noise ratio".
[0087] The early and late reflection ratio reflects the echo characteristics of the environment, and according to the size of this ratio, it is mapped to a preset discrete category, such as "strong reflection", "medium reflection" and "weak reflection". By analyzing the acoustic reflection characteristics in the audio data, the proportion of early and late reflections is calculated, and according to the set standard, it is divided into "strong reflection", "medium reflection" and "weak reflection". The generated keywords are "strong reflection", "medium reflection" or "weak reflection" respectively.
[0088] The average pitch of the speaker is analyzed, and according to the range of pitch values, it is mapped to "high pitch", "medium pitch", and "low pitch". By analyzing the fundamental frequency (F0), the pitch value is calculated, and according to the pitch interval, it is divided into "high pitch", "medium pitch", or "low pitch". The generated keywords are "high pitch", "medium pitch", or "low pitch", respectively.
[0089] The pitch standard deviation reflects the fluctuation of the pitch, and according to the fluctuation amplitude, it is mapped to "large fluctuation", "medium fluctuation", and "small fluctuation". By calculating the standard deviation of the fundamental frequency in the audio, the fluctuation amplitude of the pitch is evaluated, and according to the preset standard, it is divided into "large fluctuation", "medium fluctuation", or "small fluctuation". The generated keywords are "large fluctuation", "medium fluctuation", or "small fluctuation", respectively.
[0090] Discretization can simplify complex continuous data, making processing and analysis more efficient. For example, continuous speaking speed, pitch, or signal-to-noise ratio values may have very large value ranges, and directly processing these continuous data may be computationally complex. Mapping them to discrete categories such as "fast", "medium", etc. can reduce processing complexity and facilitate fast calculation and classification by the system.
[0091] For many users or application scenarios, discretized data is easier to understand. For example, "fast", "slow" is more intuitive than specific numerical values (such as 200 words per minute). Especially in application scenarios that need to face business decisions, the discretized results are more in line with the understanding and needs of users, helping them to quickly respond. Mapping continuous attributes to discrete categories can facilitate classification and indexing by the system. For example, when processing a large amount of audio data, classifying audio according to discrete categories such as "fast", "slow", "high pitch", "low pitch", etc. can greatly improve the efficiency of data retrieval, helping the system quickly find audio files that meet specific conditions.
[0092] Discretization of continuous data can reduce the impact of noise or extreme values. In some cases, continuous data may be affected by external factors and exhibit abnormal fluctuations, and discretization can reduce the impact of these fluctuations on subsequent processing or analysis, improving system stability.
[0093] By discretizing continuous attributes into preset categories, the data processing process is simplified, the efficiency and stability of the system are improved, and the interpretability and ease of use of the data are enhanced. The generated keywords help improve the retrieval and classification capabilities of audio data, facilitate the application of business rules, reduce the interference of noise and data fluctuations, and are particularly suitable for large-scale audio data labeling and management scenarios.
[0094] S90, generating labeling information of the audio data based on the keywords of each attribute.
[0095] In this embodiment, the system combines the keywords generated for each attribute together to form structured annotation information. For example, "male, Cantonese, fast, high pitch, high SNR" represents a segment of audio data with a male speaker, speaking Cantonese, fast speech rate, high pitch, and clear audio quality.
[0096] In one specific implementation, in the process of generating keywords, the system combines the contextual information in the audio to ensure that the keywords not only reflect a single attribute, but also enhance the accuracy of the annotation information through context. For example, more accurate annotations are generated by combining the emotional information of the speaker when analyzing the speech rate. By introducing speech emotion recognition technology, the system can reflect the emotional changes and tone characteristics of the speaker when generating keywords, further improving the comprehensiveness of the annotation information.
[0097] By generating keywords for each attribute and constructing structured annotation information for audio data, the system can greatly improve the usability and operability of audio data. This annotation method provides strong support for subsequent speech analysis, retrieval, and classification. In the financial field, the system generates detailed annotation information for customers' voice data, not only improving the accuracy of speech recognition, but also providing more intelligent support for personalized customer service, voice transaction confirmation, and other scenarios.
[0098] The present application relates to the field of artificial intelligence technology and financial technology, and discloses an audio data annotation method. By obtaining audio data, a gender classifier is used to determine the gender attribute, an accent classifier is used to determine the accent attribute, the audio quality is analyzed to obtain the signal-to-noise ratio and early-to-late reflection ratio, the pitch characteristics are analyzed to obtain the speaker's average pitch and pitch standard deviation, and the ratio of the number of phonemes to the length of the audio is calculated to determine the speaking speed attribute. For each attribute, a corresponding keyword is generated, and annotation information for the audio data is generated based on these keywords. The present application can accurately annotate various key features in the audio data by analyzing and annotating various attributes in the audio data, ensuring comprehensive expression of each attribute, and improving the processing efficiency and accuracy of voice data.
[0099] In one embodiment, the above S80 comprises:
[0100] S801, mapping the speaking speed attribute to discrete category names including fast, medium, and slow;
[0101] S802, mapping the signal-to-noise ratio attribute to discrete category names including high, medium, and low;
[0102] S803, mapping the early-to-late reflection ratio attribute to discrete category names including strong reflection, medium reflection, and weak reflection;
[0103] S804, mapping the speaker average pitch attribute to discrete category names including high, medium and low;
[0104] S805, mapping the pitch standard deviation attribute to discrete category names including high, medium and low;
[0105] S806, mapping the discrete category names as keywords of corresponding attributes.
[0106] In this embodiment, the speaking speed attribute in the audio data is analyzed by calculating the number of phonemes per unit time, which is mapped to discrete categories "fast", "medium" and "slow". The number of phonemes in the audio data is counted by a speech analysis algorithm, and based on the number of phonemes per minute, the threshold values of the three discrete categories "fast", "medium" and "slow" are set. For example, more than 200 phonemes per minute is classified as "fast", between 100 and 200 is "medium", and less than 100 is "slow". According to the calculation result, the speaking speed is mapped to the three discrete categories, and finally the corresponding keywords "fast", "medium" or "slow" are generated.
[0107] The signal-to-noise ratio in the audio data is analyzed to measure the ratio between signal strength and noise, and the signal-to-noise ratio is mapped to three discrete categories "high", "medium" and "low". By using a signal-to-noise ratio analysis tool, the strength of the effective signal and background noise in the audio is extracted, and the signal-to-noise ratio is calculated. According to the preset threshold value, for example, the signal-to-noise ratio greater than 50 dB is "high", between 30 and 50 dB is "medium", and less than 30 dB is "low". Based on this analysis result, the signal-to-noise ratio attribute is mapped to "high", "medium" and "low", and the corresponding keywords "high signal-to-noise ratio", "medium signal-to-noise ratio" or "low signal-to-noise ratio" are generated.
[0108] The early and late reflection ratio attribute in the audio data is analyzed, which is used to measure the reflection echo characteristics of the environment, and is divided into "strong reflection", "medium reflection" and "weak reflection". The acoustic analysis tool is used to calculate the ratio value of early and late reflection sound energy. According to the set threshold value, the reflection ratio value higher than a certain value is defined as "strong reflection", within a certain range is "medium reflection", and lower than a certain value is "weak reflection". According to the calculation result, the corresponding keywords "strong reflection", "medium reflection" or "weak reflection" are generated.
[0109] By analyzing the fundamental frequency in the audio data, the average pitch of the speaker is calculated and mapped to the discrete categories of "high pitch", "medium pitch", and "low pitch". Through the fundamental frequency extraction algorithm in the speech signal, the average pitch value of the speaker during the call is calculated and compared with the preset fundamental frequency range. For example, the fundamental frequency higher than 300Hz is "high pitch", between 150 to 300Hz is "medium pitch", and lower than 150Hz is "low pitch". The corresponding keywords "high pitch", "medium pitch", or "low pitch" are generated.
[0110] The pitch fluctuation in the audio data is analyzed, and the standard deviation is used to measure the fluctuation amplitude of the pitch, which is divided into "large fluctuation", "medium fluctuation", and "small fluctuation". The standard deviation of the pitch is calculated to measure the change amplitude of the pitch. According to the standard deviation value, it is divided into three discrete categories, such as greater than a certain standard deviation value for "large fluctuation", within a certain range for "medium fluctuation", and less than a certain standard deviation value for "small fluctuation". The corresponding keywords "large fluctuation", "medium fluctuation", or "small fluctuation" are generated.
[0111] After the discretization processing of each attribute, the result is mapped to the corresponding keyword. These keywords are used as the attribute labeling information of the audio data. The discrete category name after mapping each attribute is directly recorded as a keyword in the audio data. For example, the speaking speed attribute of a certain audio is mapped to "fast", and the signal-to-noise ratio attribute is mapped to "high signal-to-noise ratio", so the final keywords of the audio data will include "fast", "high signal-to-noise ratio", etc.
[0112] Example: In the financial field, customer service centers usually collect a large amount of customer call records. Through this technical solution, the system can automatically analyze the speaking speed, pitch, signal-to-noise ratio, etc. in the call record, and generate corresponding keyword labels. For example, the call record of a certain customer shows that the speaking speed is fast, the pitch is high, and the signal-to-noise ratio is low, and the system will automatically generate "fast", "high pitch", and "low signal-to-noise ratio" as keywords. This labeling method not only improves the data management efficiency, but also helps financial institutions identify the emotions or urgency of customers, so as to provide personalized services and improve customer satisfaction.
[0113] This embodiment simplifies the processing process of continuous attributes by mapping various attributes of audio data to discrete category names and generating Chinese keywords, improving the efficiency and accuracy of audio data labeling. Discretization processing makes the classification and retrieval of audio data more efficient, and the generated Chinese keywords are intuitive and easy to understand, making it easy for users to quickly understand the audio features. The automatic processing process reduces the workload of manual labeling and improves data consistency, especially suitable for large-scale audio data management scenarios.
[0114] In one embodiment, in S40, analyzing the audio quality of the audio data to obtain a signal-to-noise ratio attribute comprises:
[0115] S401, using a spectrum analysis module to separate the audio data into an effective speech part and a background noise part;
[0116] S402, respectively measuring the power of the effective speech part and the background noise part;
[0117] S403, analyzing the ratio of the power of the effective speech part to the power of the background noise part to determine a signal-to-noise ratio attribute value;
[0118] S404, according to a predetermined signal-to-noise ratio threshold range, classifying the signal-to-noise ratio attribute value into a corresponding audio quality level.
[0119] In this embodiment, the audio data is separated into an effective speech part and a background noise part by a spectrum analysis module, the power of each part is measured, and finally the signal-to-noise ratio (SNR) attribute value is determined by calculating the ratio of the two. According to the preset threshold range, the signal-to-noise ratio value is classified into different audio quality levels, helping the system to judge the intelligibility and usability of the audio.
[0120] The system first uses a spectrum analysis module to perform frequency domain processing on the audio data, decomposes the audio signal into different frequency components, and analyzes the energy distribution of each frequency band. Through the analysis, the system can identify the effective speech part (mainly concentrated in a certain frequency range) and the background noise part (usually low-frequency or high-frequency noise).
[0121] The system calculates the power of the effective speech part (i.e. the signal energy in the main speech frequency band) and the power of the background noise part. Power is an important physical quantity that reflects signal strength, which is usually estimated by the energy density function of the spectrum.
[0122] The system determines the signal-to-noise ratio attribute value by calculating the ratio of the effective speech power to the background noise power. The signal-to-noise ratio is usually expressed in decibels (dB) and is used to measure the intelligibility of the speech signal. If the signal-to-noise ratio is high, the speech signal is clear; if the signal-to-noise ratio is low, the background noise has a greater impact.
[0123] The system classifies the calculated signal-to-noise ratio attribute value into different audio quality levels according to a predetermined signal-to-noise ratio threshold range. For example:
[0124] High signal-to-noise ratio (> 30 dB): indicates high audio quality and low background noise.
[0125] Medium signal-to-noise ratio (10-30 dB): indicates average audio quality, and the background noise has a certain impact on the speech.
[0126] Low SNR (<10 dB): Indicates poor audio quality with high background noise.
[0127] The system performs spectral analysis on the input audio data using Short-Time Fourier Transform (STFT). By converting the audio signal from time domain to frequency domain, the system can clearly identify the energy distribution of different frequency bands and separate the effective speech part and background noise part according to the frequency characteristics unique to speech.
[0128] The system measures the power of the effective speech part and the background noise part respectively. The power calculation of the effective speech part mainly focuses on the middle frequency band, while the power of the background noise is usually located in the low frequency band or high frequency band. Through the Power Spectral Density (PSD) function, the system can accurately estimate the power value of each frequency band.
[0129] Through the ratio of effective speech power and background noise power, the system calculates the SNR attribute value.
[0130] The system compares the SNR attribute value with the preset threshold range to determine the level of audio quality. If the SNR is high, the system will mark the audio as high quality; if the SNR is low, it will be marked as low quality.
[0131] In one specific embodiment, the spectrum of the audio is estimated by an Autoregressive (AR) model, which can more accurately identify the effective speech part in audio containing strong noise, further improving the calculation accuracy of SNR. By introducing the autoregressive model, random background noise can be effectively suppressed, ensuring that even in a low SNR environment, the speech signal can still be accurately separated.
[0132] This embodiment can effectively evaluate the quality of audio data through spectral analysis, power measurement and SNR calculation, and provide high-quality input for subsequent speech recognition and processing. In the financial field, especially in scenarios such as telephone transactions and voice recognition identity verification, accurate calculation of SNR can help the system judge the reliability of audio data and improve the accuracy and security of speech recognition.
[0133] In one embodiment, the above S40, analyzing the audio quality of the audio data, obtaining the early and late reflection ratio attribute, comprises:
[0134] S405, using time-frequency analysis technology to identify early and late reflections in the audio data;
[0135] S406, respectively measuring the energy of the early and late reflections;
[0136] S407, analyzing the ratio of early reflection energy to late reflection energy to determine the early and late reflection ratio attribute value;
[0137] S408, according to the predetermined early and late reflection ratio threshold range, the early and late reflection ratio attribute value is classified into the corresponding reverberation degree level.
[0138] In this embodiment, the system identifies early reflections and late reflections in the audio data through time-frequency analysis techniques and measures their respective energies, and determines the early and late reflection ratio attribute through the energy ratio of the two. Early reflections and late reflections are two key components of audio reverberation, which represent the directly and indirectly reflected sounds after the sound source emits, respectively, and affect the spatial sense and clarity of the audio. By analyzing the early and late reflection ratio, the system can judge the reverberation degree in the audio data and classify it into different levels according to the reverberation degree.
[0139] The system uses time-frequency analysis techniques to convert the audio data into the time-frequency domain, and identifies the early reflections (usually the reflected sound within a short time after the voice is emitted) and the late reflections (the sound that reaches the ear after multiple reflections) in the audio by analyzing the changes in frequency and time. Early reflections generally last within 50 milliseconds, while late reflections are longer.
[0140] The system measures the power of early reflections and late reflections in the audio separately. By analyzing the energy density in the time-frequency graph, the system can estimate the energy size of early reflections and late reflections.
[0141] The system determines the early and late reflection ratio attribute by calculating the ratio of early reflection energy to late reflection energy. This ratio reflects the spatial sense and reverberation intensity of the audio. A high ratio indicates strong early reflection energy and weak reverberation, while a low ratio indicates strong late reflection energy and heavy reverberation.
[0142] The system classifies the calculated early and late reflection ratio into different reverberation degree levels according to the preset early and late reflection ratio threshold range. For example:
[0143] Light reverberation (high ratio): early reflection energy is strong, voice is clear.
[0144] Moderate reverberation (medium ratio): early and late reflection energy is balanced, voice has some reverberation but is still understandable.
[0145] Strong reverberation (low ratio): late reflection energy is strong, reverberation is severe, which may affect voice understanding.
[0146] The system uses time-frequency analysis techniques such as short-time Fourier transform (STFT) or wavelet transform on the audio data to convert the audio signal from the time domain to the time-frequency domain. By analyzing the changes in the frequency spectrum of the audio signal over time, the system can identify the time periods of early reflections and late reflections.
[0147] The time-frequency analysis obtains the frequency range and time period of early reflections and late reflections, and the system measures the energy of the two reflection regions respectively. The energy of early reflections is mainly concentrated in the high frequency band in a short time, and the energy distribution of late reflections is more extensive and the duration is longer. The system divides the total energy of early reflections by the total energy of late reflections to obtain the early-late reflection ratio. The system compares the calculated reflection ratio with the predetermined threshold range, and classifies the audio data into different levels such as slight reverberation, moderate reverberation or strong reverberation according to the result. For telephone calls or conference recordings in financial applications, audio with low reverberation is marked as high quality.
[0148] The embodiment analyzes the early reflection and late reflection energy ratio in the audio data, and the system can accurately evaluate the reverberation degree of the audio, so as to judge the intelligibility and spatial sense of the audio.
[0149] In one embodiment, the above 50 comprises:
[0150] S501, using a pitch detection tool to extract a pitch value from the audio data, and recording the pitch value to form a time-varying pitch sequence;
[0151] S502, summarizing all extracted pitch values in the audio data, analyzing the average value of all pitch values, and determining the average pitch attribute value of the speaker;
[0152] S503, analyzing the distribution of the pitch value and calculating the standard deviation of the pitch value to determine the pitch standard deviation attribute value;
[0153] S504, according to the predetermined pitch range, the average pitch attribute value is classified into corresponding pitch levels;
[0154] S505, according to the predetermined standard deviation range, the pitch standard deviation attribute value is classified into corresponding fluctuation degree levels.
[0155] In this embodiment, the pitch feature in the audio data is extracted by using a pitch detection tool, and then the average pitch and the pitch fluctuation degree of the speaker are analyzed. Pitch is one of the important attributes of speech, and by analyzing it, important references can be provided for subsequent speech recognition, emotion recognition and other scenes. The specific steps include extraction of pitch value, calculation of average pitch, calculation of pitch standard deviation, and classification of these values according to the preset standard.
[0156] The system uses a pitch detection tool (such as autocorrelation method, HPS (harmonic perceptive spectrum) method or deep learning model) to extract the pitch value from the audio data. Each pitch value reflects the fundamental frequency (F0) of the speaker at a specific time. The system will extract a series of pitch values in the entire audio signal to form a time-varying pitch sequence.
[0157] The system aggregates the extracted pitch sequence and calculates the mean of all pitch values, obtaining the average pitch attribute of the speaker. This value reflects the fundamental pitch used by the speaker throughout the speech and is an important indicator for identifying the speaker's characteristics.
[0158] The system analyzes the distribution of the pitch sequence and calculates the standard deviation of the pitch values. The pitch standard deviation reflects the degree of fluctuation in the speaker's speech. The larger the standard deviation, the greater the change in pitch and the more obvious the fluctuation. The smaller the standard deviation, the more stable the speech and the less the change in pitch.
[0159] The system classifies the average pitch attribute value into different pitch levels according to the pre-set pitch range.
[0160] For example:
[0161] Low pitch: below a certain pre-set frequency threshold.
[0162] Medium pitch: within a certain pre-set intermediate frequency range.
[0163] High pitch: above a certain frequency threshold.
[0164] Pitch standard deviation classification:
[0165] The system classifies the degree of pitch fluctuation according to the size of the pitch standard deviation. For example:
[0166] Stable: small standard deviation, indicating little change in pitch.
[0167] Moderate fluctuation: moderate standard deviation, indicating some change in pitch.
[0168] Large fluctuation: large standard deviation, indicating significant change in pitch and greater volatility.
[0169] The system processes the audio signal frame by frame through a pitch detection algorithm (such as autocorrelation or a convolutional neural network-based pitch detection tool) to extract the pitch value (F0) of each frame, forming a pitch sequence. The accuracy of pitch detection depends on the algorithm and sampling frequency. The system aggregates each pitch value in the pitch sequence and calculates its arithmetic mean. The average pitch reflects the overall pitch characteristics of the speaker, and the system judges whether the speaker's pitch is low or high based on this result. The system calculates the deviation of each pitch value from the average pitch and calculates the pitch standard deviation based on the deviation value. The standard deviation reflects the degree of fluctuation in the speaker's pitch throughout the speech and is an important indicator for judging speech stability. The system compares the average pitch attribute value with the pre-set pitch range and classifies it as low, medium or high pitch. Similarly, the system classifies the degree of pitch fluctuation into three levels of stable, moderate fluctuation and large fluctuation according to the range of pitch standard deviation.
[0170] Example: A bank's telephone banking system analyzes customer voice data through the technical solution to extract the customer's average pitch and pitch standard deviation attributes. For example, when the system detects that the customer's pitch standard deviation is large, it is judged that the customer's emotional fluctuation is large and may be in an anxious state. The system will prompt the customer service personnel to slow down the tone or adjust the communication method to improve the customer's satisfaction.
[0171] The embodiment extracts and analyzes the pitch characteristics of the speaker, including the average pitch and pitch standard deviation, through the above steps. It helps to identify the emotional state and voice characteristics of the speaker, and improves the accuracy of the speech recognition system.
[0172] In one embodiment, the above S60 includes:
[0173] S601, applying speech recognition technology to the audio data to identify and extract phoneme sequences;
[0174] S602, determining the number of phonemes in the phoneme sequence to obtain the total number of phonemes in the audio data;
[0175] S603, measuring the total duration of the audio data to determine the audio length;
[0176] S604, analyzing the ratio of the total number of phonemes to the audio length to determine the speaking speed attribute value;
[0177] S605, according to the predetermined ratio range, the speaking speed attribute value is classified into the corresponding speed level.
[0178] In this embodiment, the phoneme sequence is extracted by speech recognition technology, and the speaking speed attribute is calculated according to the ratio of the number of phonemes to the audio length. Speaking speed is an important feature of speech, reflecting the number of phonemes emitted by the speaker in a unit of time. By analyzing the speaking speed attribute, the speaker's speech speed change can be accurately identified, which has important application value in real-time speech processing, emotion analysis and other scenarios.
[0179] The system first applies speech recognition technology to process the input audio data, identify and extract phoneme sequences. Phonemes are the smallest units of pronunciation in speech, and the system converts audio into a series of phonemes by analyzing the speech characteristics of the audio signal.
[0180] The system counts the number of phonemes in the phoneme sequence to obtain the total number of phonemes in the audio data. This step ensures that the system can accurately analyze the actual number of pronunciations by the speaker in a segment of speech.
[0181] The system determines the duration of the audio by analyzing the total length of the audio file. The audio length represents the complete speech time period from the beginning to the end of the speaker.
[0182] The system divides the total number of phonemes by the length of the audio to obtain the speaking speed attribute value. The larger the ratio, the faster the speaker's speech speed; the smaller the ratio, the slower the speaker's speech speed.
[0183] The system classifies the speaking speed attribute value into different speed levels according to the predetermined ratio range.
[0184] For example:
[0185] Slow speed: The ratio is small, indicating that the speaker's speech speed is slow.
[0186] Medium speed: The ratio is moderate, indicating that the speaker's speech speed is moderate.
[0187] Fast speed: The ratio is large, indicating that the speaker's speech speed is fast.
[0188] The system uses speech recognition technology (such as HMM, DNN, etc.) to process audio data and identify the phoneme sequence in the audio. The speech recognition model accurately identifies each phoneme uttered by the speaker by analyzing the frequency, duration, and energy changes of the audio signal. The system counts each phoneme in the phoneme sequence to obtain the total number of phonemes in the audio. Through accurate calculation of the number of phonemes, the system can identify the actual pronunciation volume of the speaker in the speech. The system reads the metadata of the audio file to obtain the total duration of the audio. The length of the audio is one of the key variables for measuring the speaker's speaking speed. The system calculates the number of phonemes per second by dividing the total number of phonemes by the length of the audio. This ratio directly reflects the speaker's speaking speed. The system classifies the speaking speed according to the ratio range, into "slow speed", "medium speed" or "fast speed". For example, when the ratio is less than a certain threshold, the system marks the speech speed as "slow"; when the ratio is higher than a certain threshold, it is marked as "fast".
[0189] This embodiment can accurately calculate the speaking speed attribute of the speaker by analyzing the ratio of the number of phonemes in the audio data to the length of the audio. It has important application value in real-time speech recognition, speech analysis, and customer emotion recognition, especially in the financial field, which can help the system improve the accuracy of voice command processing and transaction confirmation.
[0190] In one embodiment, the above S90 includes:
[0191] S901, obtaining a natural language description template, the natural language description template containing placeholders for keywords of each attribute;
[0192] S902, filling the keywords of each attribute into the placeholders of the natural language description template to generate a structured prompt;
[0193] S903, input the structured prompt into a language model to generate a natural language sentence containing descriptions of all attributes;
[0194] S904, use the natural language sentence as the annotation information of the audio data.
[0195] In this embodiment, structured annotation information is generated based on each attribute of the audio data by designing a natural language description template. The system fills in the keywords of different attributes into the placeholders of the natural language description template to generate a structured prompt. Finally, the system generates a natural language sentence containing descriptions of all attributes through a language model, which is used as the annotation information of the audio data.
[0196] The system first designs a natural language description template, which contains multiple placeholders, each representing an attribute of the audio data. The attributes include gender, speech rate, pitch, signal-to-noise ratio, etc. The keywords of these attributes will replace the placeholders in the template to form a structured natural language description.
[0197] The system selects the corresponding keywords according to the attribute values of each audio data and fills them into the placeholders of the natural language description template. For example, the keywords of the gender attribute may be "male" or "female", and the keywords of the speech rate attribute may be "fast" or "slow".
[0198] After filling in the keywords of all attributes, the system generates a structured prompt. This prompt is composed of multiple attribute keywords and presents a concise and structured description of the audio data.
[0199] The system inputs the generated structured prompt into a language model (such as GPT, etc.), and the model generates complete natural language sentences based on the prompt. These sentences are more fluent and conform to the habits of natural language expression. The sentences contain descriptions of all attributes of the audio data and are presented in natural language, making it easy for users to understand and use.
[0200] Finally, the system stores the generated natural language sentences as the annotation information of the audio data. The annotation information can be used for subsequent queries, analysis, speech recognition, etc., providing structured description content for the system.
[0201] The system designs a natural language description template for each audio attribute, which contains multiple placeholders for attribute keywords. For example: Template sentence: "The audio speaker is {gender}, uses {speech rate} speed, pitch {pitch}, audio quality {signal-to-noise ratio}." Each attribute keyword will be filled into the corresponding placeholder.
[0202] The system selects appropriate keywords according to the attribute values of the audio data. For example, for an audio with a gender attribute of "male", a speech rate attribute of "fast", and a signal-to-noise ratio of "high", the system fills these keywords into the corresponding template placeholders. After all the placeholders are filled, the system generates a structured prompt, for example: "The audio is spoken by a male, with a fast speech rate, moderate pitch, and high audio quality." The system inputs the above structured prompt into the language model to generate a fluent natural language sentence. For example: "This audio is spoken by a male, with a fast speech rate, moderate pitch, and high audio quality." The system stores the generated natural language sentence as the annotation information of the audio data for subsequent retrieval, classification, and analysis.
[0203] Example: In a financial customer service center, the system generates natural language annotation information for each customer's call data through this solution. For example, the system generates the natural language sentence: "The customer is emotionally agitated, with a fast speech rate, and clear audio quality" after analyzing the speech rate, pitch, and emotional fluctuations. This annotation information can help customer service personnel better understand customer needs and adjust service strategies.
[0204] This embodiment generates more readable and understandable annotation information for audio data by designing natural language description templates and generating structured prompts and natural language sentences based on each attribute. These annotation information provides strong support for the query, retrieval, and analysis of audio data, especially in the financial field, which can improve the processing efficiency and accuracy of voice transaction systems and customer service systems.
[0205] In one embodiment, an audio data annotation device is provided, which corresponds to the audio data annotation method in the above embodiments. Referring to Figure 3 , Figure 3 The functional module diagram of a preferred embodiment of the audio data annotation device of the present application is shown. The data acquisition module 10, gender recognition module 20, accent recognition module 30, audio quality analysis module 40, pitch feature analysis module 50, speaking speed analysis module 60, gender and accent classification module 70, attribute mapping module 80, and annotation generation module 90. The detailed description of each functional module is as follows:
[0206] The data acquisition module 10 is used to acquire audio data;
[0207] The gender recognition module 20 is used to determine the gender attribute based on the audio data using a gender classifier;
[0208] The accent recognition module 30 is used to determine the accent attribute based on the audio data using an accent classifier;
[0209] The audio quality analysis module 40 is configured to analyze audio quality of the audio data, obtain a signal-to-noise ratio attribute and an early-to-late reflection ratio attribute.
[0210] The pitch feature analysis module 50 is configured to analyze a pitch feature of the audio data, and obtain a speaker average pitch attribute and a pitch standard deviation attribute.
[0211] The speaking speed analysis module 60 is configured to analyze a ratio of a number of phonemes to a length of audio in the audio data, and determine a speaking speed attribute.
[0212] The gender and accent classification module 70 is configured to take a category name of the gender attribute as a keyword of the gender attribute, and take a category name of the accent attribute as a keyword of the accent attribute.
[0213] The attribute mapping module 80 is configured to map the speaking speed attribute, the signal-to-noise ratio attribute, the early-to-late reflection ratio attribute, the speaker average pitch attribute and the pitch standard deviation attribute into corresponding preset discrete category names, and take the preset discrete category names as keywords of the corresponding attributes respectively.
[0214] The annotation generation module 90 is configured to generate annotation information of the audio data based on the keywords of each attribute.
[0215] In an embodiment, the attribute mapping module 80 is specifically configured to:
[0216] map the speaking speed attribute into discrete category names including fast, medium and slow;
[0217] map the signal-to-noise ratio attribute into discrete category names including high, medium and low;
[0218] map the early-to-late reflection ratio attribute into discrete category names including strong reflection, medium reflection and weak reflection;
[0219] map the speaker average pitch attribute into discrete category names including high pitch, medium pitch and low pitch;
[0220] map the pitch standard deviation attribute into discrete category names including large fluctuation, medium fluctuation and small fluctuation; and
[0221] take the discrete category names as the keywords of the corresponding attributes.
[0222] In an embodiment, the audio quality analysis module 40 is specifically configured to:
[0223] separate the audio data into an effective speech part and a background noise part by using a spectrum analysis module;
[0224] measure powers of the effective speech part and the background noise part respectively.
[0225] determining a signal-to-noise ratio attribute value by analyzing a ratio of power of the effective speech portion and power of the background noise portion;
[0226] classifying the signal-to-noise ratio attribute value into a corresponding audio quality level according to a predetermined signal-to-noise ratio threshold range.
[0227] In an embodiment, the audio quality analysis module 40 is specifically configured to:
[0228] identifying early reflections and late reflections in the audio data using time-frequency analysis techniques;
[0229] measuring energy of the early reflections and the late reflections, respectively;
[0230] determining an early-late reflection ratio attribute value by analyzing a ratio of the energy of the early reflections and the energy of the late reflections;
[0231] classifying the early-late reflection ratio attribute value into a corresponding reverberation degree level according to a predetermined early-late reflection ratio threshold range.
[0232] In an embodiment, the pitch feature analysis module 50 is specifically configured to:
[0233] extracting pitch values from the audio data using a pitch detection tool and recording the pitch values to form a time-varying pitch sequence;
[0234] summarizing all extracted pitch values in the audio data, analyzing an average value of all the pitch values, and determining a mean pitch attribute value;
[0235] analyzing a distribution of the pitch values, calculating a standard deviation of the pitch values, and determining a pitch standard deviation attribute value;
[0236] classifying the mean pitch attribute value into a corresponding pitch level according to a predetermined pitch range.
[0237] classifying the pitch standard deviation attribute value into a corresponding fluctuation degree level according to a predetermined standard deviation range.
[0238] In an embodiment, the speaking speed analysis module 60 is specifically configured to:
[0239] applying speech recognition techniques to the audio data to identify and extract a phoneme sequence;
[0240] determining a number of phonemes in the phoneme sequence to obtain a total number of phonemes in the audio data;
[0241] measuring a total duration of the audio data to determine an audio length;
[0242] analyzing a ratio of the total number of phonemes to the length of the audio to determine a speaking speed attribute value;
[0243] classifying the speaking speed attribute value into a corresponding speed level according to a predetermined ratio range.
[0244] In an embodiment, the annotation generation module 90 is specifically configured to:
[0245] design a natural language description template, the natural language description template containing placeholders of keywords of each attribute;
[0246] filling the keywords of each attribute into the placeholders of the natural language description template to generate a structured prompt;
[0247] inputting the structured prompt into a language model to generate a natural language sentence containing descriptions of all attributes;
[0248] taking the natural language sentence as annotation information of the audio data.
[0249] In one embodiment, a computer device is provided, which can be a server, and an internal structure diagram thereof can be as shown in Figure 4 The computer device includes a processor, a memory, a network interface and a database connected through a system bus. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium, an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The network interface of the computer device is configured to communicate with an external user terminal through a network connection. The computer program is executed by the processor to implement the functions or steps of the server side of the audio data annotation method.
[0250] In one embodiment, a computer device is provided, which can be a user terminal, and an internal structure diagram thereof can be as shown in Figure 5 The computer device includes a processor, a memory, a network interface, a display screen and an input device connected through a system bus. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The network interface of the computer device is configured to communicate with an external server through a network connection. The computer program is executed by the processor to implement the functions or steps of the user terminal side of the audio data annotation method
[0251] In one embodiment, a computer device is provided, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, the processor implementing the following steps when executing the computer program:
[0252] obtaining audio data;
[0253] determining a gender attribute using a gender classifier based on the audio data;
[0254] determining a accent attribute using an accent classifier based on the audio data;
[0255] analyzing audio quality of the audio data to obtain a signal-to-noise ratio attribute and an early-to-late reflection ratio attribute;
[0256] analyzing pitch characteristics of the audio data to obtain a speaker average pitch attribute and a pitch standard deviation attribute;
[0257] determining a speaking speed attribute by analyzing a ratio of a number of phonemes to a length of the audio data in the audio data;
[0258] taking a category name of the gender attribute as a keyword of the gender attribute, and taking a category name of the accent attribute as a keyword of the accent attribute;
[0259] mapping the speaking speed attribute, the signal-to-noise ratio attribute, the early-to-late reflection ratio attribute, the speaker average pitch attribute, and the pitch standard deviation attribute to corresponding preset discrete category names, and taking the preset discrete category names as keywords of the corresponding attributes, respectively;
[0260] generating annotation information of the audio data based on the keywords of each attribute.
[0261] In one embodiment, a computer readable storage medium is provided, having stored thereon a computer program, the computer program being executable by a processor to implement the following steps:
[0262] obtaining audio data;
[0263] determining a gender attribute using a gender classifier based on the audio data;
[0264] determining a accent attribute using an accent classifier based on the audio data;
[0265] analyzing audio quality of the audio data to obtain a signal-to-noise ratio attribute and an early-to-late reflection ratio attribute;
[0266] analyzing pitch characteristics of the audio data to obtain a speaker average pitch attribute and a pitch standard deviation attribute;
[0267] determining a speaking speed attribute by analyzing a ratio of a number of phonemes to a length of the audio data in the audio data;
[0268] the category name of the gender attribute as a keyword of the gender attribute, and the category name of the accent attribute as a keyword of the accent attribute;
[0269] map the speaking speed attribute, the signal-to-noise ratio attribute, the early-to-late reflection ratio attribute, the speaker average pitch attribute, and the pitch standard deviation attribute to corresponding preset discrete category names, and take the preset discrete category names as keywords of the corresponding attributes, respectively;
[0270] generate the labeling information of the audio data based on the keywords of each attribute.
[0271] It should be noted that the functions or steps that can be achieved by the computer readable storage medium or the computer device described above can correspond to the related descriptions of the server side and the user side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0272] Those skilled in the art can understand that all or part of the processes in the foregoing method embodiments can be completed by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer readable storage medium. When the computer program is executed, the processes of the foregoing embodiments of the method can be included. In each embodiment provided in the present application, any reference to a memory, storage, database or other medium can include a non-volatile and / or volatile memory. The non-volatile memory can include a read-only memory (ROM), a programmable ROM (PROM), an electrically programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM) or a flash memory. The volatile memory can include a random access memory (RAM) or an external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0273] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above functional units and modules is exemplified. In actual applications, the above functions can be completed by different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.
[0274] It should be explained that if the software tools or components of other companies appear in the embodiments of the present application, they are only used for example introduction and do not represent actual use. The above-described embodiments are only used to illustrate the technical solutions of the present application, but not limit the present application; although the present application has been described in detail with reference to the foregoing embodiments, it should be understood by those skilled in the art that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features can be replaced by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.
Claims
1. An audio data labeling method, characterized by, The method comprises the following steps: obtaining audio data; determining a gender attribute using a gender classifier based on the audio data; determining a accent attribute using an accent classifier based on the audio data; analyzing the audio quality of the audio data to obtain a signal-to-noise ratio attribute and an early-late reflection ratio attribute; analyzing the pitch characteristics of the audio data to obtain a speaker average pitch attribute and a pitch standard deviation attribute; determining a speaking speed attribute by analyzing the ratio of the number of phonemes to the length of the audio in the audio data; using the category name of the gender attribute as the keyword of the gender attribute, and using the category name of the accent attribute as the keyword of the accent attribute; mapping the speaking speed attribute, the signal-to-noise ratio attribute, the early-late reflection ratio attribute, the speaker average pitch attribute, and the pitch standard deviation attribute to corresponding preset discrete category names, and using the preset discrete category names as the keywords of the corresponding attributes, respectively; obtaining a natural language description template, wherein the natural language description template contains placeholders for the keywords of each attribute; filling the keywords of each attribute into the placeholders in the natural language description template to generate a structured prompt; inputting the structured prompt into a language model to generate a natural language sentence containing descriptions of all attributes; using the natural language sentence as the annotation information of the audio data.
2. The audio data labeling method of claim 1, wherein, mapping the speaking speed attribute, the signal-to-noise ratio attribute, the early-late reflection ratio attribute, the speaker average pitch attribute, and the pitch standard deviation attribute to corresponding preset discrete category names, and using the preset discrete category names as the keywords of the corresponding attributes, respectively, comprising: mapping the speaking speed attribute to discrete category names including fast, medium, and slow; mapping the signal-to-noise ratio attribute to discrete category names including high, medium, and low; mapping the early-late reflection ratio attribute to discrete category names including strong reflection, medium reflection, and weak reflection; mapping the speaker average pitch attribute to discrete category names including high pitch, medium pitch, and low pitch; mapping the pitch standard deviation attribute to discrete category names including large fluctuation, medium fluctuation, and small fluctuation; using the discrete category names as the keywords of the corresponding attributes.
3. The audio data labeling method of claim 1, wherein, analyzing the audio quality of the audio data to obtain a signal-to-noise ratio attribute, comprising: using a spectrum analysis module to separate the audio data into an effective speech part and a background noise part; measuring the power of the effective speech part and the background noise part, respectively; analyzing the ratio of the power of the effective speech part to the power of the background noise part to determine the signal-to-noise ratio attribute value; classifying the signal-to-noise ratio attribute value into corresponding audio quality levels according to a predetermined signal-to-noise ratio threshold range.
4. The audio data labeling method as claimed in claim 1, wherein, analyzing the audio quality of the audio data to obtain an early-late reflection ratio attribute, comprising: using time-frequency analysis techniques to identify early reflections and late reflections in the audio data; measuring the energy of the early reflections and the late reflections, respectively; analyzing the ratio of the energy of the early reflections to the energy of the late reflections to determine the early-late reflection ratio attribute value; classifying the early-late reflection ratio attribute value into corresponding reverberation level categories according to a predetermined early-late reflection ratio threshold range.
5. The audio data labeling method of claim 1, wherein, The pitch feature of the audio data is analyzed to obtain a speaker average pitch attribute and a pitch standard deviation attribute, including: A pitch detection tool is used to extract pitch values from the audio data, and the pitch values are recorded to form a time-varying pitch sequence; All extracted pitch values in the audio data are summarized, and the average of all pitch values is analyzed to determine the average pitch attribute value of the speaker; The distribution of the pitch values is analyzed, and the standard deviation of the pitch values is calculated to determine the pitch standard deviation attribute value; The average pitch attribute value is classified into a corresponding pitch level according to a predetermined pitch range; The pitch standard deviation attribute value is classified into a corresponding fluctuation level according to a predetermined standard deviation range.
6. The audio data labeling method of claim 1, wherein, The ratio of the number of phonemes to the length of the audio in the audio data is analyzed to determine a speaking speed attribute, including: Speech recognition technology is applied to the audio data to identify and extract a phoneme sequence; The number of phonemes in the phoneme sequence is determined to obtain the total number of phonemes in the audio data; The total duration of the audio data is measured to determine the audio length; The ratio of the total number of phonemes to the audio length is analyzed to determine the speaking speed attribute value; The speaking speed attribute value is classified into a corresponding speed level according to a predetermined ratio range.
7. An audio data labeling apparatus, characterized by comprising: The audio data labeling device includes: A data acquisition module for acquiring audio data; A gender recognition module for determining a gender attribute using a gender classifier based on the audio data; An accent recognition module for determining an accent attribute using an accent classifier based on the audio data; An audio quality analysis module for analyzing the audio quality of the audio data to obtain a signal-to-noise ratio attribute and an early-to-late reflection ratio attribute; A pitch feature analysis module for analyzing the pitch feature of the audio data to obtain a speaker average pitch attribute and a pitch standard deviation attribute; A speaking speed analysis module for analyzing the ratio of the number of phonemes to the length of the audio in the audio data to determine a speaking speed attribute; A gender and accent classification module for using the category name of the gender attribute as the keyword of the gender attribute and using the category name of the accent attribute as the keyword of the accent attribute; An attribute mapping module for mapping the speaking speed attribute, the signal-to-noise ratio attribute, the early-to-late reflection ratio attribute, the speaker average pitch attribute, and the pitch standard deviation attribute to corresponding preset discrete category names, and using the preset discrete category names as the keywords of the corresponding attributes, respectively; A label generation module for obtaining a natural language description template containing placeholders for the keywords of each attribute; filling the keywords of each attribute into the placeholders of the natural language description template to generate a structured prompt; inputting the structured prompt into a language model to generate a natural language sentence containing descriptions of all attributes; and using the natural language sentence as the labeling information of the audio data.
8. A computer device, comprising: The computer device includes a memory, a processor, and an audio data labeling program stored on the memory and executable on the processor, which, when executed by the processor, implements the steps of the audio data labeling method of any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The storage medium has stored thereon an audio data labeling program which, when executed by the processor, implements the steps of the audio data labeling method according to any one of claims 1-6.
Citation Information
Patent Citations
Speech recognition device and speech recognition method
JP2015118354A
Localization of prompts
US20070073544A1