Audio correction method and apparatus, electronic device, and storage medium
By using an audio correction model to address the issue of poor audio quality in financial scenarios, this method utilizes a pitch encoder, content encoder, and generator to correct audio features. Combined with a target discriminator to adjust model parameters, it solves the problem of poor audio quality affecting speech recognition, thereby improving both audio quality and speech recognition accuracy.
Patent Information
- Application Number
- CN202411679661.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-21
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2044-11-21
AI Technical Summary
In financial settings, inaccurate pronunciation and intonation in the audio of the person asking the question or responding to the question can lead to poor audio quality, affecting the accuracy of speech recognition and thus impacting consultation services.
By acquiring sample audio and using the original audio correction model composed of pitch encoder, content encoder and generator, pitch features and content features are extracted and corrected. Combined with target discriminator, audio quality is judged and model parameters are adjusted to generate target audio correction model, and finally target audio is corrected.
It improved audio quality, reduced the impact of poor audio quality on consultation services, mitigated the homogenization problem of corrected audio, and improved the accuracy of speech recognition.
Smart Images

Figure CN119517056B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of voice recognition and financial technology, and particularly relates to an audio correction method and device, an electronic device and a storage medium. BACKGROUND
[0002] At present, in a financial scenario, a customer service mode can be used to provide consultation services. In the above scenario, due to the inaccurate pronunciation tone of the speaking audio of the inquiring object or the replying object, the quality of the speaking audio is poor, thereby affecting the accuracy of voice recognition, and further affecting the consultation services. Therefore, how to provide an audio correction method to improve the quality of the audio and further improve the accuracy of voice recognition has become a technical problem to be solved. SUMMARY
[0003] The main purpose of the embodiments of the present application is to provide an audio correction method and device, an electronic device and a storage medium, which aims to improve the quality of the audio and further improve the accuracy of voice recognition.
[0004] To achieve the above purpose, a first aspect of the embodiments of the present application provides an audio correction method, which comprises:
[0005] obtaining a sample audio and a preset original audio correction model; wherein the original audio correction model comprises a pitch encoder, a content encoder and a generator;
[0006] extracting a pitch feature of the sample audio according to the pitch encoder to obtain an original pitch vector;
[0007] performing pitch correction processing on the original pitch vector according to the generator to obtain a corrected pitch vector;
[0008] extracting a content feature of the sample audio according to the content encoder to obtain an original content vector;
[0009] performing content correction processing on the original content vector according to the generator to obtain a corrected content vector;
[0010] performing audio fusion processing on the corrected pitch vector and the corrected content vector according to the generator to obtain a corrected audio;
[0011] performing audio quality discrimination on the corrected audio according to a preset target discriminator to obtain target discrimination data;
[0012] adjusting parameters of the original audio correction model according to the target discrimination data to obtain a target audio correction model;
[0013] According to the target audio correction model, the pre-acquired target audio is corrected.
[0014] In some embodiments, the original audio correction model further comprises a timbre encoder, and before the target discrimination data is obtained by discriminating the corrected audio according to the preset target discriminator, the method further comprises:
[0015] According to the timbre encoder, timbre feature extraction is performed on the sample audio to obtain a timbre vector.
[0016] According to the generator and the timbre vector, timbre correction is performed on the corrected audio.
[0017] In some embodiments, the content feature extraction on the sample audio according to the content encoder to obtain the original content vector comprises:
[0018] According to the content encoder, speech recognition is performed on the sample audio to obtain a text vector.
[0019] According to the content encoder, rhythm feature extraction is performed on the sample audio to obtain a rhythm vector.
[0020] According to the content encoder, sound feature extraction is performed on the sample audio to obtain a sound vector.
[0021] The text vector, the rhythm vector, and the sound vector are spliced to obtain the original content vector.
[0022] In some embodiments, before the target discrimination data is obtained by discriminating the corrected audio according to the preset target discriminator, the method further comprises training the target discriminator, specifically comprising:
[0023] Obtaining training audio;
[0024] Discriminating the training audio according to a preset original discriminator to obtain first training discrimination data;
[0025] At least one of the following adjustments is performed on the training audio: pitch adjustment, content adjustment, to obtain a comparison audio;
[0026] Discriminating the comparison audio according to the original discriminator to obtain second training discrimination data;
[0027] According to the first training discrimination data and the second training discrimination data, the original discriminator is adjusted in parameters to obtain the target discriminator.
[0028] In some embodiments, the sample audio comprises:
[0029] obtaining original audio;
[0030] dimensionally processing the original audio to obtain the sample audio.
[0031] In some embodiments, the dimensionally processing the original audio to obtain the sample audio comprises:
[0032] de-noising the original audio to obtain preliminary audio;
[0033] extracting acoustic parameters from the preliminary audio to obtain mel-frequency spectrum data, and taking the mel-frequency spectrum data as the sample audio.
[0034] In some embodiments, the obtaining original audio comprises:
[0035] obtaining initial audio; wherein the speaking object of the initial audio comprises a target object and a non-target object;
[0036] performing speaking object separation processing on the initial audio to obtain first audio and second audio; wherein the speaking object of the first audio comprises the target object, and the speaking object of the second audio comprises the non-target object;
[0037] filtering out the second audio, and taking the first audio as the original audio.
[0038] To achieve the above object, a second aspect of the embodiments of the present application proposes an audio correction device, which comprises:
[0039] an audio acquisition module, configured to acquire sample audio and a preset original audio correction model; wherein the original audio correction model comprises a pitch encoder, a content encoder and a generator;
[0040] a pitch extraction module, configured to extract pitch features from the sample audio according to the pitch encoder to obtain an original pitch vector;
[0041] a pitch correction module, configured to perform pitch correction processing on the original pitch vector according to the generator to correct a pitch vector;
[0042] a content extraction module, configured to extract content features from the sample audio according to the content encoder to obtain an original content vector;
[0043] a content correction module, configured to perform content correction processing on the original content vector according to the generator to correct a content vector;
[0044] an audio fusion module, configured to perform audio fusion processing on the corrected pitch vector and the corrected content vector according to the generator to obtain a corrected audio.
[0045] The quality discrimination module is used to perform audio quality discrimination on the corrected audio according to a preset target discriminator to obtain target discrimination data;
[0046] The parameter adjustment module is used to adjust the parameters of the original audio correction model according to the target discrimination data to obtain the target audio correction model.
[0047] An audio correction module is used to correct the pre-acquired target audio according to the target audio correction model.
[0048] To achieve the above objectives, a third aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described in the first aspect.
[0049] To achieve the above objectives, a fourth aspect of the present application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in the first aspect.
[0050] The audio correction method, apparatus, electronic device, and storage medium proposed in this application use a target discriminator to judge the audio quality of the corrected audio, obtaining target discrimination data. This target discrimination data allows for parameter adjustment of the original audio correction model, resulting in a target audio correction model with better audio quality correction capabilities. Therefore, when audio correction is performed on target audio based on the target audio correction model, the audio quality of the target audio can be improved. When applied to customer service scenarios in the financial field, this application can reduce the impact of poor target audio quality on consultation services. Furthermore, the target discriminator and the generator in the original audio correction model form a generative adversarial network, thus reducing the problem of severe homogenization of corrected audio compared to related technologies that uniformly correct all audio based on a single correction template. Attached Figure Description
[0051] Figure 1 This is a flowchart of the audio correction method provided in the embodiments of this application;
[0052] Figure 2 yes Figure 1 The flowchart of step S101 in the text;
[0053] Figure 3 yes Figure 2 The flowchart of step S201 in the text;
[0054] Figure 4 yes Figure 2the flowchart of step S202 in
[0055] Figure 5 is Figure 1 the flowchart of step S104 in
[0056] Figure 6 is Figure 1 the flowchart of another embodiment of the audio correction method;
[0057] Figure 7 is Figure 1 the flowchart of another embodiment of the audio correction method;
[0058] Figure 8 is a structural schematic diagram of an audio correction device provided by an embodiment of the present application;
[0059] Figure 9 is a hardware structural schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0060] In order to make the objects, technical solutions and advantages of the present application clearer, the present application is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.
[0061] It should be noted that although the functional modules are divided in the device schematic diagram, and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a manner different from the module division in the device or the order in the flowchart. The terms "first", "second", etc. in the specification, claims and above-described drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence.
[0062] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application, and are not intended to limit the present application.
[0063] First, the terms involved in the present application are analyzed:
[0064] Artificial intelligence (AI): is a new technical science of studying, developing the theory, method, technology and application system for simulating, extending and expanding human intelligence; artificial intelligence is a branch of computer science, artificial intelligence attempts to understand the essence of intelligence, and produce a new intelligent machine that can react in a similar way to human intelligence, the research in this field includes robots, language recognition, image recognition, natural language processing and expert system, etc. Artificial intelligence can simulate the information process of human consciousness and thinking. Artificial intelligence is also the theory, method, technology and application system of using digital computer or digital computer controlled machine to simulate, extend and expand human intelligence, perceive environment, acquire knowledge and use knowledge to obtain the best results.
[0065] Pre-emphasis: is a way of pre-processing in feature extraction. Because the energy of high frequency signal in speech signal is usually low and is greatly affected by suppression, it is necessary to increase the energy of high frequency part, that is, to compensate the amplitude of high frequency part of speech signal, so that the energy distribution of speech signal is more balanced, thereby preventing numerical calculation instability of Fourier transform.
[0066] Frame division: is a way of pre-processing in feature extraction. Speech signal is a non-stationary signal, but considering that the vocal cords move regularly when voicing, that is, the fundamental frequency is relatively fixed in a short time range, therefore, speech signal has short-time stationary characteristics.
[0067] Windowing: is a way of pre-processing in feature extraction. Considering the short-time stationarity of speech signal, each frame of signal is processed by windowing, common window functions include Hamming window, Hanning window, Blackman window, etc. Specifically, after dividing the speech signal into frames, each frame is substituted into the window function, and the values outside the window are set to 0 to eliminate the possible signal discontinuity at both ends of each frame. For example, a Hamming window of size n is added to m frames of signal. By adding a Hamming window to each frame of data, a Hamming window matrix C(m, n) is obtained.
[0068] Fourier transform: refers to the representation of a certain function that meets certain conditions as a linear combination of trigonometric functions (sine or cosine functions) or integrals thereof.
[0069] Mel filter bank: define a filter bank with M filters, where the number of filters is close to the number of critical bands, and the value of M is usually 22-26. The energy spectrum of the speech signal is passed through the triangular filter bank with Mel scale constructed by the above method to extract features from the speech signal.
[0070] Automatic Speech Recognition (ASR): used to convert the lexical content in the object speech into computer-readable input, such as key, binary code or character sequence. The speech recognition system includes four parts of signal processing and feature extraction, acoustic model, language model and decoding search. Among them, the signal processing and feature extraction take the audio signal as the input, enhance the speech by eliminating noise and channel distortion, convert the signal from time domain to frequency domain, and extract suitable representative feature vectors for the acoustic model. The acoustic model is to integrate the knowledge of acoustics and phonetics, and the features generated by the feature extraction part are inputted, and the acoustic model scores for variable-length feature sequences are generated. The language model is used to learn the relationship between words through training corpus, to estimate the possibility of the hypothesis word sequence, which is also called language model score.
[0071] Currently, in the financial and other scenarios, consultation services can be provided through customer service. In the above scenario, due to the fact that the speaking audio of the inquiring object or the replying object may have problems such as inaccurate pronunciation tone, the quality of the speaking audio is poor, thereby affecting the accuracy of speech recognition, and further affecting the consultation service. Therefore, how to provide an audio correction method to improve the accuracy of speech recognition has become a technical problem to be solved.
[0072] It should be noted that in the following embodiments, the application is applied to the intelligent customer service scenario. That is, in the following embodiments, the audio to be corrected is the audio generated by the consultation object, and the system corresponding to the intelligent customer service can perform speech recognition on the corrected audio, and obtain the reply audio or reply text from the corresponding corpus according to the speech recognition result. Among them, the consultation audio includes but is not limited to insurance purchase consultation, financial product consultation, etc. For example, when the consultation audio is to inquire about matter A, but due to the problem of inaccurate pronunciation tone of the inquiring object, the system corresponding to the intelligent customer service recognizes the consultation audio after speech recognition, and the recognized consultation content is to inquire about matter B. At this time, the intelligent customer service will match the reply audio or reply text from the corresponding corpus according to matter B, causing the phenomenon that the reply content does not match the inquiring content of the inquiring object.
[0073] Based on this, the embodiments of the present application provide an audio correction method and device, electronic equipment and storage medium, aiming to correct the consultation audio, improve the audio quality, and enable the intelligent customer service to recognize the content of the consultation audio as inquiring about matter A.
[0074] The audio correction method and device, electronic equipment and storage medium provided by the embodiments of the present application are specifically explained by the following embodiments, and first, the audio correction method in the embodiments of the present application is described.
[0075] The embodiments of the present application can acquire and process related data based on artificial intelligence technology. The artificial intelligence (AI) is a theory, method, technology and application system for simulating, extending and expanding human intelligence by using a digital computer or a machine controlled by a digital computer, perceiving an environment, acquiring knowledge and using the knowledge to obtain optimal results.
[0076] The artificial intelligence basic technology generally includes technologies such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction system, mechatronics, etc. The artificial intelligence software technology mainly includes computer vision technology, robot technology, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning, etc.
[0077] The audio correction method provided by the embodiments of the present application relates to the technical field of speech recognition and financial technology. The audio correction method provided by the embodiments of the present application can be applied in a terminal, can also be applied in a server end, and can also be software running in the terminal or the server end. In some embodiments, the terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, etc.; the server end can be configured as an independent physical server, can also be configured as a server cluster or a distributed system composed of multiple physical servers, can also be configured as a cloud server providing basic cloud computing services such as cloud service, cloud database, cloud computing, cloud function, cloud storage, network service, cloud communication, middleware service, domain name service, security service, CDN, and big data and artificial intelligence platform; and the software can be an application for implementing the audio correction method, but is not limited to the above forms.
[0078] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld devices or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment in which tasks are performed by remote processing devices connected by a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.
[0079] It should be noted that in various specific embodiments of the present application, when relevant processing needs to be performed on data related to the identity or characteristics of the user, such as user information, user behavior data, user history data, and user location information, the user's permission or consent will be obtained first, and the collection, use, and processing of such data will comply with relevant laws, regulations, and standards. In addition, when the embodiments of the present application need to obtain sensitive personal information of the user, the separate permission or separate consent of the user will be obtained through a pop-up window or a jump to a confirmation page, and after obtaining the separate permission or separate consent of the user, the necessary user-related data for the normal operation of the embodiments of the present application will be obtained.
[0080] Figure 1 is an optional flowchart of an audio correction method provided by the embodiments of the present application, Figure 1 The method in the method can include but is not limited to steps S101-S109.
[0081] Step S101, obtaining a sample audio and a preset original audio correction model; wherein the original audio correction model includes a pitch encoder, a content encoder, and a generator;
[0082] Step S102, performing pitch feature extraction on the sample audio according to the pitch encoder to obtain an original pitch vector;
[0083] Step S103, performing pitch correction processing on the original pitch vector according to the generator to correct the pitch vector;
[0084] Step S104, performing content feature extraction on the sample audio according to the content encoder to obtain an original content vector;
[0085] Step S105, performing content correction processing on the original content vector according to the generator to correct the content vector;
[0086] Step S106, performing audio fusion processing on the corrected pitch vector and the corrected content vector according to the generator to obtain a corrected audio;
[0087] Step S107, performing audio quality discrimination on the corrected audio according to a preset target discriminator to obtain target discrimination data;
[0088] Step S108, adjusting parameters of the original audio correction model according to the target discrimination data to obtain a target audio correction model;
[0089] Step S109, performing audio correction on a pre-acquired target audio according to the target audio correction model.
[0090] The steps S101 to S109 shown in the embodiments of the present application are used to determine the audio quality of the modified audio by the target discriminator, to obtain target discrimination data, so that the target audio correction model with better audio quality correction capability can be obtained after the original audio correction model is adjusted according to the target discrimination data. Therefore, when the target audio is corrected according to the target audio correction model, the audio quality of the target audio can be improved. When the present application is applied to the customer service scene in the financial field, the impact on the consultation service caused by the poor quality of the target audio can be reduced. In addition, the target discriminator and the generator in the original audio correction model constitute a generative adversarial network (GAN), so compared with the method of uniformly correcting all audios according to one correction template in the related art, the problem of serious homogenization of the corrected audios can be reduced.
[0091] In step S101 of some embodiments, a sample audio is obtained, which is an audio with poor audio quality. For example, when the present application is applied to the customer service scene in the financial field, the sample audio can be the consultation audio of the consultation object. Since the consultation object may have problems such as accent and inaccurate pronunciation tone, the audio quality of the sample audio is poor. When speech recognition is performed according to the sample audio with poor quality to realize intelligent customer service reply, the phenomenon of inaccurate speech recognition and incorrect reply may occur. For example, the ideal consultation audio of the consultation object is “understand vehicle insurance”, but due to the problems such as accent and inaccurate pronunciation tone of the consultation object, the actual consultation audio of the consultation object is “understand vehicle scrap”. The intelligent customer service will generate different reply contents according to the above two consultation audios, such as for the ideal consultation audio, the reply content includes which vehicle insurance, how much is the premium of each vehicle insurance, and how much is the claim rate, etc. However, in the actual situation, the audio recognized by the intelligent customer service is “understand vehicle scrap”, at this time the reply content includes the judgment standard of vehicle scrap. Therefore, the audio with poor quality needs to be corrected. In the embodiments of the present application, the original audio correction model is used for correction processing, and the method of correction processing will be described in other embodiments. It should be noted that the sample audio can be obtained through a corresponding application programming interface (API), such as through a customer service application API interface to obtain the sample audio. Alternatively, in order to enrich the training samples of the original audio correction model, other audios with poor quality can also be obtained from big data as sample audios, which are not limited in the embodiments of the present application.
[0092] In addition, the embodiments of the present application also need to obtain a pre-set original audio correction model, which includes a pitch encoder, a content encoder and a generator.
[0093] Referring to Figure 2 In some embodiments, the "obtaining sample audio" in step S101 includes but is not limited to steps S201 to S202.
[0094] Step S201, obtaining original audio;
[0095] Step S202, performing dimensionality increasing processing on the original audio to obtain the sample audio.
[0096] In step S201 of some embodiments, the original audio refers to audio with poor quality, for example, when the present application is applied to the customer service scenario in the financial field, the original audio can be the consultation audio of the consultation object. Since the consultation object may have problems such as accent and inaccurate pronunciation tone, the original audio has poor quality. The original audio can be obtained through a corresponding API interface, such as obtaining the original audio through a customer service application API interface. Alternatively, in order to enrich the training samples of the original audio correction model, other audio with poor quality can also be obtained from big data as the original audio, which is not limited in the embodiments of the present application.
[0097] Referring to Figure 3 In some embodiments, step S201 includes but is not limited to steps S301 to S303.
[0098] Step S301, obtaining initial audio; wherein the speaking object of the initial audio includes a target object and a non-target object;
[0099] Step S302, performing speaking object separation processing on the initial audio to obtain first audio and second audio; wherein the speaking object of the first audio includes the target object, and the speaking object of the second audio includes the non-target object;
[0100] Step S303, filtering out the second audio, and taking the first audio as the original audio.
[0101] In step S301 of some embodiments, the initial audio refers to audio with poor audio quality. For example, when the application is applied to a customer service scenario in the financial field, the initial audio can be the consultation audio of the consultation object. Since the consultation object may have problems such as accent and inaccurate pronunciation tone, the audio quality of the initial audio is poor. In addition, the initial audio is a combined audio of multiple speaking objects, that is, the initial audio includes not only the audio of the target object but also the audio of the non-target object. For example, in the customer service scenario in the financial field, when the target object performs a consultation operation in an indoor scene, the target object is an adult and the non-target object is a child. Or, when the target object performs a consultation operation in an outdoor scene, the non-target object is a passerby, and the like. It can be understood that the format of the initial audio can be any one of the following: MP3, WAV, AAC, FLAC, and the like, and the embodiments of the application are not limited in this regard.
[0102] In steps S302 to S303 of some embodiments, in order to improve the accuracy of audio quality correction and reduce the amount of calculation when correcting the audio quality, the initial audio is subjected to speaking object separation processing, that is, the audio of the target object is separated from the audio of the non-target object to obtain a first audio including only the speaking audio of the target object and a second audio including only the speaking audio of the non-target object. The second audio is filtered out, and the first audio is used as the original audio. The speaking object separation processing method can use any one of the following: a blind source separation (BSS) method, a speech segmentation method, a deep learning-based method, a time-frequency analysis-based method, a statistical modeling-based method, and the like. The blind source separation-based method is based on blind source analysis theory and performs independent component analysis (ICA) on the original audio. In the deep learning-based method, a deep neural network (DNN) or a convolutional neural network (CNN) can be used. In the time-frequency analysis-based method, the initial audio is analyzed from both the time domain and the frequency domain. In the statistical modeling-based method, a Gaussian mixture model (GMM) or a hidden Markov model (HMM) can be used to statistically model the initial audio. It can be understood that the determination method of the target object and the non-target object is related to the speaking object separation processing method, and the embodiments of the application will not be described in detail.
[0103] In step S202 of some embodiments, since the obtained original audio is one-dimensional data, in order to be able to obtain more audio features from the original audio, the original audio data is subjected to dimensionality increasing processing to obtain a sample audio with higher dimensionality. For example, a two-dimensional sample audio is obtained.
[0104] Reference Figure 4 In some embodiments, step S202 includes but is not limited to steps S401 to S402.
[0105] In step S401, the original audio is denoised to obtain preliminary audio.
[0106] In step S402, acoustic parameters are extracted from the preliminary audio to obtain mel-frequency spectrum data, which is used as sample audio.
[0107] In step S401 of some embodiments, in order to improve the accuracy of audio correction, the original audio is denoised, that is, the noise in the original audio is filtered out, to obtain preliminary audio. It can be understood that the difference between denoising and speaker separation processing is that speaker separation processing separates audio corresponding to different speakers based on the speaker as the separation reference. The denoising process separates the target speaker's speech audio from the background noise based on the target speaker's speech audio and the background noise as the separation reference. The background noise includes wind noise, siren noise, signal noise, etc. The denoising method can use any of the following methods: a frequency domain filtering-based method, a time domain filtering-based method, a deep learning-based method, and a statistical modeling-based method. In the frequency domain filtering-based method, the original audio is converted from the time domain to the frequency domain by performing short-time Fourier transform or wavelet transform on the original audio, and then the frequency domain signal is filtered and denoised. The frequency domain filtering method includes spectral subtraction, minimum mean square error (MMSE) method, etc. The time domain filtering-based method includes Kalman filtering, adaptive noise estimation (ANE), etc.
[0108] The filtering-based method, the spectral subtraction-based method, the wavelet transform-based method, the deep learning-based method, etc.
[0109] In step S402 of some embodiments, since the preliminary audio is one-dimensional time series data, in order to obtain more audio features from the initial audio, acoustic parameters are extracted from the preliminary audio to obtain mel-frequency spectrum data (MFCC) representing two-dimensional frequency domain. The mel-frequency spectrum data is used as sample audio. The acoustic parameter extraction includes the following steps: pre-emphasis, framing, windowing, Fourier transform, mel filter bank, etc. It can be understood that in addition to mel-frequency spectrum data (MFCC), two-dimensional frequency domain data in the following forms can also be extracted from the initial audio: linear prediction system (LPC), linear prediction cepstrum coefficient (LPCC), line spectrum frequency (LSF), discrete wavelet transform (SWT).
[0110] In other embodiments, the preliminary audio can also be converted into Fourier spectrum using the function in the audio processing library (librosa library), and the Fourier spectrum can be converted into mel-frequency spectrum data using other functions, thereby obtaining sample audio.
[0111] The mel-spectrum data is used as the sample audio in the embodiments of the present application, and the mel-spectrum data has better noise resistance and robustness than other two-dimensional frequency domain data, and can improve the training effect of the original audio correction model according to the mel-spectrum data.
[0112] In step S102 of some embodiments, the pitch refers to the vocal range used by the subject when speaking, and the pitch determines the vibration frequency of the sound production body. The higher the frequency, the higher the pitch. The pitch affects the accuracy and clarity of pronunciation, thereby affecting the audio quality. The sample audio is used as the input data of the pitch encoder, and the pitch feature of the sample audio is extracted according to the pitch encoder to obtain an original pitch vector.
[0113] In step S103 of some embodiments, since the original pitch vector may be a factor affecting the poor audio quality of the sample audio, the original pitch vector is used as the input data of the generator, and the original pitch vector is corrected by the generator to obtain a corrected pitch vector.
[0114] In step S104 of some embodiments, the sample audio is used as the input data of the content encoder, and the content feature of the sample audio is extracted according to the content encoder to obtain an original content vector. The content feature refers to the feature of the audio characteristics of the sample audio, such as the speaking content represented by the sample audio, the audio rhythm of the sample audio, etc. The accuracy of the content feature affects the accuracy and clarity of pronunciation, thereby affecting the audio quality.
[0115] Reference Figure 5 In some embodiments, step S104 includes but is not limited to steps S501 to S504.
[0116] Step S501: performing speech recognition on the sample audio according to the content encoder to obtain a text vector;
[0117] Step S502: performing rhythm feature extraction on the sample audio according to the content encoder to obtain a rhythm vector;
[0118] Step S503: performing sound production feature extraction on the sample audio according to the content encoder to obtain a sound production vector;
[0119] Step S504: vector splicing the text vector, the rhythm vector, and the sound production vector to obtain an original content vector.
[0120] In step S501 of some embodiments, the sample audio is subjected to speech recognition according to the content encoder to determine the speaking content represented by the sample audio, and a text vector is obtained.
[0121] In step S502 of some embodiments, the speaking rhythm is related to factors such as accent, tone, and tone of voice, so the speaking rhythm can affect the accuracy and clarity of pronunciation, that is, affect the audio quality. The content encoder extracts the rhythm features of the sample audio to obtain a rhythm vector. The rhythm vector is used to represent information such as speaking speed, pause, and connected reading.
[0122] In step S503 of some embodiments, speaking pronunciation refers to the way of pronouncing the content of speaking, the way of pronouncing words, and the like. For example, different objects may have different pronunciations for the same word, such as pronouncing “bao fei” as “bao fei” for some objects. As can be seen, the speaking pronunciation can affect the accuracy and clarity of pronunciation, that is, affect the audio quality. The content encoder extracts the pronunciation features of the sample audio to determine the pronunciation characteristics of the corresponding object, and obtains a pronunciation vector.
[0123] In step S504 of some embodiments, the text vector, the rhythm vector, and the pronunciation vector are spliced to obtain an original content vector that can represent the text vector, the rhythm vector, and the pronunciation vector as a whole.
[0124] In step S105 of some embodiments, since the original content vector can be a factor that causes poor audio quality of the sample audio, the original content vector is used as input data of the generator, and the original content vector is corrected by the generator to obtain a corrected content vector.
[0125] In step S106 of some embodiments, the generator is also used for audio fusion processing of the corrected pitch vector and the corrected content vector to obtain a new audio, that is, a corrected audio.
[0126] Reference Figure 6 In some embodiments, the original audio correction model also includes a timbre encoder. Before step S107, the method provided by the embodiments of the present application further includes but is not limited to steps S601 to S602.
[0127] Step S601: Extracting timbre features of the sample audio according to the timbre encoder to obtain a timbre vector;
[0128] Step S602: Correcting the timbre of the corrected audio according to the generator and the timbre vector.
[0129] In step S601 of some embodiments, in order to make the original audio correction model only correct the audio quality of the sample audio without changing the speaking characteristics of the corresponding object, the original audio correction model further includes a timbre encoder. The sample audio is used as input data of the timbre encoder, and the timbre features of the sample audio are extracted by the timbre encoder to obtain a timbre vector representing the speaking characteristics of the object.
[0130] In step S602 of some embodiments, the timbre vector is taken as input data of the generator, so as to perform timbre correction on the generated modified audio according to the timbre vector, so that the timbre of the modified audio conforms to the timbre corresponding to the timbre vector, that is, the modified audio retains the speaking characteristics of the corresponding object, and reduces the problem of timbre homogenization of the modified audio.
[0131] Referring to Figure 7 In some embodiments, before step S107, the method provided by the embodiments of the present application further includes training the target discriminator, including but not limited to steps S701 to S705.
[0132] In step S701, training audio is obtained.
[0133] In step S702, the training audio is subjected to audio quality discrimination according to a preset original discriminator, to obtain first training discrimination data.
[0134] In step S703, the training audio is subjected to at least one of the following adjustments: pitch adjustment, content adjustment, to obtain comparative audio.
[0135] In step S704, the comparative audio is subjected to audio quality discrimination according to the original discriminator, to obtain second training discrimination data.
[0136] In step S705, the original discriminator is subjected to parameter adjustment according to the first training discrimination data and the second training discrimination data, to obtain the target discriminator.
[0137] In step S701 of some embodiments, the training audio is obtained, which is audio used for training the target discriminator. Therefore, in order to enable the target discriminator to have the ability to discriminate audio quality, the training audio should be audio of good audio quality. It can be understood that when the expression content of an audio is clear, the pronunciation tone is accurate, etc., the audio is taken as audio of good audio quality; otherwise, the audio is taken as audio of poor audio quality. The training audio can be obtained through a corresponding API interface, or generated by a professional reading object, which is not limited in the embodiments of the present application.
[0138] In step S702 of some embodiments, the training audio is taken as input data of the original discriminator, and the audio quality of the training audio is discriminated by the original discriminator, to obtain first training discrimination data.
[0139] In step S703 of some embodiments, in order to enable the target discriminator to further have the ability to discriminate audio quality difference, the audio quality of the training audio is adjusted to obtain audio representing audio quality difference, i.e., contrast audio. The audio quality adjustment can be performed in at least one of the following ways: pitch adjustment, content adjustment, which includes rhythm adjustment, voice adjustment, etc. Since the pitch and content will affect the accuracy and clarity of pronunciation, the audio quality of the training audio can be changed according to the above adjustment methods.
[0140] In step S704 of some embodiments, the contrast audio is taken as input data of the original discriminator, and the audio quality of the contrast audio is discriminated by the original discriminator to obtain second training discrimination data.
[0141] In step S705 of some embodiments, the original discriminator is adjusted in parameters according to the first training discrimination data and the second discrimination data, respectively, to reduce the discrimination error of the original discriminator for audio of good audio quality and to reduce the discrimination error of the original discriminator for audio of poor audio quality, so as to obtain the target discriminator having the ability to discriminate audio quality good and audio quality difference.
[0142] In step S107 of some embodiments, the target discriminator capable of discriminating audio quality is pre-set, and the modified audio is taken as input data of the target discriminator to obtain target discrimination data representing audio quality good or audio quality difference. It can be understood that in some embodiments, the modified audio input to the target discriminator is the audio obtained after timbre modification.
[0143] In step S108 of some embodiments, it can be determined according to the target discrimination data whether the audio quality of the modified audio is improved compared with the audio quality of the sample audio. For example, when the target discrimination data represents audio quality difference, it indicates that the audio quality of the modified audio is not improved or the improvement degree is small; when the target discrimination data represents audio quality good, it indicates that the audio quality of the modified audio is improved. Therefore, the original audio modification model can be adjusted in parameters according to the target discrimination data to improve the modification ability of the original audio modification model for audio quality, so as to obtain a target audio modification model. It can be understood that when the target discrimination data is less than a preset threshold, it is determined that the target discrimination data represents audio quality difference; when the target discrimination data is greater than or equal to the preset threshold, it is determined that the target discrimination data represents audio quality good.
[0144] As can be seen from the above description, the modification purpose of the embodiments of the present application is to improve the audio quality of the modified audio compared with the audio before modification, without requiring the modified audio to meet a unified modification standard. The advantage of this is that it can reduce the problem of serious homogenization of the modified audio, i.e., it can preserve the speaking characteristics of the corresponding object on the basis of ensuring the accuracy and clarity of the modified audio.
[0145] In step S109 of some embodiments, the target audio is audio that needs to be corrected in actual application. For example, in a customer service scenario in the financial field, a customer service application is loaded on a terminal. The consultation object sends audio through a voice control of a consultation interface on the terminal, and the audio that can be obtained is taken as the target audio. The target audio is taken as input data of the target audio correction model, so that the target audio is subjected to audio quality correction processing through the target audio correction model. According to the corrected audio, downstream tasks such as speech recognition and intelligent customer service reply can be performed, thereby improving the accuracy of speech recognition and intelligent customer service reply. The speech recognition method includes any one of the following: a speech model-based recognition method, a neural network-based recognition method, an end-to-end-based recognition, etc. In the speech model-based recognition method, the speech model includes a hidden Markov model (HMM), a conditional random field (CRF), etc. In the neural network-based recognition method, the neural network includes a recurrent neural network (RNN), a convolutional neural network (CNN), etc.
[0146] In some embodiments, the downstream task can also include real-time translation, online consultation in a medical scenario, etc., for which the embodiments of the present application are not limited.
[0147] The audio correction method provided by the embodiments of the present application discriminates the audio quality of the corrected audio through the target discriminator, obtains target discrimination data, so that after the original audio correction module is adjusted in parameters according to the target discrimination data, a target audio correction model with better audio quality correction capability can be obtained. Therefore, when the target audio is subjected to audio correction according to the target audio correction model, the audio quality of the target audio can be improved. When the present application is applied to a customer service scenario in the financial field, the impact on consultation services caused by poor target audio quality can be reduced. In addition, the target discriminator and the generator in the original audio correction model constitute a generative adversarial network, so compared with the method of uniformly correcting all audios according to one correction template in the related art, the problem of serious homogenization of the corrected audios can be reduced.
[0148] In some specific embodiments, the consultant sends a target audio "understand vehicle scrapping" through the customer service window of the insurance APP. The target audio is obtained, and the target audio is processed by dimensionality reduction processing, object separation processing, etc. The processed audio is used as input data of the target audio correction model to correct the target audio through the target audio correction model to obtain a corrected audio "understand vehicle premium". The insurance APP calls a corpus stored locally or in the cloud according to the corrected audio, and performs semantic matching on "understand vehicle premium" and the corpus to obtain corresponding reply content. The reply content includes but is not limited to: what is vehicle insurance, how much is the premium of each type of vehicle insurance, and what is the claim rate, etc. The reply content is displayed in the customer service window in the form of text or voice, thereby realizing intelligent reply to the consultant and improving the accuracy of the reply.
[0149] Reference Figure 8 The embodiments of the present application also provide an audio correction device, which can implement the above-mentioned audio correction method. The device comprises:
[0150] An audio acquisition module 801 is configured to acquire a sample audio and a preset original audio correction model. The original audio correction model comprises a pitch encoder, a content encoder and a generator.
[0151] A pitch extraction module 802 is configured to extract a pitch feature of the sample audio according to the pitch encoder to obtain an original pitch vector.
[0152] A pitch correction module 803 is configured to perform pitch correction processing on the original pitch vector according to the generator to obtain a corrected pitch vector.
[0153] A content extraction module 804 is configured to extract a content feature of the sample audio according to the content encoder to obtain an original content vector.
[0154] A content correction module 805 is configured to perform content correction processing on the original content vector according to the generator to obtain a corrected content vector.
[0155] An audio fusion module 806 is configured to perform audio fusion processing on the corrected pitch vector and the corrected content vector according to the generator to obtain a corrected audio.
[0156] A quality discrimination module 807 is configured to perform audio quality discrimination on the corrected audio according to a preset target discriminator to obtain target discrimination data.
[0157] A parameter adjustment module 808 is configured to perform parameter adjustment on the original audio correction model according to the target discrimination data to obtain a target audio correction model.
[0158] The audio correction module 809 is configured to perform audio correction on the pre-acquired target audio according to the target audio correction model.
[0159] The specific implementation of the audio correction apparatus is basically the same as the specific implementation of the above-described audio correction method, and thus will not be described here again.
[0160] The embodiments of the present application further provide an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor implements the above-described audio correction method when executing the computer program. The electronic device can be any smart terminal, such as a tablet computer or a vehicle-mounted computer.
[0161] With reference to Figure 9 , Figure 9 The hardware structure of the electronic device of another embodiment is illustrated, which includes:
[0162] The processor 901 can be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits, and is configured to execute related programs to implement the technical solutions provided by the embodiments of the present application.
[0163] The memory 902 can be implemented in the form of a ROM (ReadOnly Memory), a static storage device, a dynamic storage device, or a RAM (Random Access Memory). The memory 902 can store an operating system and other application programs. When the technical solutions provided by the embodiments of the present application are implemented by software or firmware, the related program codes are stored in the memory 902 and are called and executed by the processor 901 to implement the audio correction method of the embodiments of the present application.
[0164] The input / output interface 903 is configured to implement information input and output.
[0165] The communication interface 904 is configured to implement the communication interaction between the device and other devices. The communication can be implemented in a wired manner (for example, USB, network cable, etc.) or in a wireless manner (for example, mobile network, WIFI, Bluetooth, etc.).
[0166] The bus 905 is configured to transmit information between various components (for example, the processor 901, the memory 902, the input / output interface 903, and the communication interface 904) of the device.
[0167] The processor 901, the memory 902, the input / output interface 903, and the communication interface 904 are communicatively connected with each other through the bus 905.
[0168] The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the audio correction method.
[0169] The memory, as a non-transitory computer readable storage medium, can be used to store non-transitory software programs and non-transitory computer executable programs. In addition, the memory can include a high-speed random access memory, and can also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state memory device. In some embodiments, the memory can optionally include a memory remotely arranged relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.
[0170] The audio correction method, the audio correction device, the electronic device, and the storage medium provided by the embodiments of the present application can determine the audio quality of the corrected audio through the target discriminator, and obtain target discrimination data, so that the original audio correction module can be adjusted in parameters according to the target discrimination data, and a target audio correction model with better audio quality correction capability can be obtained. Therefore, when the target audio is corrected according to the target audio correction model, the audio quality of the target audio can be improved. When the present application is applied to the customer service scene in the financial field, the influence on the consultation service caused by the poor quality of the target audio can be reduced. In addition, the target discriminator and the generator in the original audio correction model constitute a generative adversarial network, so compared with the method of uniformly correcting all audios according to one correction template in the related art, the problem of serious homogenization of the corrected audios can be reduced.
[0171] The embodiments described in the embodiments of the present application are used to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art can know that, with the evolution of technology and the appearance of new application scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.
[0172] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and can include more or fewer steps than the figures shown, or combine certain steps, or different steps.
[0173] The apparatus embodiments described above are merely exemplary, and the units described as separate units can or can not be physically separate, i.e., can be located in one place, or can be distributed over multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the embodiment.
[0174] Those skilled in the art can understand that all or some of the steps in the method disclosed above, the functional modules / units in the system and the device can be implemented as software, firmware, hardware and appropriate combinations thereof.
[0175] The terms "first", "second", "third", "fourth" and the like in the description of the application and in the claims of the foregoing drawings, if any, are used for distinguishing between similar objects and not necessarily for describing a particular sequential or chronological order. It is to be understood that the use of the terms so
[0176] It should be understood that in this application, "at least one" means one or more, and "multiple" means two or more. "And / or" is used to describe the relationship between the associated objects, which means that there can be three relationships, for example, "A and / or B" can mean that there are three cases: only A, only B, and A and B at the same time, where A and B can be singular or plural. The character " / " generally represents that the associated objects before and after are in an "or" relationship. "At least one of the following" or the like means any combination of these items, including any combination of single or multiple items. For example, at least one of a, b or c, can mean a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0177] In several embodiments provided in the present application, it should be understood that the disclosed apparatus and method can be implemented by other manners. For example, the apparatus embodiments described above are merely illustrative, for example, the division of the above units is merely a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units or components shown or discussed can be indirect coupling or communication connection through some interfaces, apparatuses or units, and can be electrical, mechanical or other forms.
[0178] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they can be located in one place or distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.
[0179] In addition, the functional units in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.
[0180] If the integrated unit is realized in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the part of the prior art that makes a contribution or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, including a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method of each embodiment of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various program storage media.
[0181] The preferred embodiments of the embodiments of the present application are described above with reference to the accompanying drawings, but this does not limit the scope of the embodiments of the present application. Any modifications, equivalent replacements and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of the embodiments of the present application.
Claims
1. An audio correction method, characterized by, The method comprises the following steps: obtain a sample audio and a preset original audio correction model; wherein the original audio correction model comprises a pitch encoder, a content encoder and a generator; extract a pitch feature from the sample audio according to the pitch encoder to obtain an original pitch vector; perform pitch correction processing on the original pitch vector according to the generator to obtain a corrected pitch vector; extract a content feature from the sample audio according to the content encoder to obtain an original content vector; perform content correction processing on the original content vector according to the generator to obtain a corrected content vector; perform audio fusion processing on the corrected pitch vector and the corrected content vector according to the generator to obtain a corrected audio; perform audio quality discrimination on the corrected audio according to a preset target discriminator to obtain target discrimination data; adjust parameters of the original audio correction model according to the target discrimination data to obtain a target audio correction model; perform audio correction on a pre-obtained target audio according to the target audio correction model.
2. The method of claim 1, wherein, The original audio correction model further comprises a timbre encoder, and before the step of performing audio quality discrimination on the corrected audio according to a preset target discriminator to obtain target discrimination data, the method further comprises: extract a timbre feature from the sample audio according to the timbre encoder to obtain a timbre vector; perform timbre correction on the corrected audio according to the generator and the timbre vector.
3. The method of claim 1, wherein, The step of extracting a content feature from the sample audio according to the content encoder to obtain an original content vector comprises: perform speech recognition on the sample audio according to the content encoder to obtain a text vector; extract a rhythm feature from the sample audio according to the content encoder to obtain a rhythm vector; extract a vocalization feature from the sample audio according to the content encoder to obtain a vocalization vector; concatenate the text vector, the rhythm vector and the vocalization vector to obtain the original content vector.
4. The method of claim 1, wherein, Before the step of performing audio quality discrimination on the corrected audio according to a preset target discriminator to obtain target discrimination data, the method further comprises training the target discriminator, specifically comprising: obtain training audio; perform audio quality discrimination on the training audio according to a preset original discriminator to obtain first training discrimination data; perform at least one of the following adjustments on the training audio: pitch adjustment, content adjustment, to obtain comparison audio; perform audio quality discrimination on the comparison audio according to the original discriminator to obtain second training discrimination data; adjust parameters of the original discriminator according to the first training discrimination data and the second training discrimination data to obtain the target discriminator.
5. The method according to any one of claims 1 to 4, characterized in that, The step of obtaining a sample audio comprises: obtain an original audio; perform dimensionality increasing processing on the original audio to obtain the sample audio.
6. The method of claim 5, wherein, The step of performing dimensionality increasing processing on the original audio to obtain the sample audio comprises: perform denoising processing on the original audio to obtain a preliminary audio; perform acoustic parameter extraction on the preliminary audio to obtain mel-frequency spectrum data, and use the mel-frequency spectrum data as the sample audio.
7. The method of claim 5, wherein, The obtaining the original audio comprises: obtaining initial audio; wherein the speaking object of the initial audio comprises a target object and a non-target object; performing speaking object separation processing on the initial audio to obtain first audio and second audio; wherein the speaking object of the first audio comprises the target object, and the speaking object of the second audio comprises the non-target object; filtering out the second audio and taking the first audio as the original audio.
8. An audio correction device, characterized by The device comprises: an audio acquisition module configured to acquire sample audio and a preset original audio correction model; wherein the original audio correction model comprises a pitch encoder, a content encoder, and a generator; a pitch extraction module configured to perform pitch feature extraction on the sample audio according to the pitch encoder to obtain an original pitch vector; a pitch correction module configured to perform pitch correction processing on the original pitch vector according to the generator to correct the pitch vector; a content extraction module configured to perform content feature extraction on the sample audio according to the content encoder to obtain an original content vector; a content correction module configured to perform content correction processing on the original content vector according to the generator to correct the content vector; an audio fusion module configured to perform audio fusion processing on the corrected pitch vector and the corrected content vector according to the generator to obtain a corrected audio; a quality discrimination module configured to perform audio quality discrimination on the corrected audio according to a preset target discriminator to obtain target discrimination data; a parameter adjustment module configured to perform parameter adjustment on the original audio correction model according to the target discrimination data to obtain a target audio correction model; an audio correction module configured to perform audio correction on a target audio obtained in advance according to the target audio correction model.
9. An electronic device, comprising: The electronic device comprises a memory and a processor, the memory stores a computer program, and the processor implements the method of any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium storing a computer program, the computer program comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 9. The computer program is executed by the processor to implement the method of any one of claims 1 to 7.
Citation Information
Patent Citations
Audio signal generation method, device and equipment and storage medium
CN112712812A
Audio processing method and device, electronic equipment and readable storage medium
CN113470699A